From Prompts to Systems That Act
Building real automation with LLMs, VLMs, and VLAs
AI systems are no longer just responding. They are planning, acting, and operating in environments where the cost of getting it wrong is real.
Last week, Data & AI Stockholm brought the community together for an evening at Proxify around exactly this question: what does it actually take to build these systems, not in demos, but in practice?
We are grateful to Proxify for hosting us. The venue carried that distinctive Stockholm character, older architectural details paired with a warm, modern interior, and the whole team made the evening feel genuinely welcoming. As always, much of the value happened between and after the talks, in the conversations that do not make it into any agenda.
From Seeing to Acting
The evening opened with Joana Fonseca, AI Engineer at TRATON, who works on vision-language models (VLMs) and vision-language-action models (VLAs) for autonomous driving. Beyond her work at TRATON, Joana is a board member of Stockholm AI, a non-profit community for AI professionals and researchers in Stockholm.
Before the technical content, she did something simple but effective:
She asked the room how they felt about using these systems in real-world environments.
The options covered a range, and the responses reflected it. Most people landed on optimism about the potential, with safety as the primary concern, particularly in high-stakes domains like autonomous driving. A notable share also selected the more cautious position: that these models should not be the default answer to every problem simply because they are available.
That framing set up the core challenge well.
She opened with two concrete problems her team works with.
First, how do you identify the meaningful or high-value segments out of millions of hours of raw driving data?
Second, how do you scale the system without needing to gather data for every single corner case? These are not framing questions.
They are active engineering challenges that shape how the whole system gets designed.
Traditional approaches hit a ceiling here. Rule-based systems can only handle what was anticipated. Data-driven models improve with volume, but remain limited to what they have seen. VLMs and VLAs change this by combining visual understanding, language reasoning, and action outputs, allowing systems to generalize to new situations rather than pattern-match to known ones.
The Q&A sharpened this further, with several questions comparing these approaches to classical computer vision pipelines and probing where generalization actually holds under real operating conditions.
By the end of the session, after walking through what these systems are actually capable of, she ran the pulse check again. The room had not dramatically shifted, but it had become more grounded. Optimistic about the potential, and more precise about what safety-conscious deployment actually requires.
From Requirements to Executable Systems
Some of you might remember Erisa Dervishi from our previous session at Spotify, where she and her colleague walked us through Scaling Metrics at Spotify: The Story Behind Data Engineering Challenges. This time, back representing Spotify, she shared the team’s journey of experimenting with AI-driven planning and what they actually learned along the way.
The starting point was straightforward. Give the model a high-level summary, some context, a few references, and let it break the work down.
The instinct that followed was equally natural: when the output was not good enough, add more. More context, more instructions, more detail.
The results fell short.
More input did not produce more clarity. It produced noisier output, incorrectly bundled tasks, blurred boundaries, and a model that was harder to steer, not easier. The problem was not the model’s capability. It was the absence of structure around it.
The team stepped back and redesigned the approach around three components: a shared design document defining what is being built, why, and in what order; implementation-ready tasks that could be picked up and executed without additional clarification; and a parallelization layer to identify which workstreams could run independently.
The flow became
approach ->units of work ->ordering
One distinction that came through clearly was between artifacts that are disposable and knowledge that is durable. The design documents, the task breakdowns, the Jira epics, these are temporary. The rules, patterns, and planning skills refined through each iteration compound over time. The system gets better not by accumulating more content, but by improving its structure.
And in the Q&A, that million dollar question came up:
Will this level of automation make us jobless?
The answer was honest: nobody is drawing final conclusions yet. It is still a period of active experimentation, and as Erisa put it, a genuinely fun ride for the teams deep in it.
Where both talks landed
The surface topics were different, but the underlying tension was the same. More capability does not automatically produce better outcomes. What matters is how that capability is structured, constrained, and evaluated over time.
Both talks circled around the same set of questions:
What boundaries does the system operate within?
What does it hand back to a human?
How do you know when it is actually working?
These are not questions teams are preparing for. They are questions teams are already navigating.
And that is a wrap of a night spent asking hard questions, in a room full of people who clearly still have plenty to say.
Thank you to everyone who joined us at Proxify, brought their questions, and made the evening what it was.
See you at the next one!
Data & AI Stockholm 💙
Stay Connected
If you’d like to stay connected with the community and hear about future events, the best places to find us are:






