Logo Marius Högger

Bootcamp AI-Agents - Transforming Organizations with AI Agents

AI Agents in Production

26.7.2026ca. 30 participants

On 26 July 2026 I was once again a guest at the FHNW AI Agents Bootcamp to deliver my talk "AI Agents in Production: What the Platform Must Deliver". I had given the same talk back in March at the first run of the bootcamp. The core message remained the same: a single agent is quick to build, but from the third use case onward a central platform is needed to solve cross-cutting concerns such as knowledge base, governance, transparency, evaluation, LLM gateway and PII protection once for everyone.

A Different Group, a Different Dynamic

What made this run special was the discussion. Participants came with concrete plans and wanted to know one thing above all: how much effort is really behind an AI use case? How long does it take until a RAG system is production-ready? How large does a team need to be to operate an AI platform?

I tried to give honest assessments from my experience across various projects. The truth, however, is that almost every answer has to start with "it depends": on data quality, on the existing infrastructure, on the requirements for accuracy and compliance. I probably said "it depends" a few too many times that morning, but it does reflect reality: blanket effort estimates for AI projects are almost always misleading.

Practical Experience as the Common Thread

My experiences from concrete use cases were particularly in demand. Participants wanted to know which projects had actually delivered the expected value and where the typical stumbling blocks were. Topics such as knowledge base quality, dealing with hallucinations and the question of when a PoC makes the leap to production dominated the conversations. A pattern emerged: the biggest challenges rarely lie in the technology itself, but in data preparation, organisational anchoring and realistic expectations.

Evaluation as an Eye-Opener

As at the first run, the evaluation module attracted considerable interest. Many participants had barely thought about how to measure the quality of an AI system systematically. The idea that AI quality does not have to be left to chance, but can be checked automatically with reference datasets and LLM-as-judge, was a genuine eye-opener for some. The discussion quickly turned to the question of who in the organisation maintains these evaluation datasets and how to get subject-matter experts on board.

Summary

The second run of this talk reminded me again that the AI platform topic strikes a nerve with organisations. The questions are becoming more concrete, the plans more ambitious, and the appetite for honest, first-hand accounts from practice is greater than ever. Even if I said "it depends" more often than I would have liked: that kind of nuance is exactly what prevents AI projects from failing on unrealistic expectations.