Speed isn't quality

Opinion5 min

AI makes it easy to mistake speed for quality. A model produces an analysis in two minutes, a prototype stands up in an afternoon, and in the demo an agent does exactly what it should. What you see on screen then shapes how you judge the work.

Vibe coding amplifies that. You describe what you want, the model writes the code, and after a few hours you have something that works as long as you are the only user.

What is the difference between a vibe-coded prototype and a production system?

A prototype shows that something can be done; a production system keeps working when many people use it every day and the builder is not around. A prototype does not need to run under five hundred real users or handle exceptions. It can do without logging, a permission structure, a fallback, a handover, and an owner. A production system cannot.

When I review the code of a prototype, the foundation is often better than expected. What is missing is almost always the production layer: tests, a place where someone other than the builder can manage the system, and a design that holds beyond the first users. That is no criticism of whoever built the prototype. The prototype had a different job.

With my own product Signal Match I saw the same thing from the other side. In July most of the work went into tests, billing, and automatically recovering stalled processing. New features came second that month. Users never see that work, but they notice straight away when it is missing.

Rebuild or harden?

Harden what is there if the foundation holds, and rebuild if it does not. If security and structure are sound, the work goes into the production layer and the rest stays. If security is missing or the code cannot be followed, starting over is usually cheaper than repairing.

Timing matters too. As long as there are only test accounts, you can rework the data model without bothering anyone. With real users and real data, every change becomes a migration. So take the step to production before the first customer logs in.

What does the production gap cost?

I call the distance between demo and daily use the production gap. It holds the work fast projects tend to push forward: compliance, maintenance, and making sure people actually use the system.

That bill arrives later. At launch, when compliance blocks: since 2 August 2026, Article 50 of the EU AI Act requires a chatbot to tell you that you are talking to AI, unless that is already obvious. A prototype rarely accounts for that. What else applies is on my EU AI Act page. After three months, the bill arrives when nobody owns the system. Or on a Tuesday morning, when the only builder is gone and the system stops.

Gartner expects more than 40 percent of agentic AI projects to be cancelled by the end of 2027, due to rising costs, unclear business value, or inadequate risk controls. Those are exactly the items a demo does not show.

How do you take an AI prototype to production?

Start by writing down what “done” means, then test on the client’s data. This is the order I keep to:

  1. A definition of done in week one. Which questions does the system answer, at what error rate, in which system, and who judges that?
  2. A baseline measurement. How does the process run now, in minutes per task or errors per batch? Without that measurement every result is an opinion.
  3. Testing on real data. An application can work well on the builder’s test set and behave differently on data from daily work. That is why the evaluation set runs on the client’s machine.
  4. The production layer. Logging, permissions, a fallback, and points where a person can check and correct before the system moves on.
  5. An owner and a runbook. One name in the organisation, and one page on what the system does, what it must not do, and where to look when it stops.

Each step comes from a lesson in my memo on 35+ implementations, with the examples.

When does building fast help?

When you use the speed to learn sooner. This is not a case for months of discovery: a small scope in four weeks is often better than a hundred-page roadmap.

Testing a hypothesis quickly is good, and so is building fast with the future owner next to you. It goes wrong when a demo is sold as a production system, or when the build comes without a handover. Then the client stays dependent on the builder.

An early decision point is cheaper than a surprise in week three. In my projects it falls on day five. If the case is simple, we keep building in weeks three and four. If it is average, the client chooses between a simpler automation now or a larger project. If it is complex, the short format is dropped and a longer project starts straight away.

So ask of every fast project what the speed left out.

Why depth is often faster

A system with clear scope, one owner, documentation, and measurement points moves through an organisation faster than a polished demo nobody makes a decision about. You lose less time on repair work, on winning trust back, and on explaining why something that “already worked” has to be rebuilt.

Trust is the expensive item here. AI agents now run without errors 95 to 99 percent of the time, but people remember the one mistake. If nobody explains it, they start working around the tool.

That is why I would rather build one workflow properly than five prototypes halfway. One system people use says more than five ideas that get applause on Slack.

Want to read on?

The memo collects the lessons from 35+ AI projects: where pilots got stuck and what I do differently because of it. A 14-minute read.

30 minutes · free, no obligation