Back

What 35+ AI implementations taught me.

Memo · 14 min read
Thijs BongertmanHonest AI

Since February 2025, in my previous role, I ran more than 35 AI implementations, and not every one went well. This memo holds the fourteen lessons I took from them, in three phases: before you sign, while building and after go-live. The examples stay general so the organisations involved cannot be identified. At the bottom are ten questions to ask an agency before you sign.

Part 1. Before you sign.

The most expensive mistakes happen before a single line of code exists: in the intake, the scope and the platform choice, in what nobody says out loud about what "done" means, and in the measurement nobody took.

Lesson 1The problem is rarely where the client says it is.

First conversations start with a solution. "We want a chatbot." Or: "We want to make our documents searchable with AI." Rarely with the problem itself.

My first step is to take that solution back to the question. Which problem does this solve, for whom, how much time does it save, and is this the biggest problem there is?

In one project, a few more questions showed that a different bottleneck deserved more attention than the original build request. We changed the order of the work accordingly.

In another engagement, the problem was how people used existing software, and building more would not fix that. First we had to understand why the features already available went unused.

First ask whether this is the problem you need to solve. Whether you can build it comes after.

Lesson 2Price on assumptions, pay in hours.

I have seen proposals assume less material and fewer exceptions than actually existed. Once price and scope are fixed, that forces difficult choices. Work that seemed out of scope came back during delivery anyway.

So I audit the material first and name a price after. How many documents are there really, how many exceptions sit in the process, and who works with it? I write down the exclusions, and we discuss changes before anyone builds them. One day of looking prevents a quarter of overrun.

Lesson 3The platform sets the ceiling. The client's own platform too.

During one implementation, the chosen platform turned out not to support everything the design needed. A better prompt cannot fix that. The limitations should have been on the table up front; we adjusted the design and the expectations.

The limits also sit outside the AI platform. Existing systems, licences and internal decision-making determine just as much what is possible. Since that project, I run a platform assessment with the client before kickoff, covering those four: the AI platform, their own systems, the licences and who has to decide. Sometimes the outcome is a different platform, sometimes a different home for this use case. You only see the ceiling if you ask about it.

Lesson 4Expectations are asymmetric. Write down what "done" is, and how you measure it.

Clients hear "PoC" and think "working product". They hear "four weeks" and think "everything done".

That is not bad faith. Marketing creates expectations a first version cannot meet yet, and I have sometimes discussed those expectations too little myself. Technically working and useful in daily work are two different judgements.

So in week one I write a definition of done. "A working chatbot" is too vague. The definition states which questions it answers, at what error rate, in which system, and who judges that. The document is a fixed part of project start, with a formal acceptance moment.

Agree on the measurement method too. I have seen a builder and a client judge the same application differently. Besides the target, agree on the data, exceptions and calculation rules the measurement uses.

That conversation is uncomfortable, and it is the most useful one you can have early in a project.

Lesson 5Without a baseline, every result is an opinion.

Most AI projects end with a feeling. "It really saves time." "People are enthusiastic." Those are moods. A result is a delta: this is how it was, this is how it is now, this is the difference.

So the baseline is a fixed step in the plan. In week two I measure how the process runs today: minutes per task, errors per batch, the number of people who actually use the existing tooling. In week three I measure the delta. You need it when someone asks in month four why this cost money.

A baseline can change the scope. If existing software is barely used, guidance sometimes achieves more than a new application.

Lesson 6Billing by the hour rewards overrun.

Unclear phases and missing evaluation points make overruns hard to contain. I have watched extra work pile up when nobody had agreed in advance when to decide again.

That is why I work on a fixed price and a fixed scope. It comes down to the incentive. Whoever bills by the hour earns from scope creep. Whoever agrees a fixed price earns from tight scoping, and that is what you want as a client.

The second reason is new. An analysis that used to take two days takes two hours with AI. Code that took three days is done in one. Bill by the hour and the client gets the benefit of your speed while you gain nothing, so an agency stops getting faster. A fixed price on a fixed result rewards the agency that gets faster, and the client knows up front what they get.

Do not start delivering before the agreement is in place. A good relationship does not replace a clear scope and a signed engagement confirmation.


Part 2. While building.

The next four lessons cover what goes wrong between kickoff and delivery, and how to spot it early.

Lesson 7Small scope, big trust. And a checkpoint on day five.

The projects that turned out best were the most tightly bounded: one problem, one team, four weeks, then measure and decide whether to continue.

When someone asks for a complete platform, I prefer to start with one bounded application. You see what works and what is missing before you add more components.

That lesson applies to my own offer too. The four-week kickstart from my previous role was, for a long time, too ambitious for the price. Training, analysis and an automation live in four weeks only works for simple cases. On more complex projects it only became clear in week two or three that things were harder, and that is where scope creep starts. So the decision point now sits on day five, with three outcomes. Simple: go ahead and build in weeks three and four. Medium: the client chooses between a simpler automation now or a larger engagement. Complex: several things have to happen at once, so the kickstart format is dropped and a longer engagement starts straight away. One early decision point with the client is cheaper than a surprise in week three.

Lesson 8Internal green is not proof. Client data is the truth, and the test should belong to the client.

An application can work well on the builder’s test set and behave differently on everyday data. I only found those differences by testing with the client.

Watch the input too. An unexpected value or exception can make processing fail; include those cases in the assessment.

Build against your own test set and judge the result on real client data. Report deviations and limitations, so everyone has the same picture of what works.

Then go one step further and give the client the test. I publish the evaluation set I build with, so it runs on the client's machine and the client judges the result. A claim you cannot recompute is marketing. That goes for "95 percent accurate" as much as for "35+ implementations".

Lesson 9The security blocker is often a phantom blocker.

When an integration was blocked, another route to the required information turned out to exist. So establish what information you need first, and choose the technical route after.

The reverse happens too: legal and IT take months over an AI policy while employees use personal ChatGPT for work documents. Without guardrails, your most careful people are the ones who drop out, because they wait for permission.

At every blocker, first ask what data you minimally need, and whether it already sits somewhere you are allowed to go. And put guardrails in place before the tool arrives. A one-page policy on day one beats a legally watertight document in month six.

Lesson 10The champion is the key. And the bus factor.

Every successful project had a champion: an employee, not the IT manager or the director, who saw what AI meant for their own work and brought colleagues along.

Projects without a clear champion stall, even when the tool is good. The tool goes unused, the feedback dries up, and after three months nobody remembers what was built.

One champion is also a risk. If only one person understands the application, further development and maintenance become fragile. A team becomes independent when several people understand the application and can guide others, and that is part of the handover.

Find the champion in week one and give them access, time and a say. By week four, make sure there are two.


Part 3. After go-live.

These four lessons cover what decides, after go-live, whether a tool gets used or sits idle by month three.

Lesson 11Adoption doesn't start after launch.

The most common mistake is planning adoption as a separate step after implementation. "We build the tool, then we run a training, then people will use it."

That doesn't work. Adoption starts at problem analysis. Involve the people who will use the tool in defining the problem: their input makes the tool better, and people who helped frame the problem use the solution sooner.

Support from a manager shows up as time to practise, ask questions and help decide. That time has to exist during the project already.

Lesson 12The silent majority decides adoption, not the believers.

In every organisation you see three groups. The accelerators are experimenting on day one. The holdouts, fifteen to twenty percent, are sceptical or too busy and only move once they see proof. In between sits the middle group: sixty to seventy percent of people, who think it's fine but don't start on their own.

The mistake I saw most often: all the attention goes to the accelerators. They are enthusiastic, give feedback and show up to the sessions, but they don't bring the middle group along by themselves. A tool that twenty enthusiasts use and two hundred others don't is a hobby club.

The middle group needs examples from their own work. Let colleagues guide each other and discuss where the application helps and where it does not.

Lesson 13One mistake weighs more than 99 good answers.

AI agents now run without errors 95 to 99 percent of the time. Still, people remember that one mistake. If it comes without an explanation, distrust grows, and a tool that is distrusted gets bypassed.

Let the application state what it does not know and show its sources. Build in moments where a person can check and correct it before the system continues. Then you can see what you can and cannot trust.

Lesson 14Never leave an agent behind without an owner.

Ownership determines what happens after you leave. Dan Shipper of Every puts it this way: "every agent needs a human". Gartner expects over 40 percent of agent projects to be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls.

That is why, since August 2026, my handover is a fixed list of five. One name that owns it. The evaluation set from lesson 8, running on the client's machine and not on mine. A one-page runbook: what it does, what it must not do, where to look when it stalls. An agreement on who steps in when it errs, with the button from lesson 13. And a path for the next model version: the model that works today can be replaced within six months, and then a retest has to run.

That list is younger than most of the projects in this memo. The first engagement where the client has ticked all five points is still running. That is why it sits here as a lesson and not as proof. Ask every agency what it leaves behind and who owns it afterwards.


The thread.

35+ implementations mostly prove that I have done enough wrong to know what works.

The successful projects followed the same pattern every time: measure first, then build, and show how you measure. The ten questions below test each step.

Most companies skip at least one of those steps, usually because the agency doesn't ask.

Ten questions before you sign.

Ask them of every agency, including me.

  1. Which problem are we solving, for whom, and how much time or money does it save per week? If the answer is "a chatbot", it isn't an answer.
  2. Did the agency see the material before naming a price? How many documents, how many exceptions, how many systems?
  3. What can the chosen platform not do, and what can our own systems and licences not do? Ask for three concrete limits; "none" is a red flag. And who at the works council, security or legal still has to say yes?
  4. What is the definition of done, on paper, with error rate, measurement method and reviewer?
  5. What is the baseline, when is it taken, and who measures the delta in week three?
  6. Is the price fixed or hourly? Who pays for overrun?
  7. When is the first moment we can stop without losing face? If that is only after four weeks, it's too late.
  8. Which data is it tested on: the agency's or ours? And can we run the test ourselves?
  9. Who is the internal owner, and who is the second?
  10. What happens in week one with the people who will use the tool, and what is the plan for the sixty percent who won't start on their own? If the answer is "training after delivery", read lessons 11 and 12 again.

Which one are you skipping?

Want one AI engagement reviewed?

Request a free 30-minute introductory call, with no obligation. We discuss your question and a possible next step. Pick a time in my calendar.

Let’s talk