Few organisations today have run no artificial-intelligence experiment at all: a drafting assistant here, a document-search prototype there, a demonstration shown to the executive committee. Yet the pattern we see in the field keeps repeating: pilots accumulate, impress for a few weeks, then quietly fade away.
This is not a technology problem. For the vast majority of enterprise needs, available capabilities are more than sufficient. What blocks the transition lies elsewhere — how the need is framed, whether the data is available, how the output fits into the real process, and whether teams trust the result at all.
Why pilots never become usage
A pilot proves that something is possible. Usage proves that it is useful, repeatable and sustainable. The gap between the two is organisational far more than technical, and the reasons for stalling repeat themselves.
- The project started from an available technology rather than an identified business problem.
- No success criterion was defined before launch, leaving nothing to judge but general impressions.
- The required data exists in theory, but is scattered, incomplete or unusable as it stands.
- Nobody owns the use after the demonstration: no business owner, no maintenance budget.
- The output is integrated nowhere: users switch tools, copy, paste and check — and the gain dissolves into friction.
Industrialising is therefore not about "moving to a better model", but about working through those five points.
Choosing a use case: four questions before you start
The choice of the first use case determines much of the outcome. We put it through four questions.
- How frequent is the task? Something repeated daily by several people delivers a return no rare, spectacular case can match.
- What does an error cost? A clumsy sentence in an internal note is fixed in ten seconds. An error in a regulatory calculation is not.
- Is the data actually available? Accessible, current, structured — and permitted to leave its original perimeter.
- Is there a measurable success criterion? Handling time, rework rate, volume processed at constant headcount: without a measure, the project can never be arbitrated.
A good first use case combines high frequency, low error cost, available data and a clear measure. It is rarely the most impressive one in a demo.
When AI is not the right answer
Let us be clear: a significant share of the "AI projects" we audit are in reality rules problems or data-quality problems. Where a task follows stable, verifiable logic — checking a format, applying a rate table, routing a file — deterministic automation does it better: faster, cheaper, testable and auditable. Introducing a probabilistic model there adds uncertainty where none existed.
Generative AI becomes relevant when the input is unstructured or ambiguous, when the cases are too numerous to express as rules, or when the expected output is language. And no model compensates for an inconsistent reference base: if your data is wrong, AI will produce errors faster.
Humans in the loop: calibrating oversight to the cost of error
The question is not "should there be a human in the loop?" but "where, and how intensively?". We reason in levels, according to the cost of an error.
- Systematic validation: nothing goes out unreviewed. Mandatory for any published content, any contractual commitment, any decision affecting a person.
- Sample-based control: the process runs, a defined percentage is reviewed periodically, and a drift threshold returns it to systematic validation.
- Autonomous execution: reserved for internal, reversible, low-stakes tasks, with logging and the ability to undo.
This only has value if the review is real. Validation "for form's sake", on excessive volumes, creates an illusion of control more dangerous than no control at all.
Evaluation: what does "it works" mean?
This is the step most often skipped, and the one that separates a serious approach from a fashion. It requires three things. A representative test set: real cases, deliberately including edge cases, with the expected answer established by the business. A before/after measure on indicators defined upfront and across the whole process, checks included, not on the assisted step alone. And reproducibility over time: a system's behaviour shifts with its versions and its data, so what was validated once must be revalidated.
A system that is right nine times out of ten, with nobody able to spot the tenth, is not a productivity gain: it is a transfer of risk.
Build in a user-feedback channel too, embedded in the tool, so an inaccurate result can be flagged in one gesture. Without it, defects surface late and trust erodes quietly.
Risks worth facing squarely
A credible approach names its risks. Five deserve particular attention.
- Plausible but wrong answers. A well-phrased output can still be inaccurate. This is the most insidious risk: it disarms the reader's vigilance.
- Data leaking to third-party services. An employee pasting a confidential document into an unapproved service exposes the organisation without any malicious intent.
- Bias. A system learned from past practice reproduces it, imbalances included — a major issue in recruitment or in access to a service.
- Vendor dependency. Proprietary models, formats and connectors: exit costs build up fast. Reversibility is designed in from the start.
- Opaque decisions. If you cannot explain why a case was handled a certain way, you will be able neither to defend it nor to justify it to a regulator.
Governing use: internal rules and the regulatory frame
Governing does not mean forbidding. Without an explicit framework, usage spreads anyway, beyond any control, and the organisation discovers too late what has circulated.
An internal usage policy
Four elements form a minimum base: a data classification stating what may and may not be submitted to an AI service; a list of approved tools, with the route to have a new one cleared; logging of sensitive uses; and a rule requiring human validation of any published content. Decide explicitly, too, what will never be automated: some decisions should stay human by choice, not by technical limitation.
A regulatory framework still taking shape
The emerging European framework on artificial intelligence follows a risk-tiered approach: the more a use can affect people's rights or safety, the higher the requirements on documentation, oversight and data quality, with some uses ruled out entirely. A general transparency expectation sits alongside it, notably where a person interacts with an automated system or reads generated content. This articulates with existing data-protection rules, which continue to apply in full. Without pre-empting the detail of future obligations, document your uses and your controls now.
Moving to controlled use
The real cost, routinely underestimated
Decisions are too often made on the cost of access to the technology alone, which is only a fraction of the total. The decisive cost lies elsewhere: integration with existing tools, data preparation, oversight time, periodic evaluation, version management. A quiet but well-integrated use creates more value than a brilliant demonstration nobody maintains.
Training and bringing teams along
Adoption cannot be decreed. Teams need to understand what the tool can do, what it cannot, and how to spot an error — training in critical judgement far more than in the interface. They also deserve an honest answer about how their role will change; while that question stays implicit, resistance stays silent. Involve future users from the framing stage: they know the edge cases.
The next step
If you take away one thing: choose a use case with a low cost of error, high frequency and measurable value; write its success criterion before you launch; design in user feedback and oversight from the outset; and put on record what will never be automated. One successful, adopted use is worth more than ten abandoned pilots — and makes the next one far easier to fund. Our consultants can help you identify that first case and put the framework around it.



