Plenty of organisations now have an impressive generative AI demo. Far fewer have one in production. The distance between the two is not usually a technical problem — the model works, the demo lands, the room is excited. The distance is governance: the evidence, controls and oversight that a regulated business needs before it will let software touch real decisions and real customers.
If your pilots keep stalling at the same point, it's worth naming what's actually blocking them. In our experience it is rarely model quality. It's that nobody can yet answer the questions a risk committee is obliged to ask.
The questions that stop a pilot
- What data went into this, and were we allowed to use it that way?
- When it gets something wrong, who notices, and how quickly?
- Can we show an auditor why it produced a given answer?
- What stops it doing something it shouldn't — and who is accountable when it does?
A demo doesn't need to answer these. Production does. The work of getting past the pilot is largely the work of being able to answer them with evidence rather than assurances.
Build the guardrails, not just the model
Scope it narrowly. The pilots that ship are the ones aimed at a single, well-bounded task with a clear owner — summarising a document type, drafting a first-pass response, retrieving and citing policy. Broad "assistant for everything" pilots are exciting and almost impossible to govern. Narrow is shippable.
Keep a human in the loop where it matters. Not every output needs human sign-off, but the consequential ones do. Decide deliberately which decisions the system can make alone, which it can only recommend, and make that boundary explicit and enforced rather than implied.
Ground answers in your own sources. A model that retrieves from your approved knowledge base and cites what it used is far easier to trust, correct and audit than one answering from memory. Citation isn't a nicety in a regulated setting — it's how a reviewer checks the work.
Log everything. Inputs, outputs, the sources used, the version of the model and prompt, and any human override. When something goes wrong — and occasionally it will — the difference between a contained incident and a crisis is whether you can reconstruct exactly what happened.
Evidence is the deliverable
Treating evaluation as a one-off pre-launch gate is the classic mistake. Models, data and usage all drift, so the controls have to be continuous: a defined set of tests the system runs against regularly, monitoring for the failure modes you care about, and a clear route for users to flag bad outputs that feeds back into improvement.
This is also where the UK's direction of travel matters. The regulatory environment increasingly expects organisations to demonstrate that AI systems are safe, fair and accountable — not merely to assert it. The good news is that the evidence which satisfies a regulator is the same evidence that makes the system genuinely safer to run. Governance and value are not in tension here; the controls are what let you scale with confidence.
A sequence that works
Pick one narrow, valuable use case. Wrap it in retrieval, human oversight and logging from day one rather than bolting them on later. Define what "good" looks like and test against it continuously. Prove it in a contained pilot with real users and real evidence. Then expand — reusing the same governance scaffolding for the next use case, so each one ships faster than the last.
The organisations getting value from generative AI aren't the ones with the best demos. They're the ones who built the boring machinery — evaluation, oversight, auditability — that lets a regulated business say yes.
General guidance only, not legal or compliance advice. The right controls depend on your sector, use case and risk appetite — we're happy to help map them.