A convincing AI demo takes days. A system that a finance team, a clinic or an operations floor will rely on every morning takes engineering. The difference is rarely the model itself — it is everything around it: the data it sees, the decisions it informs, the way it fails, and how anyone would know.
Across the AI features we have built into production software, the projects that make it share a handful of habits. None of them are exotic. All of them are deliberate.
Start with the decision, not the model
The most useful question at the start of an AI project is not “which model?” but “which decision gets better, and for whom?” A support team triaging tickets, an accountant reconciling statements, a coordinator scheduling appointments — each makes a decision with a cost when it goes wrong and a benefit when it goes right.
Framing the work around that decision gives you a scope, a success measure and an honest fallback: what does the person do today, and what should happen when the AI isn't confident? If you can't answer that yet, you aren't ready to choose a model.
Treat data like an interface
Models are only as dependable as the data flowing into them, and in production that data changes. A new form field appears, a vendor alters an export format, a department starts using a free-text column for something it was never meant for.
- Define explicit contracts for every source the model reads — schema, freshness and ownership.
- Validate inputs before they reach the model, and log what was rejected and why.
- Version the retrieval corpus the way you version code, so every answer can be traced back to the documents that produced it.
The model is the smallest part of a production AI system. The data contracts, guardrails and feedback loops around it are where the engineering happens.
Design the human in the loop on purpose
“Human in the loop” is often added as a disclaimer. It works far better as a product feature. Show the evidence behind a suggestion, make accepting or correcting it a single action, and route low-confidence cases to a review queue instead of guessing.
Every correction is also a labelled example. Capturing it cleanly turns everyday use into a steady stream of evaluation data — which is exactly what you need next.
Evaluate continuously, not once
A one-off accuracy test before launch says little about week twelve. Production AI needs the equivalent of a regression suite: a curated set of real cases with expected outcomes, run automatically whenever a prompt, model version or retrieval index changes.
- Keep a golden dataset drawn from real, anonymised usage.
- Track quality, latency and cost per request side by side.
- Alert on drift — in the inputs as well as the outputs.
Build the boring parts first
Authentication, audit logging, rate limits, retries, cost controls and a kill switch are not glamorous. They are also the first things a security review, a compliance officer or an incident will ask about. Building them before the clever parts is what allows the clever parts to ship.
None of this slows a team down. It is what lets them move quickly without betting the business on a demo.
