AI engineering
How to ship AI features that survive real users
Hoang Minh Nguyen

An AI demo takes an afternoon. A prompt, a model call, a text box, and suddenly the product can summarise, draft or answer. Then real users arrive, with their typos, their edge cases and their habit of asking for things nobody planned for, and the feature that impressed the room starts producing answers nobody can stand behind.
The gap between the demo and the dependable feature is not a modelling problem. It is an engineering one.
Start with the failure, not the prompt
Before writing a prompt, write down what a bad answer costs. A wrong product recommendation is an annoyance. A wrong refund amount is a support ticket. A wrong dosage is a lawsuit. The acceptable error rate decides the architecture: whether a human reviews the output, whether the model may act or only suggest, and how much you are willing to spend per request to be right.
Build the evaluation set first
You cannot improve what you cannot measure, and "it feels better" does not survive a model upgrade. Collect a few hundred real inputs with known good outputs and run every change against them: prompts, models, retrieval settings, all of it.
A useful evaluation set includes:
- The happy path, so regressions are caught immediately.
- Known failures, added every time a user reports a bad answer.
- Adversarial inputs, including prompt injection and requests outside the feature's scope.
Treat prompts like code. They need version control, review and tests, because they break in exactly the same ways.
Ground the model in your own data
Most disappointing AI features are answering from the model's general knowledge when they should be answering from yours. Retrieval gives the model the facts; the model supplies the language. Invest in the unglamorous parts: clean chunking, good metadata, and permission checks so the model never retrieves a document the user could not open themselves.
Design the fallback
Every model call can be slow, wrong or unavailable. Decide in advance what the product does in each case:
- Timeouts. Stream partial results or degrade to a simpler path rather than spinning forever.
- Low confidence. Say so, and offer the manual route.
- Outages. Keep the core workflow usable without the AI layer.
A feature that fails gracefully earns more trust than one that is occasionally brilliant.
Watch it in production
Log inputs, outputs, latency and cost per request, with user feedback attached. Review a sample every week. The questions people actually ask are always different from the ones you designed for, and that difference is where the next version comes from.
The teams that ship AI well are not the ones with the cleverest prompts. They are the ones who treat a model like any other unreliable dependency: measured, constrained and wrapped in software that knows what to do when it is wrong.
KeepReading

Designing multi-tenant SaaS backends that scale
The isolation, data and billing decisions that let a platform grow without a rewrite.

Velocity without shortcuts: how we ship weekly
The pipeline, review habits and release rituals that let a small senior team ship every week.