Skip to content

AI engineering

How to ship AI features that survive real users

Hoang Minh Nguyen

An AI demo takes an afternoon. A prompt, a model call, a text box, and suddenly the product can summarise, draft or answer. Then real users arrive, with their typos, their edge cases and their habit of asking for things nobody planned for, and the feature that impressed the room starts producing answers nobody can stand behind.

The gap between the demo and the dependable feature is not a modelling problem. It is an engineering one.

Start with the failure, not the prompt

Before writing a prompt, write down what a bad answer costs. A wrong product recommendation is an annoyance. A wrong refund amount is a support ticket. A wrong dosage is a lawsuit. The acceptable error rate decides the architecture: whether a human reviews the output, whether the model may act or only suggest, and how much you are willing to spend per request to be right.

Build the evaluation set first

You cannot improve what you cannot measure, and "it feels better" does not survive a model upgrade. Collect a few hundred real inputs with known good outputs and run every change against them: prompts, models, retrieval settings, all of it.

A useful evaluation set includes:

  • The happy path, so regressions are caught immediately.
  • Known failures, added every time a user reports a bad answer.
  • Adversarial inputs, including prompt injection and requests outside the feature's scope.

Treat prompts like code. They need version control, review and tests, because they break in exactly the same ways.

Ground the model in your own data

Most disappointing AI features are answering from the model's general knowledge when they should be answering from yours. Retrieval gives the model the facts; the model supplies the language. Invest in the unglamorous parts: clean chunking, good metadata, and permission checks so the model never retrieves a document the user could not open themselves.

Design the fallback

Every model call can be slow, wrong or unavailable. Decide in advance what the product does in each case:

  1. Timeouts. Stream partial results or degrade to a simpler path rather than spinning forever.
  2. Low confidence. Say so, and offer the manual route.
  3. Outages. Keep the core workflow usable without the AI layer.

A feature that fails gracefully earns more trust than one that is occasionally brilliant.

Watch it in production

Log inputs, outputs, latency and cost per request, with user feedback attached. Review a sample every week. The questions people actually ask are always different from the ones you designed for, and that difference is where the next version comes from.

The teams that ship AI well are not the ones with the cleverest prompts. They are the ones who treat a model like any other unreliable dependency: measured, constrained and wrapped in software that knows what to do when it is wrong.

More articles
All articles