STACKOPTIMA Editorial · Published · Reviewed

Define success at the product boundary

An AI service can return an HTTP success code and still fail the user. It may omit required evidence, produce invalid structured data, exceed the deadline, or answer from an outdated record. Production reliability therefore needs a definition of successful work that extends beyond server availability.

Choose a small set of outcomes that matter. For a document extractor, success might require valid fields and delivery before a queue deadline. For an assistant, it might require a supported answer or an appropriate escalation within the conversation's time limit. These are proposed examples, not universal service-level targets.

Google's SRE guidance uses service-level indicators and objectives to connect reliability measurements with engineering priorities. Its error-budget approach makes the permitted shortfall explicit. The practical question for an AI product is which user events count as good, which count as failures, and over what observation window.

Measure the slow and unsuccessful work

Track complete task attempts, including retrieval, tools, model calls, and validation. Measure initial response time separately from completion time. Add outcome categories for timeouts, provider errors, invalid output, missing evidence, and cancellations. Without this separation, teams may spend time tuning generation speed while the real delay comes from retrieval.

Report sample sizes and the monitoring window. A calm hour does not establish a monthly reliability rate. A provider's status page also does not measure your application: your region, access configuration, quotas, database, and network path create additional dependencies. Treat external status information as context for investigation.

For an original hypothetical target of 99% successful tasks over 10,000 eligible tasks, the allowed shortfall is 100 tasks. Decide in advance what action a rapid rise in failures triggers. An objective without an operating response is a dashboard decoration rather than a management tool.

Give retries a budget

Retrying can recover a transient failure, but repeated attempts consume time, money, and capacity. Set a maximum attempt count and a total task deadline. Distinguish failures that may recover from failures that require a different action, such as invalid credentials or an unsupported request.

Space retry attempts instead of sending synchronized bursts. Use a delay policy appropriate to the provider's guidance and honor any explicit retry timing it returns. Ensure that retrying a tool action cannot duplicate a purchase, message, or other side effect. A failed network response does not necessarily mean the underlying action failed.

Calculate the financial consequence. In a deliberately simplified example, 50,000 tasks with one extra billed attempt on 8% of tasks create 4,000 additional attempts. Whether this is acceptable depends on the price, the recovered success rate, and the user deadline. Log the reason for each retry so the system can improve rather than repeatedly pay for the same mistake.

Design fallback as a product behavior

A backup model is useful only if it can satisfy the same essential constraints. Check supported inputs, output schema, evidence handling, privacy terms, and tool behavior. A technically available substitute can still create a worse failure if it silently drops capabilities or changes the meaning of the result.

Sometimes the appropriate fallback is a saved result with a visible timestamp, a reduced feature set, or a clear escalation. Decide what users can safely do with degraded output. Do not present stale information as current merely to keep the interface looking healthy.

Make recovery testable

Create a short runbook for common failures: source unavailable, quota exhausted, malformed output, delayed refresh, and database restoration. Assign an owner and describe how to roll back a change. Rehearse selected failures in a controlled environment without affecting real users.

Reliability improves when incidents produce a specific change: a better timeout, a corrected alert, a narrower permission, or a regression case. Keep the connection between observed failure and corrective work visible. The goal is dependable completion under the conditions your users actually experience, with an honest response when those conditions cannot be met.

Sources and further reading

Apply this to your workload →

All insights · Editorial policy and corrections