From guarantees to quality margins: the governance shift AI requires
This article is part of Framna's Trust by Design series, exploring what it takes to build trustworthy AI systems beyond the model itself. If you haven't read the previous articles yet, you can find the first one on grounding here, and the second on evaluation here.
The previous articles in this series explored two important layers for reducing AI risk. Grounding connects a model to an organization's own sources, reducing the likelihood of invented information. Evaluation tests outputs through deterministic checks, model-based evaluation and human review.
Together, these layers make quality more visible. They provide signals about where a system performs well, where it falls short and where stronger controls may be needed. Because even with both layers in place, some uncertainty remains.
Establishing the quality margin
Some outputs will still be incomplete or subtly wrong. Others may appear convincing but are incorrect. Evaluation can identify many of these issues, but it cannot predict every way a system will behave once people start using it in real situations.
The quality margin therefore needs to be established together. Framna can advise on what is technically realistic by evaluating the system against representative use cases, available source material and agreed quality criteria before launch.
The organization brings the business context: how the output will be used, the consequences of different types of errors and the requirements the product needs to meet. Together, these perspectives determine which controls are needed and what level of quality is realistic for the use case.
From financial risk to quality margin
The financial industry offers a useful comparison. Financial institutions know that risk cannot be eliminated entirely, so they use measures such as Value at Risk, or VaR, to understand potential exposure within defined parameters.
AI requires a similar mindset, although the measurement itself is different. The aim is not to promise that a system will never produce an incorrect answer. It is to understand the quality that can realistically be achieved, the types of deviation that could occur and the consequences when they do.
We call this a quality margin: the acceptable range of deviation for a given use case, together with the controls required to stay within it.
Not every mistake has the same impact. A wrong order amount matters more than an awkward sentence in a customer service reply. The quality margin should reflect that.
The margin therefore cannot be expressed as one generic accuracy percentage. It needs to reflect the product, its users, the available data and the consequences of getting something wrong.
Quality is not a fixed point
After launch, the conditions around an AI product rarely stay the same. Users start asking questions that were not part of the test phase, business rules change and source content can become outdated or only partly updated. The system itself may also change, for example through model updates, adjustments to retrieval or new product flows. And as a use case becomes more important or sensitive, the organization may no longer accept the same level of deviation. That is why the quality margin agreed at launch needs to be reviewed over time.
Mistake-free output is rarely a realistic standard. What matters is whether teams understand how the product is performing, can identify meaningful changes early and know when intervention is needed. Keeping AI quality within the agreed margin therefore requires more than monitoring. It requires clear responsibility for what happens after launch.
Effective AI governance depends on a clear division of roles after launch
The technology partner has an important role in the system itself: the architecture, control measures, monitoring and technical steering needed to keep performance within the agreed margin. The organization brings the business context the system depends on: the source content, operational knowledge, the agreed tolerance for the use case and understanding of how the output will be used in practice.
When performance changes, the first step is to understand where that change comes from. Sometimes the system or its controls need adjustment. Sometimes the underlying content, business rules or context have changed. In practice, both sides need to stay involved, because AI quality depends on the relationship between the technical setup and the environment in which the product is used.
This division should be agreed before launch and revisited over time. As the system, usage and business context change, the margin may need to be reviewed as well.
What this means for organizations building on AI
Grounding and evaluation make AI systems easier to inspect, measure and improve. They also show why AI quality cannot be treated as a one-time launch decision.
There will always be a margin that needs to be managed. Leadership needs to create the governance around that reality: what quality means for the use case, which deviations are acceptable, who reviews the signals and when the system, the content or the controls need to change.
Building AI into a product means taking responsibility for a system that will keep moving after launch.
Subscribe
Join our newsletter and stay up-to-date