Blog

Before and After. Never During.

Introducing Inference Assurance — the missing layer in the AI stack.

The AI stack has two mature disciplines for reliability, and a hole between them.

Before deployment: evals. You benchmark the model, red-team it, score it, and decide whether to ship. After deployment: observability. You log the prompt, the response, the latency, the cost.

Neither one watches the inference that actually happened.

The eval tested questions that were not this one. The log recorded an answer already sent. Between them sits the event that carries all of the liability — this model, this input, this generation — and the industry’s answer for that moment is a second model’s opinion of the finished text, or a human reading it before it goes out.

Neither one looks inside the inference itself before the result is acted on.

Who Pays for the Gap Today

Regulated industries pay for it, and they pay in stalled adoption.

Where a wrong answer creates legal exposure — credit, claims, care, advice — supervisors expect the institution to account for how a decision was reached. When the AI leaves no evidence about a specific decision for the risk function to assess before the result is acted on, or to re-examine afterward, institutions compensate with a person. Human-in-the-loop is often not a design choice; it is the control institutions fall back on when they lack a measured basis for letting the AI act on its own. Regulated institutions allow traditional models to make consequential decisions because compliance can test, explain, and re-score every decision. LLMs remain outside that decision loop because compliance cannot yet do the same reliably.

That human-in-the-loop person is the compensating assurance layer. The layer is slow and expensive, and it means AI in regulated institutions often advises rather than acts. The ROI case improves when institutions can reserve human review for the decisions that need it instead of applying it by default. What has been missing is a measured basis for making that distinction, and a record that can be re-examined afterward.

Where firms have moved ahead anyway — pushing autonomous AI into production without sufficient assurance — failures have led to regulatory action and litigation. Untrustworthy or simply wrong outputs reached real people and carried real consequences for them and for the institution.

Look closely at those cases. The problem wasn’t simply that the model got something wrong—something a newer or more accurate model might have prevented. The deeper problem was that, when it did, the operator couldn’t explain why or provide evidence for the decision. The real damage was both the error and the inability to account for it.

That is an assurance failure, not just an accuracy failure. The industry keeps trying to solve it with accuracy.

What Fills the Gap Today

Four categories sit next to this problem. None of them occupies it.

Evals measure a model, not an inference. They tell you what a system does across a distribution. They tell you little about the specific generation that is about to be acted on.

Guardrails and LLM-as-judge read the output. LLM-as-judge is a second model’s opinion of the first model’s text; other guardrails use rules, classifiers, or validators. These can be useful checks, but they operate outside the model that made the decision. They do not see the internal computation that produced the result, and in advanced opaque systems that is the missing evidence.

Observability records what was asked and what was returned. It is indispensable for forensics and it is mainly after the fact. It does not assess what happened inside the model before the result is acted on.

Post-hoc explainability approximates why a model did something, once it has already done it. U.S. adverse-action law requires creditors subject to Regulation B to provide specific reasons for an adverse action, and the regulation’s official commentary says those reasons must relate to and accurately describe the factors actually considered or scored. A post-hoc explanatory approximation therefore cannot be assumed to satisfy the requirement merely because it produces a plausible rationale. That matters because research on model-explanation techniques has identified sensitivity to methodology, reference data, model specification, and other choices that can cause materially different explanations of the same or similar decisions.

Regulatory model-risk guidance has not yet resolved this issue for newer AI architectures. On April 17, 2026, the OCC, Federal Reserve Board, and FDIC issued revised interagency model-risk guidance that expressly states that generative AI and agentic AI models are “novel and rapidly evolving” and therefore outside the scope of that guidance. The agencies said they intend to consider AI, including generative and agentic AI, in subsequent work. That is the gap.

Introducing Inference Assurance

Inference Assurance is the layer that produces evidence from inside the model about a specific inference, at the moment it happens, in a form a risk function can act on and an auditor can re-examine.

Three criteria. A system is doing Inference Assurance only if it meets all three.

Three criteria for Inference Assurance: reads the model's internal computation, renders a verdict before the output is acted on, and retains replayable evidence independently of re-running the model.

Most tools that sound like they belong in this category fail at least one of the three. That is what makes the list useful: it is a test you can run. When a vendor claims to catch unreliable AI output, ask which of the three criteria they meet.

How φ-Lattice Delivers It

Model inference flows through Phi-Lattice measurement and assessment to evidence consumed by model risk, audit, controls and compliance reporting. Verdicts range from No Issues through Noisy, Caution and Suspicious to Unreliable. Measurement and assessment are working today.

φ-Lattice is the first system built to all three criteria.

It sits beside the model at inference time and reads the internal signals the model produces while it generates — not the finished text, and not a log of it. Those signals become a calibrated verdict — from No Issues to Unreliable — together with the risk factors that drove it. The verdict and its evidence persist independently of the model, so they can be re-examined later without re-running the inference that produced them.

That output is built to be consumed, not admired. φ-Lattice does not replace your model risk function, your audit trail, or your compliance reporting. It supplies the input those systems have never had: per-inference evidence in a form they can already ingest. A model risk team gets a scored population to validate against. An auditor gets a specific decision they can re-open. A control owner gets a signal they can write a threshold against.

Measurement and assessment are working today. Mapping the verdict directly into a workflow action — hold, escalate, route to a human — is being built now with our first design partners.

We are naming the category because we believe the gap is structural rather than an artifact of immature models. Better models will not close it. A more accurate system that still cannot account for a specific decision is a more useful system with an unchanged liability profile.

Regulated industries do not need faster AI. They need AI they can stand behind when someone asks them to explain one decision. That is what Inference Assurance is for, and it is what we are building.

→ Interested in becoming a design partner? Get in touch.


Research and sources

Underlying evidence for the claims above.