Inspect
Open the stored result and recorded context. A result-only record does not necessarily support recomputation.
From a claim to an inspectable result
What ran, what was observed, and what can be inspected or recomputed.
Recorded demonstration · September 13, 2026
A Llama-3.2-3B-Instruct answer to a question about Nixon's “I am not a crook” statement contained inaccurate historical specifics. Reliability Assessment recorded a Suspicious assessment. A later supported scoring calculation used retained inputs without invoking the assessed model again. The original answer was not regenerated or rewritten.
| Observation | Recorded result | Meaning |
|---|---|---|
| Original assessment | Suspicious | Demonstrated today The original output was flagged for further verification. |
| Reassessment under a different calibration profile | Suspicious | Demonstrated today The retained inputs supported a later calibration-profile calculation. |
| Model calls during reassessment | 0 | Demonstrated today No language-model invocation was needed for this operation. |
Why it matters: an investigation can revisit the assessment of the answer that was actually recorded, rather than substitute a new answer from a new model run.
Controlled synthetic demonstration
In a controlled underwriting demonstration, we held the financial inputs constant and changed only the applicant’s race. The deliberately biased model produced different decisions for the two otherwise identical cases, creating a known reference against which the Decision-Pathway Audit could be tested.
| Variant | Financial inputs | Race | Model outcome |
|---|---|---|---|
| Applicant A | Credit 670; DTI 30% | Race A | DENY |
| Applicant B | Credit 670; DTI 30% | Race B | APPROVE |
The financial profile was identical. Only race changed. The model’s decision changed with it.
The Audit adds supported margin-oriented readouts and representation comparisons to investigate where the model's treatment diverged internally, not just that the final outcomes were different. Its decision-rule mapping demonstration also recovered the planted thresholds within a predefined rule family.
Why it matters: the controlled example makes a model-response difference visible and connects it to an explicit reference. A customer evaluation then asks whether those measurements improve a real investigation.
For regulated lenders, this is the type of differential treatment that fair-lending, model-risk, and compliance teams may need to investigate and document.
This controlled case demonstrates differential model response against a known reference. It does not establish real-world discrimination, causal layer attribution or a legal fairness conclusion. The decision-rule mapping is limited to the predefined tested family.
Controlled synthetic DPA demonstration with deliberately planted rules. Recorded example; not a customer audit.
Replay recomputes a named, supported operation from retained inputs. Opening a result, checking its integrity and recomputing an assessment are different activities.
Open the stored result and recorded context. A result-only record does not necessarily support recomputation.
Check available integrity information. A matching digest does not establish factual correctness or independent origin.
Use the required retained inputs and compatible computation to produce a linked derived assessment or measurement result.
| Operation | Required evidence | Scope |
|---|---|---|
| Reliability Assessment score / reassessment under a different calibration profile | Required features and profile inputs | Demonstrated today Supported scoring and threshold comparisons, not new generation. |
| Geometry recomputation | Required internal signals and execution context | Demonstrated today A qualified measurement stage, not automatic reproduction of every original verdict. |
| DPA contrast replay | Recorded baseline and variant margin trajectories | Demonstrated today Stored-case contrasts, not a new applicant or a new decision-rule sweep. |
Reassessment preserves the original result and links the new result to it. Evidence still has a retention and access lifecycle. Zero model calls does not mean zero compute or storage cost. Missing inputs cannot be created later by changing retention settings.
One real workflow. One conversation to start.
Apply for Early Access