Architecture

Enterprise Agent Audit Trails: What to Record Beyond the Prompt

Orisdale Editorial Team · Published · 6 min read

A reviewer needs more than the prompt and the answer. Record the request, the allowed tool, the calculation version, the result reference, the narrative claims, and the human decision.

A reviewer should be able to reconstruct an agent-assisted answer from the record, not from memory of the chat. The prompt and the final paragraph are not enough. The useful record names the request, the allowed tool, the calculation or query version, the result that was cited, the claims in the narrative, and the human decision that followed.

This article describes that evidence record. The trust-layer article explains why controls sit outside the model. The business-logic contract explains how a calculation is specified. The approved-tool pattern on the architecture page is the sequence this record follows. None of those pieces is a substitute for the record itself.

Record attached to each stage
Question, tool, result, narrative, and review with a parallel recordQuestionRecordscopeToolRecordschema versionResultRecorddefinitionNarrativeRecordclaim linksReviewRecorddecision

Illustrative record. Event correlation is not cryptographic proof, and not every platform exposes a table commit id.

Why the prompt is a weak record

A prompt is an instruction. It is not the identity of the caller, the scope that was accepted, the tool schema that ran, or the number that came back. Two people can send similar wording and receive different results because the period, the entity, or the metric definition differed. If the stored artifact is only the message thread, a later reviewer cannot tell which of those differences mattered.

The failure mode to avoid is a story that treats a fluent answer as if it were an audit. A narrative can be clear and still cite the wrong result, omit a missing field, or imply a cause the calculation never tested. The record should make those gaps visible. It should not promise that the organization will pass a control framework, a certification exam, or a regulatory review.

What to keep with the request

Start with the question the system actually accepted, not only the text the person typed. Unclear parameters should remain unresolved until someone supplies them. A record that silently fills “last quarter” or “the main customer” has already lost the review.

Useful request fields, when they exist in the system you are designing, include:

  • Who asked, in the identity system the platform already uses.
  • The time the request was accepted.
  • The business question, including any clarification that changed the scope.
  • The period, entity, and metric that were validated.
  • The parameters that were rejected, and why.

Identity enforcement is not the same control as parameterization. A correct user can still send an invalid period. An invalid parameter should stop the path before a calculation runs. That stop is itself an event worth keeping.

Tool, contract, and result

The next layer is the allowed tool. Record the tool name and the schema version that accepted the parameters. If the tool calls a governed calculation, record the contract or definition version and the result reference the narrative will cite. The Finance proof is a small published example of this idea: a contract version, a result file, and claim checks that stay separate from the sentences.

Do not assume every warehouse exposes a table commit identifier. Some platforms can correlate a query with a result set. Some can hash an input file. A hash shows that a particular artifact was the one referenced. It does not show that the formula was the right formula, or that a cited driver caused the variance. Event correlation and cryptographic proof are different claims. Publish only the correlation you can actually produce.

A compact synthetic envelope, with no secrets and no realistic tokens, looks like this:

{
  "requestId": "synthetic-req-1001",
  "scope": { "period": "2026-09", "department": "EX-FIELD-OPS", "currency": "USD" },
  "tool": { "name": "variance.lookup", "schemaVersion": "1" },
  "contractVersion": "1.0.0",
  "resultRef": "results.json#EX-FIELD-OPS",
  "narrativeClaims": ["C3"],
  "humanDecision": "not recorded"
}

The department identifier is the same synthetic key used in the published Finance example. It is not a customer. humanDecision stays “not recorded” until a person records a decision. The envelope does not invent a confirmation.

The answer text should be stored with the claims it makes, not only as a blob. A numeric claim can be compared with the result fields. A causal claim cannot be confirmed by arithmetic. If the narrative says a variance was caused by a particular driver, the record should point at the evidence item and leave human confirmation unset until a reviewer assesses that item.

Explicit linkage means the reviewer still has work to do. A matching department and period does not prove the sentence. Unsupported claims stay unsupported. Missing evidence stays missing. The record should not upgrade either state because the prose was confident.

Human decision, access, and failure

Consequential findings stay with a person. The record should be able to hold an approve, reject, or request-evidence outcome, plus who recorded it and when. Until that happens, the status is not a default approval.

Access, redaction, and retention are design questions for the organization, not features this website operates. Decide who can read the prompt, who can read the result, which fields are redacted in a shared review, and how long the record is kept. A failure to call the tool, a denied policy check, and a timeout are different events. Each should be stored as that event, not rewritten into an answer.

A short review checklist

Use this as a design checklist, not as evidence that a control has been tested:

  1. Can a reviewer see the accepted scope without opening a chat transcript?
  2. Is the tool name and schema version present?
  3. Is the calculation or definition version present when a number is cited?
  4. Does each narrative claim point at a result or remain marked unsupported?
  5. Is human confirmation absent until someone records it?
  6. Are denials and missing evidence stored as their own states?

How a reviewer walks the synthetic record

Take the envelope above as a desk exercise, not as a system log. The request id is invented. The scope matches one row in the published Finance example: September 2026, department EX-FIELD-OPS, US dollars. The tool name is illustrative. The contract version matches the published calculation contract, so a reviewer can open that file and see the thresholds. The result reference points at the same department in results.json.

The narrative claim id is a pointer. It does not repeat the sentence, because the sentence belongs with the claim assessment, where its status and human-confirmation field already live. If the assessment says evidence is linked and confirmation is not recorded, the walkthrough stops there. The reviewer reads the evidence description and decides. The record does not finish that decision.

If the same envelope were missing resultRef, the numeric sentence would have nothing to cite. The correct state is an unsupported claim, not a smoother paragraph. If the tool schema version were absent, the reviewer could not tell whether a later schema would have rejected the parameters. That is why the version sits next to the tool name even when the call succeeds.

Limits

This pattern does not make an agent audit-proof. It does not certify a control environment. It does not require every query to be precompiled, and it does not ban every form of generated SQL. Bounded, validated requests can still produce SQL that an approved engine runs, as Snowflake’s own Cortex Analyst documentation describes for its semantic views. The requirement is that the record shows which path ran and which result was cited.

Orisdale does not publish this envelope as a deployed product. The architecture page shows the reference sequence. Discuss one decision, and the record you would need for it, from the contact page.

Keep the record boring on purpose. A reviewer who must decode a novel schema for every answer will not use it. Prefer a stable set of fields and an explicit empty value over a new column for every demo. Add a field when a review actually failed without it, and write down that reason. That discipline is what makes the trail inspectable six months later, when the prompt is no longer interesting and the result still is.

Discuss how this applies to your environment