Noxia

Field note · Advice & compliance

In a regulated firm, the audit trail is the product.

Every AI demo shows you the result. Almost none show you the record. In a regulated firm that is backwards, because the record is the part you will actually be asked for.

3 min read Sources checked 8 September 2026

The short answer

Log six things for every agent action: what triggered it, what inputs it read and from where, what it produced, which human confirmed it, what changed between draft and final, and what it refused or escalated. If a supplier cannot show you that record in a demo, they have built the interesting half and left you the liable half.

Ask a supplier to demo the audit trail rather than the agent. Watch what happens.

Most demos are built around the moment of output — the report appears, the call is booked, the pack is assembled. The record of how that happened is usually a log file somebody will "expose in the API".

For a firm outside regulation, fair enough. For you, that is the deliverable.

What one agent action must leave behindSix records for every agent action: trigger, inputs and their sources, output, the human who confirmed it, the difference between draft and final, and anything refused or escalated.1Triggerwhat started it2Inputsand their sources3Outputwhat it produced4Confirmed bya named human5Diffdraft vs final6Refusedand escalated to whomWHAT ONE AGENT ACTION MUST LEAVE BEHINDSix records. If a supplier can only show you number 3, you are buying a demo.
The six records. Number six is the one nobody builds.Noxia’s specification, not a regulatory standard. Written to satisfy Consumer Duty evidence questions.

Why is "the model decided" not an answer?

Because it is not a decision anyone can be accountable for, and accountability is the unit the regulator works in.

Consumer Duty asks you to evidence good outcomes. SM&CR asks who was responsible. Neither has a category for a process nobody owns. If your file cannot name a person at the point of decision, the answer is not "the AI" — the answer is you.

The purpose of the log is not to prove the agent was right. It is to prove a human was there when it mattered, and to make it cheap to find out when they were not.

What does each record actually need to contain?

RecordMust containThe question it answers later
TriggerEvent, timestamp, which rule firedWhy did this happen at all?
InputsEach document, its origin, its version and dateWhere did that figure come from?
OutputThe artefact, immutable, with a hashIs this the version that was sent?
Confirmed byNamed individual, timestamp, what they sawWho is accountable?
DiffWhat the human changed before it wentDid anyone actually read it?
RefusedWhat it would not do, and who it went toDid the guardrail work, or was it decorative?

Why does the refusal record matter most?

Because a guardrail nobody can see is indistinguishable from no guardrail.

If your agent has never refused anything, one of two things is true: the work is genuinely always in scope, or the limit does not exist. The log is how you tell those apart, and it is the first thing we would look for in someone else's system.

How do you specify this without a compliance project?

  1. Write the six records into the statement of work. One page. Not an annex, a requirement.
  2. Ask to see them in the demo, populated with a real run, before you sign.
  3. Require export. If the trail lives only in a supplier's dashboard, you do not have a record — you have a subscription to one.
  4. Set retention to match your file retention, not the supplier's default of ninety days.
  5. Test it once, deliberately. Pick a completed case at random and try to reconstruct the decision from the log alone. If you cannot, neither can anyone else.

The uncomfortable part

Building this makes the automation less impressive and more expensive. A confirmation step is friction, and friction is exactly what the demo was designed to remove.

That is the trade, and in a regulated firm it is not really a trade at all. The thing you are buying is not the time saved. It is the time saved that you can still account for.

Questions people actually ask

What should an AI agent log in a regulated firm?

Six records per action: the trigger, the inputs and their sources, the output, the named human who confirmed it, the difference between draft and final, and anything the agent refused or escalated and to whom.

Is "the AI decided" an acceptable answer to a regulator?

No. Consumer Duty asks a firm to evidence good outcomes and SM&CR asks who was responsible. Neither has a category for a process nobody owns, so accountability returns to the firm by default.

Why does logging refusals matter?

A guardrail nobody can observe is indistinguishable from no guardrail. If an agent has never refused anything, either the work is always in scope or the limit does not exist — and only the log tells you which.

Sources

  1. FCA Consumer Duty (in force 31 July 2023) and SM&CR — the accountability requirements this specification is written against. source ↗ — regulator, primary re-checked every 6 months
  2. The six-record specification is Noxia’s own and is not a regulatory standard. It is what we build to. — our opinion, labelled

Checked 8 September 2026. Next scheduled check 7 March 2027. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.

Ask us to demo the audit trail, not the agent.

We will show you the six records on a real run before you sign anything, and the export path out of our system on day one.

Talk to us about this