Noxia

Industries · Financial services

AI agents for UK mortgage brokers, IFAs and wealth managers

The work either side of the advice — packaging a case, drafting the report a human then challenges, chasing the document that is holding everything up. Not the advice. Not the suitability decision. Not anything a person has to own.

Sources checked 8 September 2026 1130 words

The short answer

Noxia builds AI agents for UK mortgage brokers, IFAs and wealth managers. They do the work on either side of advice — packaging a case for a specific lender, drafting a suitability report for a human to challenge, chasing the outstanding document — and they write an evidenced record of every step. They do not give advice, assess suitability, or decide anything a person must own.

Every firm in this market has the same shape of problem and almost none of it is the advice. The advice is the part your people are good at and the part they are paid for. The problem is everything that has to happen before and after it, and how much of that is chasing.

In May 2026 the messaging firm Nivo published research from conversations with more than seventy UK specialist lenders and brokers. One secured loan case: fifteen to twenty rounds of messages, more than five people, upwards of five hours of administration on the broker-to-lender leg alone, and a thirty-five per cent chance of coming back for corrections anyway. Half of submitted information was right first time.

Where a case actually goesA typical secured loan case involves 15 to 20 rounds of messages, more than five people, over five hours of broker-to-lender administration, with 35 per cent of cases requiring corrections.50%right first timesubmitted information15–20message roundsper case5+ hrsadminbroker→lender leg35%need correctionsor repeatsWHERE A CASE ACTUALLY GOESNivo, May 2026 · 70+ UK specialist lenders and brokers · vendor research, method and sample stated
None of that is advice.Nivo, reported 7 May 2026. Vendor research, but with a named method and sample size.

What we automate

The jobWhat the agent doesWhat stays human
Case packagingAssembles the pack to one named lender's current requirements, flags what is missing before submissionThe submission decision, and anything the lender's underwriter asks that is a judgement
Suitability draftingProduces a first draft from the fact-find and the research, with every figure traced to the document it came fromThe suitability assessment itself. The adviser challenges the draft and owns the file
Document chasingAsks the client for the outstanding item, in the client's channel, until it arrives or a human takes overAnything the client asks that is not "where do I send this"
Enquiry handlingAnswers, qualifies, books into the real diary, hands to a named personThe conversation the moment it stops being logistics
File review preparationAssembles the evidence for a file check and lists what is not thereThe review, and every conclusion in it

The pattern is the same in all five: the agent moves paper and the human owns the decision. That is not caution for its own sake — it is the only shape that survives a file check, and it is the shape the benchmarks say actually works.

Why we scope it that narrowly

Because the honest reading of the evidence demands it. On OSWorld — 369 bounded desktop tasks — the best agent now scores 86.1%, above the 72.36% human baseline. On TheAgentCompany, a simulated firm whose work runs long and spans applications, the most competitive agent completes 30% of tasks unaided.

The same technology, two task shapesBest agent on OSWorld scores 86.1 per cent against a human baseline of 72.36 per cent. On TheAgentCompany, long-horizon cross-application work, the most competitive agent completes 30 per cent of tasks autonomously.OSWorld — bounded desktop tasks, best agent86.1%above the human baselineOSWorld — human baseline72.4%original OSWorld paperTheAgentCompany — long-horizon, cross-application30%most competitive agent, unaided
Scope is the whole product.OSWorld-Verified leaderboard, re-checked 8 September 2026. TheAgentCompany, Carnegie Mellon, arXiv:2412.14161.

Anyone selling you an agent that runs a case end to end is selling you the 30% number and quoting the 86% one. We would rather sell you five tasks that land in the 86% shape and tell you plainly where the other job is. There is a longer version of this argument in the note on the gap between those two numbers.

What we will not automate

  • The advice. Not the recommendation, not the suitability assessment, not the risk profile.
  • Anything a named person has to own under SM&CR. If a regulator would ask who decided, a person decided.
  • Client communication after it stops being logistics. "Send me your last three payslips" is logistics. "Should I fix for five years" is not.
  • Vulnerability assessment. Every credible signal here is contextual, and a missed one is a Consumer Duty failure, not an efficiency loss.
If an agent could safely give the advice, you would not need the firm. The reason this vertical is a good market is exactly the reason the scope is narrow.

The record every agent writes

Consumer Duty turned "we gave good advice" into "we can evidence that we gave good advice", and an automated step with no record is worse than a manual one, because at least the manual one had a person who remembers. Six records per agent action, in the statement of work, not an annex:

RecordContainsAnswers
TriggerEvent, timestamp, which rule firedWhy did this happen at all?
InputsEach document, its origin, its version and dateWhere did that figure come from?
OutputThe artefact, immutable, with a hashIs this the version that was sent?
Confirmed byNamed individual, timestamp, what they sawWho is accountable?
DiffWhat the human changed before it wentDid anyone actually read it?
RefusedWhat it would not do, and who it went toDid the guardrail work, or was it decorative?

The refusal record is the one nobody asks for and the one that matters most: it is the only evidence that a guardrail is load-bearing rather than decorative. The full specification is here, including how to test it on a real case before you sign.

How we price it

Flat monthly, per scope. Never per minute, per message or per case. Usage pricing makes your busiest month your most expensive one and makes the quote impossible to compare against a salary — which is the comparison you are actually making. Why per-minute pricing punishes growth.

What we would tell you not to buy

Most firms who ring us about AI reception have been shown a frightening statistic about missed calls. It is somebody else's statistic. Before you buy anything — from us or anyone — do the arithmetic on your own call log:

Open the missed-call calculator

It runs entirely in your browser, sends nothing anywhere, and is perfectly capable of telling you the answer is "do not buy this". The method behind it is here.

How an engagement starts

  1. One job, named. Not "automate the back office". "Package a case to this lender's current requirements."
  2. The six records in the statement of work, before anything is built.
  3. Four to six weeks to a working agent on real cases, with a human confirming every output.
  4. Ninety-day review against the number we agreed at the start. If it did not move, we say so and you stop paying.

We publish what does not work, including our own. Here is the audit where we checked five claims on our own site and removed three of them, including one that credited Gartner with a figure Gartner never published.

Questions people actually ask

Can AI write a suitability report for an IFA?

It can produce a first draft from the fact-find and the research, with every figure traced back to the document it came from. It cannot make the suitability assessment. The adviser challenges the draft, changes what is wrong, and owns the file — and the record shows what they changed, which is the evidence a file check actually wants.

Is using AI compliant under Consumer Duty?

Consumer Duty does not prohibit automation; it requires you to evidence good outcomes. That makes the record the compliance question, not the technology. An automated step that writes trigger, inputs, output, confirmation, diff and refusal is easier to evidence than a manual one nobody logged. An automated step with no record is worse than doing it by hand.

What can AI agents actually do reliably for a mortgage broker?

Bounded, well-specified tasks: assembling a pack to one lender’s stated requirements, chasing a named document, answering and booking an enquiry, preparing evidence for a file review. On bounded desktop tasks the best agents now beat the human baseline. On long-horizon work spanning applications they finish about 30% unaided, which is why scope decides whether this works.

Will an AI agent talk to my clients?

Only about logistics, and only where it can identify itself as an assistant and hand to a named person the moment the conversation stops being logistics. Anything touching advice, suitability or vulnerability goes to a human immediately, and the handover is one of the six records.

How much does this cost?

Flat monthly per scope, quoted against the job rather than usage. We will tell you when the arithmetic does not support buying anything — the missed-call calculator on this site exists to help you reach that conclusion without us.

Sources

  1. Nivo, research with 70+ UK specialist lenders and brokers, reported by Mortgage Solutions, 7 May 2026 — 50% right first time, 15–20 message rounds, 5+ people, 5+ hours broker-to-lender admin, 35% corrections. — vendor research, named method and sample re-checked every 6 months
  2. OSWorld-Verified leaderboard — 86.1% best agent (Qwen3.8-Max) across 369 tasks, against the 72.36% human baseline stated in the original OSWorld paper. Leaderboard data verified 4 September 2026; re-checked by us 8 September 2026. source ↗ — public leaderboard, primary re-checked every 30 days
  3. TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks, Carnegie Mellon, arXiv:2412.14161 — "the most competitive agent can complete 30% of tasks autonomously". source ↗ — peer-reviewable preprint, primary
  4. FCA Consumer Duty, in force 31 July 2023, and SM&CR — the accountability requirements the six-record specification is written against. source ↗ — regulator, primary re-checked every 6 months
  5. The six-record specification and the four-item exclusion list are Noxia’s own. Neither is a regulatory standard. They are what we build to. — our opinion, labelled

Checked 8 September 2026. Next scheduled check 8 October 2026. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.

One job, named, with the six records in the statement of work.

Tell us the job that is eating the hours. If the arithmetic does not support building it, we will say so in the first conversation rather than the fourth.

Start the conversation