Noxia

Field note · Buying AI

Software that sells, and software that just costs.

Three quarters of UK financial services firms already use AI. Most of them cannot tell you what it earned them. That is not a failure of the software, it is a failure of the question asked before buying it, and the question is a short one.

4 min read Sources checked 20 September 2026

The short answer

Software pays for itself when the buyer can name, before purchase, one number it will move, who already counts that number, and who will check the output. Software that only costs is bought against a category rather than a number — "we need AI", "everyone has one". The test is not the technology or the price: it is whether the number being moved was already being measured before the tool arrived.

On this page · 6 sections

The adoption figures are settled. The Bank of England and the FCA surveyed 118 firms and found 75% already using AI, with another 10% planning to within three years. The ONS puts wider UK business adoption lower, at around one in three. Nobody needs another article telling them AI is here.

The interesting question is the one underneath: of the firms that bought, which ones can point at what it did?

The one-sentence test

Before buying anything, complete this sentence honestly:

"This moves [a number], which [a named person] already counts, and [a named person] checks its output."

Three blanks. Each one fails differently, and each failure has a smell.

No number. The purchase is defensive — bought because a competitor has one, or because a board asked. These are the tools that get renewed for three years and opened twice. They are not always wrong; sometimes the defence is real. But they should be bought with that named as the reason, and budgeted as marketing rather than as efficiency.

A number nobody counts yet. The most common failure, and the most avoidable. If you cannot state today's value, you will never be able to demonstrate the improvement — and in twelve months the renewal conversation will be a discussion of vibes. The fix is free: start counting before you buy. A fortnight of tallying is worth more than any pilot.

Nobody checks the output. Either the output does not matter, in which case reconsider the whole purchase, or it does and the supervision cost has been left out of the business case. For anything client-facing or regulated that cost is not optional — the safeguards that came into force in February require a route to human intervention on significant automated decisions, and a route is a rota.

Two purchases, same priceSoftware that sells is attached to a number that was already being counted, to a named person who counts it, and to a named person who checks its output, with a measurable before and after. Software that only costs is attached to a category rather than a number, has no baseline, no named owner and no checker, and is renewed on the strength of an impression.Software that sellsthe sentence completesMoves a number counted before it arrivedA named person owns that numberA named person checks the outputA before, and therefore an afterSoftware that only coststhe sentence stallsBought against a category, not a numberNo baseline, so no proof either wayNobody owns the outcomeRenewed on an impressionTwo purchases, same priceSoftware that sells is attached to a number that was already being counted, to a named person who counts it, and to a named person who checks its output, with a measurable before and after. Software that only costs is attached to a category rather than a number, has no baseline, no named owner and no checker, and is renewed on the strength of an impression.Software that sellsthe sentence completesMoves a number counted before it arrivedA named person owns that numberA named person checks the outputA before, and therefore an afterSoftware that only coststhe sentence stallsBought against a category, not a numberNo baseline, so no proof either wayNobody owns the outcomeRenewed on an impression
The difference is in the sentence, not the software.Noxia’s own test.

Why case admin is the standard example

The clearest wins we see are boring, and they are boring for a reason: the number was already there.

Research published by Nivo in May 2026, from conversations with more than seventy specialist lenders and brokers, found that only half of submitted information is right first time, that a typical secured loan case takes fifteen to twenty rounds of messages, and the broker-to-lender leg alone consumes upwards of five hours of admin, with a 35% chance of coming back for corrections.

Every one of those is already counted by somebody, usually in their head and bitterly. That means the sentence completes on the first try: this moves the correction rate, which the case manager already tracks, and the case manager checks the output. A firm buying against that can tell in six weeks whether it worked. The full breakdown is here, and the system we build against it is deliberately narrow.

The two honest exceptions

Buying a capability you do not have. Answering a question you currently cannot answer at all has no baseline by definition. That is legitimate — but say so, and set a date by which the new number will exist.

Buying to remove a risk. An insurer's requirement, a regulator's expectation, a single-person dependency. The value is the tail you cut off, and you will never observe it. Also legitimate; also worth writing down as the reason, so nobody judges it on efficiency later.

What is not legitimate is buying against a number, failing to move it, and quietly re-describing the purchase as one of these two afterwards.

What this does not tell you

It does not tell you the tool will work. A perfectly framed purchase can still fail on the product, the data or the people. The test filters out the purchases that cannot be judged, which is a different and larger category than the purchases that fail.

It also does not settle whether the saving is big enough to be worth having. A tool can move a real, counted number and still lose money, because adoption costs eighty hours before the subscription — that is the £200 floor. Run this test first to see whether the purchase is judgeable, then run the floor to see whether it is worth judging. The triage tool does both in six questions, and it says no more often than yes.

Questions people actually ask

How do you know if business software will pay for itself?

Complete one sentence before buying: this moves a named number, which a named person already counts, and a named person checks its output. If any of the three blanks will not fill, the purchase cannot be judged after the fact — which is different from saying it will fail, but it means you will never know.

What percentage of UK financial services firms use AI?

The Bank of England and FCA survey published on 21 November 2024 found 75% of the 118 firms surveyed already using AI, with a further 10% planning to within three years. Wider UK business adoption measured by the ONS is lower, at around one in three.

Should you ever buy software without a measurable return?

Yes, in two cases: buying a capability you do not currently have at all, where no baseline can exist by definition, and buying to remove a risk such as an insurer requirement or a single-person dependency, where the value is a tail you will never observe. Both are legitimate if named as the reason at the time, rather than reached for afterwards when the efficiency case fails.

What should you do before buying an AI tool?

Start counting the number you intend to move, before you buy. A fortnight of tallying costs nothing and is worth more than a pilot, because without a baseline the renewal conversation twelve months later has no evidence in it at all.

Sources

  1. Bank of England and Financial Conduct Authority, “Artificial intelligence in UK financial services – 2024”, 21 November 2024: of 118 firms surveyed, 75% already using AI and a further 10% planning to within three years; 55% of use cases involved some automated decision-making; 46% of firms reported only partial understanding of the AI they use. bankofengland.co.uk ↗ — primary; the regulators’ own survey re-checked every 6 months
  2. Nivo, research published 7 May 2026 from in-depth conversations with more than 70 specialist lenders and brokers, reported by Mortgage Solutions: only half of submitted information right first time; for a typical secured loan case, 15–20 rounds of messages, more than five people and upwards of five hours of admin on the broker-to-lender leg; 35% of cases requiring corrections. mortgagesolutions.co.uk ↗ — vendor research, but with a named method and sample size, and used here only as an example of numbers a firm already tracks re-checked yearly
  3. The one-sentence test, the three failure modes and the two exceptions are ours. No survey underlies them; they are a description of the purchases we have watched succeed and fail. — our own argument, labelled as such

Checked 20 September 2026. Next scheduled check 19 March 2027. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.

Cite this note

Noxia, “Software that sells, and software that just costs”, Field notes, 20 September 2026; sources checked 20 September 2026. https://www.noxia.co.uk/field-notes/software-that-sells

Try the sentence on the thing you are about to buy.

If it will not complete, send it to us anyway — working out which number a job actually moves is most of the value, and it is the part vendors skip because it is the part that can say no. We will tell you what to start counting, and whether it is worth counting at all.

Talk to us about this