Free tool · Decision tool
Should you automate this? Six questions that settle it.
Not “should we use AI”, which is unanswerable, but “should this particular task be automated”, which is not. Six questions, in your browser, sending nothing anywhere. It is perfectly capable of telling you to leave a task alone permanently.
The short answer
Two questions decide most of it: if the task goes wrong, can you undo it, and would you notice. Reversible and visible is where automation belongs. Irreversible and invisible is where it never does. The other four questions — whether the task is written down, how often it happens, whether it turns on a judgement about a person, and whether a named person owns the output — decide the order and the guardrails.
“Should we use AI” has no answer, because it is not a question about anything. “Should this be automated” has an answer, and usually a short one.
The two axes below come from our note on what to automate first. The other four are the questions that stop a good idea failing for boring reasons: nobody wrote the task down, it happens twice a year, or nobody would own the output when it went wrong.
One task. Six questions.
Answer for one specific task, not for a department. “Automate the back office” has no answer; “package a case to this lender’s requirements” does.
—
—
Why “is it written down” is on the list
Because it is the most common cause of failure and the least discussed. A task that lives in one experienced person’s head is not a process, it is a performance. Automating it means writing it down first — and firms routinely discover at that point that three people were doing it three different ways, which is worth knowing whether or not you automate anything afterwards.
Half the value of an automation project is often the specification. That part is free and you can do it without us.
Why “a judgement about a person” is disqualifying
Not because a model cannot produce an output — it can, fluently. Because when the decision affects someone’s money, housing, health, employment or credit, the standard is not accuracy, it is accountability: a named person who decided, who can say why, and who can be challenged. That is a person by definition.
Automate the paperwork on both sides of that decision, all of it. Never the decision.
What we do with a “yes”
Scope it to verbs, not nouns. “Answer, qualify, book, hand over” rather than “handle enquiries”. Write the six audit records into the statement of work before anything is built — the specification is here. Agree the number you are trying to move, and review it at ninety days against that number rather than against a feeling.
And with a “no”
Then you have saved the money, which is the more common outcome and the one nobody publishes a tool for.
Questions people actually ask
What tasks should you not automate?
Any task where a mistake is both irreversible and hard to spot, and any task that turns on a judgement about a person — money, housing, health, employment or credit. In those cases the standard is accountability rather than accuracy, and accountability requires a named human. Automate the work either side of the decision instead.
How do you decide what to automate first?
Start where a mistake is reversible and would be noticed quickly, where the task is already written down, and where it happens often enough for the saving to be real. That combination is where current agents are reliable. Rare, undocumented, irreversible work is where projects fail.
Why does it matter whether a process is documented before automating it?
Because automating an undocumented process means writing it down first, and firms frequently discover at that point that several people were doing it differently. The specification is often more valuable than the automation, and it costs nothing but attention.
Does this tool send my answers anywhere?
No. It runs entirely in your browser and there is no server behind this page. The page counts its own views — an anonymous page count — the URL, the referrer, and a coarse country and device type, with no cookie and no identifier of any kind — and your answers are never part of that. Nothing is sent while you use the tool: turn your connection off once it has loaded and it still works.
Sources
- The two-axis framework (reversible × detectable) and the four supporting questions are Noxia’s own. They are not a standard, a regulation or a published methodology. — our opinion, labelled
- The reliability premise — that bounded tasks succeed and long-horizon cross-application work does not — rests on the OSWorld-Verified leaderboard (86.1% best agent, 72.36% human baseline) and TheAgentCompany, Carnegie Mellon, arXiv:2412.14161 (30% completed autonomously). source ↗ — public leaderboard and preprint, both primary re-checked every 30 days
- No answers, scores or results leave the page. The page counts its own views, anonymously and first-party, and nothing you click is part of that count. — verifiable in your browser’s network tab
Checked 8 September 2026. Next scheduled check 8 October 2026. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.
Got a “yes”? Bring us the one task, not the department.
One job, named, with the six audit records in the statement of work and a number agreed before anything is built.
Start the conversation