Field note · What agents can do
AI agents now beat humans at desktop tasks. They still finish 30% of real work.
Two benchmarks, two very different answers, and almost every AI pitch you receive quotes the flattering one. Here is what each actually measures, and which one describes your firm.
The short answer
On OSWorld — 369 bounded desktop tasks — the best agent scores 86.1%, above the 72.36% human baseline. On TheAgentCompany, a simulated firm with long-horizon work spanning applications, the most competitive agent completes 30% of tasks autonomously. Same technology, different job shape. Scope decides which number you get.
You will be shown the 86%. You will not be shown the 30%. Both are real, both are current, and the difference between them is not a difference in the models.
What is OSWorld actually measuring?
369 tasks on a real desktop, each one execution-verified: open this, change that, save it there. Bounded, single-sitting, checkable by a script.
An agent scoring 86.1% on that is genuinely impressive, and it is above what the original paper measured humans at. If your work looks like those tasks, buy the thing.
What is TheAgentCompany measuring?
A simulated company. Software, HR and finance work that runs long, crosses applications, and requires the agent to notice that something earlier went wrong.
The Carnegie Mellon paper's own summary is blunt: the most competitive agent completes 30% of tasks autonomously, and "more difficult long-horizon tasks are still beyond the reach of current systems."
Why does the number collapse?
Because errors compound and nobody is watching.
A bounded task fails visibly — the file did not save, you can see it. A long-horizon task fails at step four and produces a confident, complete-looking result at step eleven that is built on the step-four mistake. Nothing errors. Something is simply wrong.
| What is being compared | Bounded task | Long-horizon task |
|---|---|---|
| Example | Reformat this spreadsheet | Package and submit this case |
| Steps | Under ten, one sitting | Dozens, over days |
| Failure looks like | An error you can see | A plausible wrong answer |
| Who notices | You, immediately | The lender, in a week |
| Benchmark score | 86.1% | 30% |
So how do you get the 86% number in a real firm?
By refusing to buy the long-horizon version. You cut the long task into bounded ones and put a named person at each join.
- Find the joins. Where does the work change hands, systems or days? Those are your checkpoints, and they already exist informally.
- Scope each agent to one leg. Not "handles the case". "Assembles the pack for lender X and stops."
- Make the handover explicit. A person confirms, or it does not proceed. This is the whole difference between the two columns above.
- Log every action. When something is wrong in a week, you need to see which leg produced it.
- Measure per leg. An agent that is right 95% of the time on one bounded leg is worth buying. The same agent trusted end-to-end is not.
What should you ask a vendor?
One question: on what benchmark, and is the task shape like mine?
If the answer is a percentage with no benchmark named, it is marketing. If the benchmark is bounded and your work is long-horizon, the number is real and irrelevant. Neither is dishonest exactly — but only one of them is about you.
Questions people actually ask
How accurate are AI agents in 2026?
It depends entirely on task shape. On OSWorld, 369 bounded desktop tasks, the best agent scored 86.1% as of August 2026, above the 72.36% human baseline. On TheAgentCompany, which tests long-horizon work across applications, the most competitive agent completed 30% of tasks autonomously.
What is OSWorld?
A benchmark of 369 execution-verified desktop tasks used to measure computer-use agents. Because the tasks are bounded and checkable, scores on it are much higher than on benchmarks that test long, multi-application work.
Why do AI agents fail at long tasks?
Errors compound without being visible. A mistake at step four produces a confident, complete-looking output at step eleven. Nothing throws an error, so nobody notices until a downstream party does.
Sources
- OSWorld-Verified leaderboard — 86.1% (Qwen3.8-Max), top five 83.4–86.1%, against the 72.36% human baseline stated in the original OSWorld paper for its 369 tasks. Leaderboard data verified 4 September 2026; re-checked by us 8 September 2026. source ↗ — public leaderboard, primary re-checked every 30 days
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks, Carnegie Mellon, arXiv:2412.14161 — "the most competitive agent can complete 30% of tasks autonomously". source ↗ — peer-reviewable preprint, primary
Checked 8 September 2026. Next scheduled check 8 October 2026. Numbers that move — leaderboards, live indices — are re-checked every 30 days; annual datasets and rules in force every six months; dated research once a year. If something here has gone stale before we got to it, tell us and we will correct it and say what changed.
Want your long process cut into legs an agent can actually finish?
That is the work: find the joins, scope one leg, put a person on the handover. We quote it as a fixed build, and you can measure it per leg.
Talk to us about this