Platforms

Why we prefer Claude

Claude is the platform I recommend to most Australian organisations, and the one I run my own business on, because it finishes the work you give it: hand it a long task and it comes back done, which matters more than how well any single paragraph reads.

Updated

This article is not sponsored by Anthropic or Claude, and I have no financial interest in either. I have been a paying user of the platform since early 2025, and I run my business on it. The recommendation is based on that experience, and the argument is made from what I have actually built with it rather than from benchmarks.

The even-handed version, where Claude, Copilot and ChatGPT each get their case argued properly, is in Claude vs ChatGPT vs Copilot.

What changed in November 2025

I had been building with these models since early 2025, and until late that year the honest description of the output was sloppy. It worked often enough to be interesting and failed often enough that I checked every step, which meant no task was ever truly handed over. I was a ChatGPT user for years and had no reason to move.

Anthropic released Claude Opus 4.5 (opens in a new tab) on 24 November 2025 and the difference showed up immediately in ordinary use. Give it a task and it finishes the task, the way a competent professional does, and the output was consistently good rather than occasionally impressive. It was the first model over 80% on SWE-bench Verified (opens in a new tab), and it arrived with a 67% cut to the Opus price. I switched that month and did not agonise over it.

That is one practitioner's account, and the reason I trust it beyond my own enthusiasm is that the market moved the same way within months, with clients arriving already asking for the platform by name.

OpenClaw and what an assistant actually is

The second thing that changed came from outside the labs. OpenClaw (opens in a new tab) appeared in November 2025 as an open-source assistant that runs on your own machine and meets you in the apps you already use, so you text it and the work gets done. Every commercial product at that point still felt like a chatbot in a browser tab. This felt like an assistant, and the difference was presence: it was already in the apps I use, it kept running when I closed the laptop, and it could act on a request rather than answer it.

The idea won on every front. Anthropic shipped the same shape into its own product through messaging channels, OpenAI hired the creator in February 2026 to lead personal agents (opens in a new tab), and the project itself moved to a foundation. That last fact cuts against a lazy version of my own argument, because the assistant idea now belongs to everyone and both labs are chasing it. What Anthropic had was a year of building for the operator while the assistant was still a curiosity, and product direction of that kind takes longer to copy than a benchmark score.

Why a coding breakthrough matters at the board table

Most people file "best coding model" under engineering and stop reading, but the benchmark is measuring whether a model can hold a long chain of dependent steps, and a lot of board and finance work has exactly that shape. A long refactor and a month-end close are the same kind of problem, because both run through 40 dependent steps where an error at step 3 quietly corrupts everything after it while nobody watches. The improvement showed up in coding first because code fails loudly, since it either runs or it doesn't, so a wrong step surfaces straight away.

Board and executive work has no compiler, so the same capability arrived later and quieter in the places directors care about: a diligence read, a pack analysis, a reconciliation, a market scan that has to hold its own definitions across 60 documents. The consequence for a buyer is that the unit of work has moved from the answer to the task. A platform that writes a lovely paragraph and loses the thread at step 9 is worth less than one that finishes, and most evaluation exercises still test the paragraph.

What I run on it

I use Claude every day, in Claude Desktop and mostly in Claude Code. Claude Code is sold as a developer tool and works in practice as a general harness for getting things done, which is the single most useful thing I can tell a business reader about the platform, because almost nothing I do in it is software.

  • My financial position. Organised, reconciled and visualised, from the actual source files.
  • Overseas travel. Planned end to end, including the parts that involve reading 20 pages of conditions nobody reads.
  • Marketing and content. Drafted, edited and shipped against standards I wrote once and now reuse.
  • This website. Built and maintained in it, including the page you are reading.

The point of that list is that the tasks have nothing to do with each other, because a tool that handles finance, logistics, writing and software with the same set-up is working as an operating layer, and that is why our fluency work starts with configuring the desktop rather than teaching prompts.

Where it still loses

Three things irritate me, and they are all about buying and running the platform rather than about what the model can do.

  • Frontier price. Opus 5 runs at $5 per million input tokens and $25 per million output, Fable 5 at $10 and $50, against a capable GPT-5.6 tier at $2 and $12. At the working tier the prices have converged, but the best Claude work happens on the expensive models and that is where the multiple bites.
  • Commercial terms that will not sit still. 50% extra usage for a week, then extended, then permanent. Sonnet 5 priced as an introduction until 31 August, then made permanent (opens in a new tab). Fable 5 (opens in a new tab) free for Pro subscribers, then not, then permanent for higher tiers at half the usual limits. Each of these is defensible alone and the pattern is tiring, and combined with usage limits that are hard to translate into a forecast, it means a CFO cannot tell you what a heavy user costs in November. That is a procurement problem, and procurement is how organisations actually buy.
  • Refusals on legitimate work. It still declines requests it should handle, less often than it used to but often enough to notice, and it lands hardest on exactly the regulated and sensitive material some teams work with all day.

What would change my mind

Three things, and I would say so publicly if any of them happened. A competitor matching completion rate on long multi-step work at the current price gap, which is the most likely of the three and would end the argument on economics alone. Anthropic pricing or rate-limiting the platform out of reach of the mid-market, which the last year of terms makes worth watching. Or a serious governed path inside the Microsoft estate that removed the boundary advantage Copilot currently holds, which would change the recommendation for a large share of Australian organisations without changing anything about the model.

Until one of those happens, the recommendation stands, dated, because the reason has held all year: it finishes the work you give it. If you want the version of this argument where the other two get to answer back, read Claude vs ChatGPT vs Copilot.