The decision an organisation faces is narrower than the reasoning, writing and coding scores most comparisons run on: which platform do we standardise on, what do we do about the licences we already pay for, and what does it cost us if we are wrong. Disclosure: Veriet runs on Claude and recommends it in most engagements, and the three cases below are still written to win, with the Copilot one the case I have argued hardest in client rooms.
How each platform arrived without a decision
Start with how each platform arrived, because in most organisations none of them was selected. Copilot arrived through the Microsoft agreement, bundled into an enterprise renewal where the uplift looked modest against the total, and the seats landed months before anyone designed a way of working around them. ChatGPT arrived through staff, on personal accounts and personal credit cards, before there was a policy telling them not to. Claude arrived through the person who reads release notes, normally one engineer or one analyst.
So all three arrived without anyone deciding, which matters because an organisation that never chose cannot explain why it is where it is, and it will not be able to explain the next move either. What has shifted is that buyers now arrive with a preference of their own.
Every engagement I have taken in 2026 has come to me asking for Claude. Through 2025 they were asking for ChatGPT.
That is 6 engagements, so the sample is small and it is mine, but the direction matters more to me than the count.
The case for Copilot
Copilot is the only one of the three that is already inside your compliance boundary, which for a regulated Australian organisation is most of the argument. It reaches your documents, mail, meetings and chat without a connector, it inherits the permissions and sensitivity labels you have already spent years getting right, and a person using it sees exactly what that person could already see, whereas the other two reach the same data through connectors you have to assess and govern yourself.
The governance questions are also answered before anyone asks them. Tenant, retention, eDiscovery, residency and the training terms were all settled in an agreement your legal team has already read, so there is no new vendor assessment, no new data processing agreement and no new board paper. The seats are paid for, which means the marginal cost of the rollout is close to zero, and the alternative is paying twice.
At its strongest the argument is that a governed tool people use returns more than a better one nobody has approved, and most disappointment with Copilot traces back to the rollout rather than the product, because the licences arrived through a procurement decision before anyone designed the ways of working around them, so the second half of the job was never done. Out of the box an enterprise seat knows nothing about your systems, your data or your standards, and most people stop asking after the third mediocre answer, so that gap is configuration and working habits, and it closes quickly enough that I would fix it before concluding anything about the platform.
The case for ChatGPT
ChatGPT took the world in a way no enterprise software ever has, which means your people already know it. That is the cheapest change management available to you, and change management is where most of these programmes fail, since nobody needs convincing to open the tool they already use at home.
It also wins on price at volume, and by a margin that grows with headcount. GPT-5.6 Terra runs at $2 per million input tokens and $12 per million output, which undercuts the Claude tier doing equivalent work, and the gap widens against the frontier models where the best Claude work happens. At 2,000 seats that difference is a line item somebody will notice. It also carries the widest surface area of the three, with voice, images, data analysis and the largest third-party ecosystem.
I weigh the pace argument highest, because OpenAI has closed every capability gap it has fallen behind on, repeatedly, so I would build no 3-year plan on the assumption that it stays behind. If you want one platform your whole workforce will use, at a cost you can defend to a board, ChatGPT is a defensible answer.
The case for Claude
The case for Claude is that it finishes long, multi-step work with the fewest interventions. What improved through late 2025 was reliability across chains of dependent work: the 40-step task where an error at step 3 quietly corrupts everything after it, which is the shape of a month-end close, a diligence read or a pack analysis. It holds its own definitions across 60 documents, and it runs where the work happens, in the terminal, the editor and the desktop app, instead of waiting in a browser tab.
A year of running a business on it, and what changed to make that possible, is in Why we prefer Claude. The short version is that it finishes the task more often than the other two.
The three side by side
| Copilot | ChatGPT | Claude | |
|---|---|---|---|
| Where it sits | Inside the Microsoft agreement you have already signed | A new vendor assessment and data agreement | A new vendor assessment and data agreement |
| Access to your data | Documents, mail, meetings and chat with no connector, inheriting your permissions and labels | Connectors you assess and govern yourself | Connectors you assess and govern yourself |
| Price | Bundled in the agreement, a sunk cost in most estates | GPT-5.6 Terra at $2 and $12 per million tokens | Sonnet 5 (opens in a new tab) at $2 and $10, Opus 5 at $5 and $25, Fable 5 at $10 and $50 |
| Strongest at | Work that lives inside Microsoft 365 | Adoption, breadth and unit cost at volume | Long multi-step work finished without intervention |
| Where it loses | Tasks that span systems or run long, and the slowest release pace of the three | Consistency across long work, and guardrails you build yourself | Frontier price, usage limits you cannot forecast, and refusals on sensitive material |
| Best fit | Deep Microsoft estates with residency and labelling already settled | High volume work that is well specified and cheap to check | High stakes work that is expensive to check |
Standardising on Copilot also ties your AI capability to one vendor's release pace, and ChatGPT is pulled toward the consumer because that is where the larger market is, so both trade away something you would otherwise control.
The list prices have converged at the working tier and the gap opens at the top, where the best Claude work runs, so what you pay depends on the work you are buying for. High volume, low stakes, well specified and cheap to check, meaning classification, summarising, first drafts and extraction at scale: buy on price, because paying frontier rates there is waste. Low volume, high stakes, long chains and expensive to check, meaning board papers, diligence and anything where a quiet error surfaces 3 weeks later: buy on completion rate, because the cheaper model that needs 3 attempts and a careful human review is the more expensive one.
How to run the decision
Run the comparison on real work over 2 weeks, and 4 steps get you an answer you can take to a board.
- Settle the data boundary first. Decide what company, customer and personal information may leave the business and under what terms. That narrows the field before any demonstration, and it is the part a director will be asked about.
- Pick 3 pieces of real work. One long analytical task, one recurring operational task, and one piece of writing that carries your standards, all with the real inputs and real deadlines the work normally carries, because cleaned-up samples hide the problems you are choosing a platform to handle.
- Give every platform the same context. The same instructions, the same source documents, the same definition of done. Most comparisons quietly hand the incumbent a configured environment and the challenger a blank page.
- Measure completion rather than sentiment. Count how often each platform finished without intervention, how many corrections every attempt needed, and what the fully loaded cost becomes at 10 times the volume, because enthusiasm in the room tells you how the demo went rather than how the work will go.
The verdict, as at August 2026
Claude, for most Australian organisations, because it finishes long work with the fewest interventions and that is what decides whether people keep using a tool after the demo. The margin is roughly a year old, it is narrower than its advocates suggest, and it is the kind of margin one release can close.
2 groups should not follow the recommendation.
- Deep Microsoft estates where the boundary is already settled. If the agreement is signed, residency is answered and labelling is in place, stay on Copilot. The switching cost belongs in the decision as a number, and where your data sits under each vendor is its own question that What Australian compliance requires answers.
- Organisations where per-seat cost is the binding constraint. At scale, on work that is well specified and cheap to check, ChatGPT is the answer a CFO can defend and I would not argue with it.
Choosing so the decision survives the answer changing
The competition is fierce, GPT-5.6 is strong at a fraction of frontier pricing, and the order could change with one release. To me that argues for choosing well rather than waiting, because the organisations that waited through 2025 ended up with whatever their vendor agreement and their staff chose for them.
Make the decision survivable by keeping the expensive part portable. Your instructions, context, standards and worked examples are the asset, so hold them as documents you own rather than as configuration trapped in one vendor's console. Settle the data boundary separately from the model, since the boundary question outlives any model. Avoid multi-year exclusivity where the commercial terms let you. Then review the choice on a fixed annual cadence instead of every time a competitor posts a benchmark.
Building the capability to change platforms cheaply is worth more than the platform choice itself, because once you have it the next verdict costs you a week.