Every new car on a US lot carries a number that was produced indoors. Fuel-economy estimates come out of a laboratory, where the drive wheels sit on a dynamometer and a federal test procedure decides what counts as driving (fueleconomy.gov). The number is real, measured, and honestly obtained. It is also the reason the phrase your mileage may vary had to be invented.

Claude Opus 5 shipped Friday with the same breed of number, and explaining that well is worth more than the news itself. Because sometime this week — in a status email, in a proposal, on a call — someone will tell you their team already upgraded to the latest model, that it beats everything, and that it costs half. All of that is defensible. It is equally defensible that the outside labs which have measured it so far are reporting ties.

Both fit inside the same week. The distance between them is what a buyer needs to understand to avoid buying a headline.

What did Anthropic actually announce on July 24?

Start with what carries a name and a date. Anthropic released Claude Opus 5 on July 24, 2026, at the same token price as the model it replaces and below what its own flagship, Claude Fable 5, costs. The exact rates are tabulated further down, in the section on cost (Anthropic's pricing page, read July 26, 2026).

On capability, the central claim of the announcement is verbatim: "On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks" (official announcement). The sentence names its own exception, which is more than most launch copy manages.

The technical documentation adds what can be checked line by line: a far larger context window, reasoning switched on by default, an effort ladder that runs from low to max, and a fast mode that trades money for speed (official Opus 5 documentation).

None of that is in dispute. It is a company describing its own product, with its name on it.

What did the independent evaluators find?

Three organizations measured it on their own within days of the release, and none of them landed where the announcement did.

Epoch AI, the most conservative of the three, placed it below Fable 5 on its general capabilities index and level with it on the software-engineering component (reported by Yellow, July 2026).

Artificial Analysis published the friendliest read, and also the most uncomfortable line inside it: factual accuracy improved while the hallucination rate moved the wrong way (Artificial Analysis, July 24, 2026). Worth knowing before you weigh it: this shop states openly that it supported Anthropic in evaluating the model ahead of release. That is not disqualifying. It is context a reader deserves to have on the table.

CodeRabbit reviews other people's code for a living, and its test was the closest thing here to a workday: error patterns pulled from verified issues in real open-source pull requests, each configuration run several times against the baseline mix it already has in production (CodeRabbit, July 24, 2026). The new model wrote more accurate comments, caught fewer of the known bugs, and generated four times the noise.

That contrast deserves a second read. Anthropic's own documentation lists among Opus 5's capability improvements "code review and bug-finding, surfacing real bugs at a high rate per pass with few false positives." The company that reviews code for a living measured fewer findings and more noise than its own baseline. Both statements are checkable. They point in opposite directions.

Here it is in one place, with no intermediary:

What was measured Who measured it Result
Coding and knowledge work (Frontier-Bench, GDPval-AA) Anthropic "The new state-of-the-art." Figures published as chart images, not as text tables
CursorBench 3.2 (max effort) Anthropic "Within 0.5%" of Fable 5's peak score — a technical tie
Cybersecurity tasks Anthropic States it stays behind its own Mythos 5
Capabilities Index (ECI) Epoch AI Opus 5: 159 · Fable 5: 161 → Fable 5 ahead
Software engineering (SWE-ECI) Epoch AI Tied. No component scores were disclosed
Intelligence Index (max effort) Artificial Analysis Opus 5: 61 · Fable 5: 60 · GPT-5.6 Sol: 59 → technical tie
AA-Briefcase, long agentic work (Elo) Artificial Analysis Opus 5: 1720 · Fable 5: 1574 → Opus 5 ahead
Humanity's Last Exam Artificial Analysis 53% for both → tie
Terminal-Bench v2.1 (max effort) Artificial Analysis Opus 5: 89%, level with leader GPT-5.6 Sol
Factual knowledge and hallucination (AA-Omniscience) Artificial Analysis Accuracy +7 points, but the hallucination rate rises 14 points to 50%; still behind Fable 5 on factual knowledge
Known-bug detection (~100 error patterns from real PRs, 3 runs each) CodeRabbit Opus 5 (x-high): 55.2% · production baseline: 61.1%
Precision of actionable comments (same test) CodeRabbit Opus 5: 39.3% · baseline: 35.2%
Noise, "nitpicks" (same test) CodeRabbit Opus 5: 92 · baseline: 23
SWE-bench Pro, public board with a standardized harness Scale AI (SEAL) Neither Opus 5 nor Fable 5 appears on the board

Sources: Anthropic and its Opus 5 documentation; Epoch AI via Yellow; Artificial Analysis; CodeRabbit; Scale AI's public SWE-bench Pro board. All consulted July 26, 2026.

How can both versions be true at the same time?

This is the part almost nobody will explain to you, and it is the reason this article exists.

A language model is never evaluated by itself. It is evaluated inside a harness: the program that hands it the problem, gives it tools, lets it retry, decides how many steps it gets and when it gives up. The harness is not packaging. It can account for half the result.

When the company that trained a model publishes a score, it normally runs that model inside a harness it built and tuned for it — the laboratory version, with the test cycle written around the car. When a neutral organization measures the same model, it drops it into a standardized rig, identical for every model, precisely so the numbers can be set side by side. The two do not land in the same place.

The distance is documented. An analysis published on June 16, 2026 cites Scale AI's own work putting the swing from harness choice at 10 to 20 points on the same test, with the model untouched. In that same analysis: of the hundred models listed on one popular leaderboard that day, all but one carried scores submitted by the vendors themselves (Digital Applied, June 2026).

Two scores from "the same test" are therefore not necessarily comparable. Subtracting one from the other is the arithmetic of measuring one car on the dyno and the other on a January commute, then announcing a winner.

There is a second reason, simpler, and this one we checked ourselves. Anthropic published nearly all of its figures as chart images rather than text tables. The body of the announcement carries relative claims — beats every model, within a fraction of a percent of the peak — while the numbers live inside the pictures. We opened the official page to confirm it, and it matches what Vellum reported on launch day: "Anthropic's charts for these benchmarks are published as images on the announcement page, not readable comparison tables."

The practical consequence is awkward. A good share of the exact figures circulating this week came from someone looking at a chart and estimating a number, not from tabulated data. That is why Opus 5 figures contradict each other from one aggregator to the next, and why several numbers you have probably already seen on LinkedIn are missing from the table above. We left them out deliberately: we could not confirm them against the original evaluator.

The third reason is the human one. The company that trains a model and the outfit that measures it from outside are not answering the same question. Anthropic answers "what is this model capable of on its best day?" The independent evaluator answers "what does this model do when I put it on the work I am already doing?" Both questions are legitimate. The second one is yours.

Where does Opus 5 lose?

Neither side hides this, but it is worth laying out.

Anthropic itself checks one box: cybersecurity, where it stays behind Mythos 5. Artificial Analysis adds one that was not in the script — a hallucination rate moving the wrong way, with factual knowledge still behind Fable 5.

Four more categories surfaced in the readings that well-known creators in the space did of the announcement's charts. We take those for exactly what they are — a third party's interpretation of images published by Anthropic, not a measurement of their own — and that is why we report them without figures:

We stop there on purpose. Which of those categories matters depends on what your company does, and that part we don't know.

What nobody can claim yet

Neither Opus 5 nor Fable 5 appears on Scale AI's public SWE-bench Pro board, the standardized rig that runs every model under identical conditions. Three days after launch, that is normal.

It still means something concrete: the industry's most-cited coding test has no neutral measurement of Opus 5. Anyone telling you otherwise this week is quoting somebody's harness.

Did the price drop, or will your bill drop?

Entirely depends on where you are coming from, and that is where most of the confusion this month is going to happen.

Item Claude Opus 5 Claude Fable 5 Claude Sonnet 5
Price per million tokens (input / output) $5 / $25 $10 / $50 $2 / $10 through Aug 31, 2026; $3 / $15 after
Fast mode (research preview) $10 / $50 Not listed Not listed
Batch API (50% discount) $2.50 / $12.50 $5 / $25 $1 / $5; then $1.50 / $7.50
Average cost per task, Intelligence Index (max effort) $2.03 $2.75 (with fallback) $1.53

Prices in USD. Sources: Anthropic's pricing page and Artificial Analysis, consulted July 26, 2026. For reference, the predecessor Opus 4.8 averages $1.80 per task on that same index. Two details that rarely get mentioned: output token consumption varies roughly eightfold between the lowest and the highest effort setting (Artificial Analysis), and Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text than 4.6 and earlier models do (Anthropic documentation) — so comparing per-token prices across generations understates real spend.

Read the last row slowly. "It costs half" is true against one specific model: the flagship. If your vendor was running that, there is a saving. If they were running the mid-tier model from the same house — which is almost anyone not burning laboratory budget — the new model is more expensive per completed task. Against its own predecessor, it also goes up.

None of which makes the spend wrong. It means "the price dropped" and "I will spend less" are two different sentences, and the second does not follow from the first.

There is also a spending lever nobody discusses, because it does not appear in announcements: the effort level. It is a parameter that decides how much the model thinks before answering, and it moves output token consumption across a wide range. Anthropic's documentation says to start at the default and adjust in either direction based on your own evals.

If someone is quoting you "AI-assisted work" without telling you what effort level it runs at, you are being quoted a cost range, not a cost.

The anchor nobody can prove

A reading that circulated among well-known creators in the space belongs on the table, labeled as what it is — a commercial hypothesis, not a verifiable fact: that shipping the flagship months earlier at a very high price may have worked as an anchor. Set a high reference, and the model that lands under it later feels like a discount.

We have no way to test that, and the people who floated it did not present it as certainty. It earns its place here because the technique is old and you have watched it in your own category: the top-of-the-line item almost nobody buys, which exists mostly so the one in the middle looks sensible.

Where do we stand when we say all this?

Plainly, because it is the only thing that gives this text any value: at Kynoz we use these tools every day to build our clients' software, and we pay those token bills ourselves. We write and review code with AI assistants, we run agents on long tasks, and the invoices come out of company money. We are talking from the garage, not from the grandstand.

That commits us to two things. One is a disclosure: none of the companies named here pays us, we are not partners or resellers of any of them, and we earn nothing if you end up using one over another. The other is less comfortable: because we use these tools on work that gets delivered and invoiced, we cannot afford to believe an announcement. A model that promises to review code and in practice lets bugs through does not cost us a headline. It costs us the call from a client whose system stopped working.

Which is why this piece will not tell you which model to use. We don't know your operation, your volume, or what happens in your business when something breaks — and anyone recommending a model without those three things is selling, not advising.

You buy software. What do you do with all this?

Nothing dramatic. Move four questions earlier in the conversation.

When your vendor announces they have upgraded to the latest model, this is what separates an informed decision from buying a headline:

  1. Who measured that number? If the company that trained the model published it, that is legitimate product information, measured on its own equipment. If it came from an independent evaluator, ask which harness they ran it in.
  2. How close is that test to what my company actually does? A high score on public-repository problems says nothing about your sales-tax filings in the states where you have nexus, the EDI feed your largest customer requires, or the paperwork that keeps a shipment USMCA-eligible.
  3. Who reviews the output before it reaches my operation? This is the expensive question. The evidence of this week — a company in the business of code review measuring fewer findings and far more noise than its own baseline — does not make the model bad. It does settle whether the human review step is optional. CodeRabbit's own conclusion is that the output has to be filtered before it reaches developers.
  4. What changes on my invoice? If the answer includes "it's cheaper," ask cheaper than what. You have just seen that the comparison depends entirely on the starting point.

And a fifth that is really for you, not for the vendor: what happens in your business the day one of these tools gets it wrong? If the error is caught within the hour and fixed before anyone outside notices, go ahead. If the error travels to a customer, to your books, or to a filing, you already know which part of the process needs a person standing in it.

Does any of this change if the team is nearshore?

Two things, and both cut in your favor if you use them.

The first is timing. A vendor half a world away answers questions like these overnight, in writing, after the decision has already been made. A team inside your own working hours answers them on the call where the question comes up. Ask which effort level their agents run at, and who reads the output before it reaches your branch, while everyone is still awake. Shared business hours get sold as a feature of working across the USMCA region; they are worth precisely what you use them for, and this is one of the things to use them for.

The second is currency. Token bills are denominated in dollars no matter where the team sits, so a lower hourly rate does not absorb a model choice that multiplies the volume somebody has to filter, and it does not absorb the rework when a missed bug reaches your customer. In a cross-border contract the clause that matters is not the rate. It is who signs for the result when the output is wrong, and that answer should be a name, not an uptime percentage.

And if you're the one building the agents?

A short section for readers with a technical team, because some of the changes matter and did not make the coverage.

The most useful one is free: the official documentation asks you to remove inherited instructions like "include a final verification step" or "use a subagent to verify," because the new model verifies its own work and those lines cause over-verification — tokens burned for nothing. If your team is carrying prompts written for earlier models, there is money sitting on the floor.

Three more, also from the documentation. Tools can now be added or removed mid-conversation while the prompt cache survives (in beta), which matters if you run long agents where rebuilding the cache is a line item. The fallbacks parameter has a new default mode that applies recommended fallback models by refusal category instead of a list you maintain by hand. And the minimum cacheable prompt length came down, so prompts that were previously too short to cache now qualify.

Two warnings from the same page, because migrating blind is expensive: reasoning is on by default — revisit your output token limit, which now has to cover the thinking as well — and disabling it at the top effort levels returns an error. That last one breaks compatibility with whatever you already have written.

The only things we can state as fact

That a new model shipped, that it undercuts its own house's flagship on price, and that two serious evaluators with visible methodology found it level with the model it supposedly beat: one in software engineering, the other on its overall intelligence index.

That the industry's most-cited coding test still has no neutral measurement of it.

And that the useful question was never which model is ahead this week. It is how far the laboratory sits from your road. Your company does not operate on the dyno, and a vendor who shows you the number without telling you the conditions is showing you the window sticker, not the keys.

Before you approve a change in price, scope or timeline justified by this week's announcement, ask for the four answers above in writing. If you want another set of eyes on what comes back, that is what the contact form and WhatsApp are for. We don't sell AI models — we build software and straighten out processes, and for that it helps if you understand what you are being charged for.

Frequently asked questions

Does Claude Opus 5 beat Claude Fable 5 at coding?

Depends who is measuring. Anthropic states that its new model is state of the art on coding evaluations. Epoch AI, which runs them independently, has the two level on software engineering and puts Fable 5 ahead on the general capabilities index. Artificial Analysis has them in a technical tie on general intelligence, with a clear Opus 5 advantage on long agentic tasks. The defensible reading of what is published as of July 26, 2026 is that they tie on coding and the new one is cheaper.

Why don't Anthropic's numbers match the independent evaluators'?

Because a model is evaluated inside a harness — the program that gives it tools, lets it retry and decides when it stops — and the harness changes the result. A June 2026 analysis attributes swings of 10 to 20 points on the same test, with the same model, to that variable alone. On the leaderboard that analysis reviewed, nearly every score had been reported by the company that trained the model. Two figures from "the same test" are not necessarily comparable to each other.

Is it true that it now costs half?

Against the flagship of the same house, yes on the per-token rate (see the pricing table above). But real cost is measured per completed task, not per token, and there the new model averages more than the mid-tier model of the same family and more than its own predecessor, according to Artificial Analysis. Coming from the flagship, you save. Coming from something cheaper, you spend more. Fast mode also bills above the base rate.

Can I let AI review code without a human looking at it?

On this week's evidence, that is not a defensible promise. CodeRabbit — a company whose business is reviewing code — measured the new model at its maximum effort setting detecting fewer known bugs than the baseline it already had in production, and generating several times more low-value comments (exact figures in the results table above). It works as a precision-oriented second reviewer, and somebody still has to filter what it produces before it reaches developers.

What do I ask my software vendor when they tell me they're on the latest model?

Four things: who measured the number they are quoting, how close that test is to your company's actual work, who reviews the output before it touches your operation, and against what starting point the promised saving is measured. If the team works in your time zone, ask all four live and get the answers in one call instead of four email threads.

Does Kynoz use these tools?

Yes, daily, to build our clients' software, and we pay those bills ourselves. Which is why this article does not recommend a model: we don't know your operation or what an error costs you. What we do is read the original sources before repeating a headline, and that is exactly what a client is paying for.