DeepSeek V4 Flash 0731 vs Claude Opus 4.8: The 1 Cent Agent Run

DeepSeek shipped V4-Flash-0731 on July 31 with the same architecture and 7.4x the DeepSWE score. We priced the benchmark table nobody prices: one cached agent run costs $0.0096 against $0.73 on Claude Opus 4.8.

DeepSeek V4 Flash 0731 vs Claude Opus 4.8: The 1 Cent Agent Run

TL;DR: Is DeepSeek V4 Flash 0731 Worth Switching To?

On 2026-07-31 DeepSeek shipped V4-Flash-0731 with the same architecture and the same parameter count as the preview, and its published agent scores land between 78% and 98% of Claude Opus 4.8 depending on which of the seven public benchmarks you read. The median is 92%. The interesting number is not the score, it is the price of a cached agent run: about $0.0096 on V4 Flash against $0.7324 on Opus 4.8, a 76x gap. That gap is twice as wide as the 38x you get from list prices alone, because DeepSeek charges 2% of its input price for a cache hit while Anthropic charges 10% and adds a cache-write premium. Agent workloads are the most cache-heavy shape there is, so the discount compounds exactly where agents live.

Two caveats travel with that number and are covered below. DeepSeek ran its own scores on an unreleased harness, and DeepSeek has announced peak-hour pricing at 2x with no effective date.

What Actually Shipped on July 31

The upgrade is narrower than the benchmark table suggests, and DeepSeek said so in its own follow-up note.

Fact Detail
Release date 2026-07-31, public beta
Model architecture Unchanged from V4-Flash-Preview
Parameter count Unchanged from V4-Flash-Preview
Scope of upgrade The deepseek-v4-flash API only
V4 Pro API Unchanged
App and Web models Unchanged
New format support Responses API, natively
Harness adaptation Codex
Context length 1M tokens
Max output 384K tokens

The architecture line is the one worth sitting with. Same shape, same size, and the DeepSWE score moved from 7.3 to 54.4. Whatever produced that came from post-training and harness work rather than scale, which is a different cost curve for DeepSeek and a different upgrade cadence for anyone building on it.

Quotable finding: DeepSeek-V4-Flash-0731 kept the exact architecture and parameter count of its preview and still moved DeepSWE from 7.3 to 54.4, a 7.4x gain. Agent capability in 2026 is being bought with post-training, not parameters.

How Much Did the Same Model Improve?

Every row below is the same weights class, preview against 0731.

Benchmark V4-Flash-Preview V4-Flash-0731 Gain
DeepSWE 7.3 54.4 7.45x
AutomationBench (Public) 10.8 25.1 2.32x
Cybergym 38.7 76.7 1.98x
Agents' Last Exam 15.8 25.2 1.59x
Toolathlon-Verified 49.7 70.3 1.41x
NL2Repo 39.4 54.2 1.38x
Terminal Bench 2.1 61.8 82.7 1.34x

0731 also beats DeepSeek's own V4-Pro-Preview on every published row, including Terminal Bench 2.1 at 82.7 against 72.1 and DeepSWE at 54.4 against 12.8. The cheap tier now outscores the expensive tier's preview, which is a reason to re-run your Flash-versus-Pro decision rather than inherit it from the migration.

How Close to Claude Opus 4.8 Is It Really?

Here is the honest version, sorted by how favorable the row is. The percentage is V4-Flash-0731 as a share of Opus 4.8 on the same benchmark.

Benchmark V4-Flash-0731 Opus 4.8 Share of Opus
Agents' Last Exam 25.2 25.7 98.1%
Terminal Bench 2.1 82.7 85.0 97.3%
DeepSWE 54.4 58.0 93.8%
Cybergym 76.7 83.1 92.3%
Toolathlon-Verified 70.3 76.2 92.3%
AutomationBench (Public) 25.1 27.2 92.3%
NL2Repo 54.2 69.7 77.8%

Range: 78% to 98%. Median: 92%.

Pick Terminal Bench and you can write that V4 Flash is at 97% of Opus. Pick NL2Repo and it is at 78%. Both sentences are true and neither is the story. Five of the seven public rows cluster between 92% and 98%, and NL2Repo is the clear outlier, which reads as a specific weakness in translating natural language into whole-repository changes rather than a general capability gap.

DeepSeek published two further rows, DSBench-FullStack (68.7 against 71.6) and DSBench-Hard (59.6 against 71.7). Its own footnote says both are internal benchmark sets, so they are excluded from the range above rather than blended into it. A vendor's private eval is evidence about the vendor's priorities, not a comparable score.

The Footnote Most Write-Ups Will Skip

DeepSeek's benchmark image carries a footnote that changes how much weight the table can hold:

For public Code Agent tasks, V4-Flash was tested using our upcoming DeepSeek Harness (minimal mode) framework. Settings: max tier, topp=0.95, temperature=1.0.

Three things follow. The harness is DeepSeek's own. It is described as upcoming, so it was not publicly available when the scores were published. And there is no statement that GLM-5.2 and Opus 4.8 were run through the same harness at the same settings.

Agent benchmark scores are a property of the model and the scaffold around it. A better harness raises scores without touching weights, which is the same mechanism that plausibly explains part of the preview-to-0731 jump. None of this makes the numbers wrong. It makes them vendor-reported under conditions you cannot yet reproduce.

Quotable rule: A vendor-run agent benchmark on a private harness measures the model and the harness together. Until the harness ships, treat the score as a claim about the pair, not the model.

If the decision matters, run your own 20-task set. Our guide to how benchmark scores are produced has the method.

First-Party Prices, Including the Cache Column

Every figure here comes from the vendor's own published pricing page on 2026-08-01, per 1M tokens.

Model Input Cache read Output Cache read as share of input
DeepSeek V4 Flash $0.14 $0.0028 $0.28 2.0%
DeepSeek V4 Pro $0.435 $0.003625 $0.87 0.8%
GLM-5.2 $1.40 $0.26 $4.40 18.6%
Claude Opus 4.8 $5.00 $0.50 $25.00 10.0%

The cache column is where this gets interesting, and it is the column nobody puts in the comparison table.

Anthropic prices a cache read at 0.1x base input, and charges 1.25x base input to write a 5-minute cache entry. DeepSeek prices a V4 Flash cache read at 0.02x its input, five times deeper in relative terms and 179x cheaper in absolute terms, and its published pricing shows no cache-write premium. V4 Pro's cache read is deeper still at 0.8% of input.

That asymmetry does nothing for a one-shot chat call. It does a great deal for an agent.

The Artifact: What One Agent Run Actually Costs

An agent run replays its own history. Every step resends the system prompt, the tool schemas, and every prior tool result, which is why the cache-hit rate on a well-built agent loop is far higher than on chat traffic. Here is a deliberately ordinary coding-agent run, priced four ways.

Assumptions, stated so you can change them:

Parameter Value
Steps per run 20
Fixed prefix (system prompt + tool schemas) 8,000 tokens
Appended per step (tool result + assistant message) 1,500 tokens
Output per step 600 tokens
Total input processed across the run 445,000 tokens
Total output 12,000 tokens
Tokens served from cache 408,500 (91.8%)
Tokens billed as fresh input 36,500

The 91.8% cache rate is not optimism. It falls straight out of the arithmetic: only the 8,000-token prefix and the 1,500 new tokens per step are ever fresh.

Two pricing assumptions sit underneath the table. Anthropic is billed with its published 5-minute cache-write rate of $6.25 per 1M, because it charges separately to write. DeepSeek is billed with fresh tokens at its published cache-miss rate, which its pricing page states as an explicit column, so no write premium applies. For GLM-5.2 we found no published cache-write rate and have billed fresh tokens at base input; if Z.ai does charge to write, its row is slightly optimistic, which does not affect the DeepSeek-to-Opus comparison.

Cost per run:

Model Cached run Uncached run Saving from cache
DeepSeek V4 Flash $0.0096 $0.0657 85.4%
DeepSeek V4 Pro $0.0278 $0.2040 86.4%
GLM-5.2 $0.2101 $0.6758 68.9%
Claude Opus 4.8 $0.7324 $2.5250 71.0%

Worked example for the two ends, so the numbers are checkable:

DeepSeek V4 Flash, cached
  cache reads   408,500 x $0.0028 / 1M = $0.001144
  fresh input    36,500 x $0.14   / 1M = $0.005110
  output         12,000 x $0.28   / 1M = $0.003360
  total                                = $0.009614

Claude Opus 4.8, cached (5-minute cache writes)
  cache writes   36,500 x $6.25   / 1M = $0.228125
  cache reads   408,500 x $0.50   / 1M = $0.204250
  output         12,000 x $25.00  / 1M = $0.300000
  total                                = $0.732375

The finding: caching does not narrow this gap, it widens it.

Comparison Ratio
Opus 4.8 to V4 Flash, list prices only 38x
Opus 4.8 to V4 Flash, both models caching 76x
GLM-5.2 to V4 Flash, both models caching 22x

Both models were given the same cache-friendly workload and both were credited with their own vendor's cache discount. The gap still doubled, because a 2%-of-input cache read beats a 10%-of-input cache read that also carries a write premium. Anyone comparing agent models on the input and output columns alone is reading the wrong two columns.

At 10,000 agent runs a month, the same table reads: $96 on V4 Flash, $278 on V4 Pro, $2,101 on GLM-5.2, $7,324 on Opus 4.8.

One number we deliberately did not adjust for

Anthropic's pricing documentation notes that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text." Read the baseline before you use it. The same page says Claude Sonnet 4.6 and earlier use the previous tokenizer, so that 30% is measured against Anthropic's own older tokenizer, not against DeepSeek's. It is a real number if you are costing a move from Opus 4.6 to Opus 4.8. It says nothing about how many tokens the two vendors count for the same prompt.

Token prices are quoted per token, but your prompt is text, and no published figure compares DeepSeek's tokens-per-text to Anthropic's. So the 76x above is left unadjusted in both directions. If that difference matters to your budget, it is cheap to settle: send the same 20 prompts to both APIs and read usage.prompt_tokens off each response.

Against GLM-5.2, the Other Cheap Agent Model

Z.ai's GLM-5.2 has been the default recommendation for cost-conscious agent work, including in our own Kimi K3 versus GLM-5.2 comparison. V4-Flash-0731 beats it on all six comparable public rows.

Benchmark V4-Flash-0731 GLM-5.2 Delta
AutomationBench (Public) 25.1 12.9 +12.2
Toolathlon-Verified 70.3 59.9 +10.4
DeepSWE 54.4 46.2 +8.2
NL2Repo 54.2 48.9 +5.3
Terminal Bench 2.1 82.7 81.0 +1.7
Agents' Last Exam 25.2 23.8 +1.4

Two of those margins are inside noise for a single run, and the harness caveat applies to this pairing too. The cost side is not close: 22x on the cached agent run above. A team that moved to GLM-5.2 for agent cost reasons in June has a reason to re-measure.

The Risk on Every Number Above

DeepSeek's pricing page carries a forward-looking notice: the API "will soon adopt a peak/off-peak pricing policy. During peak hours, prices will be 2x the regular prices, applicable to all billing items." Peak hours are listed as 9:00 to 12:00 and 14:00 to 18:00 Beijing time, daily, with the effective date pending an official announcement.

At 2x on all billing items, the cached run goes to roughly $0.0192 and the Opus ratio halves to 38x. That is still a wide gap, and it is a real scheduling variable: six hours a day of doubled pricing rewards moving batch and non-interactive agent work outside the window. Anyone budgeting on the numbers in this post should check whether that date has been announced.

Two other things to hold loosely. The model is in public beta. And 0731 is a dated build, so the behavior you test today is the behavior of a specific snapshot.

Which Model ID, and Where

The direct route is DeepSeek's own API on the OpenAI-compatible shape, with the explicit model ID rather than a retired alias:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_DEEPSEEK_API_KEY",
    base_url="https://api.deepseek.com",
)

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "List the files you would read first."}],
)

If you are running an agent harness, the 0731 build also speaks the Responses API natively and is adapted for Codex, which removes a translation layer that previously sat between DeepSeek and several agent frameworks.

On OpenModels, DeepSeek V4 models are reachable through one OpenAI-compatible endpoint and one key alongside the rest of the catalog:

client = OpenAI(
    api_key=os.environ["OM_API_KEY"],
    base_url="https://api.getopenmodels.com/v1",
)

Marketplace prices, directional only

Our live listing currently shows DeepSeek V4 Flash from $0.124 input and $0.249 output per 1M across 4 provider routes. Treat that as directional. A marketplace "from" price is a starting figure composed across routes, the lowest input and the lowest output can come from different providers, and route availability moves. Our own July migration post recorded 6 routes at $0.124 and $0.247, which is how quickly this surface changes. None of the arithmetic in this post uses these figures, and neither should your budget until you have confirmed a specific route.

Frequently Asked Questions

What changed in DeepSeek V4 Flash 0731?

DeepSeek upgraded the deepseek-v4-flash API on 2026-07-31 with substantially stronger agent and coding performance, native Responses API support, and Codex adaptation. The model architecture and parameter count are unchanged from the preview build, so the gains come from post-training and harness work. The V4 Pro API and the App and Web models were not changed in this release.

Is DeepSeek V4 Flash as good as Claude Opus 4.8?

Not quite, and the answer depends on the benchmark. Across the seven public agent benchmarks DeepSeek published, V4-Flash-0731 scores between 78% and 98% of Opus 4.8, with a median of 92%. It comes closest on Agents' Last Exam (98%) and Terminal Bench 2.1 (97%), and is furthest behind on NL2Repo (78%), which tests translating natural language into repository-wide changes. All scores are vendor-reported on DeepSeek's own unreleased harness.

How much does an agent run cost on DeepSeek V4 Flash?

About one cent. On a 20-step coding-agent run processing 445,000 input tokens and generating 12,000 output tokens with a 91.8% cache-hit rate, V4 Flash costs roughly $0.0096 at first-party list prices. The same run costs about $0.0278 on V4 Pro, $0.21 on GLM-5.2, and $0.73 on Claude Opus 4.8. Your figures will differ with prompt size, step count, and cache-hit rate.

Why does prompt caching favor DeepSeek more than Anthropic?

Because the discounts are different sizes. DeepSeek charges $0.0028 per 1M cached input tokens against $0.14 uncached, so a cache read costs 2% of the input price. Anthropic charges $0.50 against $5.00, so a cache read costs 10%, and it adds a 1.25x premium to write a 5-minute cache entry. On a cache-heavy agent workload the DeepSeek-to-Opus cost ratio widens from 38x on list prices to 76x once both models cache.

Should I use DeepSeek V4 Flash or V4 Pro for agents?

Start by re-testing Flash. V4-Flash-0731 outscores the V4-Pro-Preview on every published benchmark row, so a Flash-versus-Pro decision made during the July migration may now be inverted. Pro remains roughly 2.9x the cost of Flash per cached agent run, which is a much smaller gap than the list prices suggest because Pro's cache read is only 0.8% of its input price.

Will DeepSeek prices go up?

DeepSeek has announced a peak and off-peak pricing policy under which peak-hour prices will be 2x regular prices across all billing items, with peak hours listed as 9:00 to 12:00 and 14:00 to 18:00 Beijing time. No effective date has been announced. If it lands as described, the cached agent run above moves from about $0.0096 to about $0.0192 during those six hours, and scheduling non-interactive work outside the window becomes a real lever.

The Bottom Line

The headline that will circulate is that a model priced at $0.14 per 1M input tokens got within a couple of points of Claude Opus 4.8 on Terminal Bench. The more useful reading is that it landed between 78% and 98% depending on the task, on a harness DeepSeek has not released, and that the cost gap on realistic agent traffic is roughly 76x rather than the 38x the price sheet implies.

That does not make it the right model for every job. NL2Repo is the row that should give a repository-scale agent team pause, and vendor-run scores are not a substitute for your own eval. It does mean the cost of being wrong is now very low. At one cent a run, testing V4-Flash-0731 against your existing agent on 20 real tasks costs less than the meeting about whether to test it.

Route it through OpenModels if you want DeepSeek, GLM-5.2, and a fallback behind one key and one OpenAI-compatible endpoint. Attribute the spend per agent run before you scale it, which is the job an agent finance gateway does through cost attribution and per-agent margin.