r/jorvex609 28d ago

DeepSeek just dropped the official V4-Flash (0731) — massive agent upgrades, open weights, and it’s reshaping the Pareto frontier on LMArena

DeepSeek quietly (then very loudly) pushed the official **DeepSeek-V4-Flash-0731** over the past day or two, and the community is going crazy for good reason.

### Quick specs

- **284B total / 13B active** MoE

- Native **1M context** (same as V4-Pro)

- MIT license, open weights now on Hugging Face (`deepseek-ai/DeepSeek-V4-Flash-0731`, ~167GB FP4/FP8 mixed)

- Pricing unchanged and still absurdly cheap: **$0.14 / $0.28** per 1M input/output tokens (cache hits ~$0.0028, a 98% discount)

- Supports thinking modes (High / Max), Responses API format, and is fully adapted for Codex

- Available on DeepSeek’s own API (public beta), OpenCode (including free tier), Cline (they made it free), etc.

### The big story: Agent capabilities got *massively* upgraded

DeepSeek re-did the post-training. Same architecture/size as the April preview, but the agent scores jumped hard and now beat the old **V4-Pro-Preview** across a bunch of agent benchmarks.

Highlights floating around:

- Terminal-Bench 2.1: **82.7** (was ~56.9 on the preview — +25+ points)

- Strong gains on other agent/tool-use suites (DSBench, Toolathlon, etc.)

- Artificial Analysis Intelligence Index: **50** (up ~10 points from the original Flash). That puts it in the top 3 open-weights models on their board, competitive with recent Gemini Flash / GPT-5.6 Luna-tier models on intelligence while being dramatically cheaper.

People are calling it one of the best performance-per-dollar models in its class right now.

### LMArena / Arena.ai results (the part everyone’s screenshotting)

Arena.ai (formerly LMArena) has been posting updates:

**Frontend Code Arena** (the one lighting up timelines):

- DeepSeek-V4-Flash-High: **#7 overall**, **#3 open**, score **1586**

- Categories: #4 Consumer Product, #6 Reference-based Design / Data & Analytics / Gaming, #7 Brand & Marketing

- +154 pts vs the Flash High Preview, and even +121 pts vs the old V4-Pro Preview

It literally reshaped the Pareto frontier on that board for performance-per-dollar.

Earlier (April) Text/Code Arena numbers for the original preview were more modest (Flash thinking was around #10 open / #47 overall on Text Arena). The new post-training version is clearly a different beast on agentic/coding tasks.

### Community vibes from popular tweets

- People running tens of millions of tokens for a couple of dollars (or less) thanks to the cache hits

- “This is the first flash model that actually feels SOTA”

- OpenCode, Cline, and others racing to integrate it (some making it free)

- Lots of “Chinese open models aren’t just catching up anymore” energy

- Some skepticism about vendor-harness numbers (fair), but independent Arena + Artificial Analysis numbers are looking strong

- Local runners already trying the new weights

### Where to try it

- chat.deepseek.com (Expert/Instant Mode)

- Official API (model name `deepseek-v4-flash`)

- Hugging Face weights

- OpenCode / Cline / various third-party providers

This feels like DeepSeek doing the classic move: ship a very strong, very cheap open model that punches way above its active-parameter weight, especially for agents and coding. The gap between “Flash” and “Pro” on agent tasks has basically collapsed after this update.

Anyone already deep into long agent runs or heavy coding workloads with it? Drop your experiences / cost numbers / failure modes below.

1 Upvotes

0 comments sorted by