DeepSeek V4.1 Flash Weights Go Live: 552B, 1M Context, MIT

Abhishek GautamAbhishek Gautam15 min read
DeepSeek V4.1 Flash Weights Go Live: 552B, 1M Context, MIT

Quick summary

Open weights on ModelScope, live on Alibaba Qwen, 890 bytes of KV cache per token. Up to 42x cheaper than GPT-6 Astra on output.

Advertisement

A 552-billion-parameter model that stores only 890 bytes of memory per token of context is now free to download under an MIT licence. DeepSeek V4.1 Flash has been on DeepSeek's API since September, and on October 6, 2026 its weights landed on Alibaba's ModelScope hub while the model went live on Alibaba's Qwen platform. It reads text and images, handles a 1,048,576-token context window, and costs $0.30 per million input tokens and $1.20 per million output tokens at peak on DeepSeek's own API.

Independent evaluator Vals AI ranks it the number one open-weight model on its index. It also found DeepSeek's headline coding score was inflated by its own harness. Here is what V4.1 Flash is, what the numbers really say, and when it makes sense to switch.

What DeepSeek V4.1 Flash Is

DeepSeek V4.1 Flash is an open-weight multimodal Mixture-of-Experts language model from Hangzhou-based DeepSeek, positioned as its "lightweight flagship": cheaper and faster than V4-Pro while matching or beating it on agentic coding.

SpecValue
Total parameters552B (plus about 196B reported "Engram" parameters)
Active parameters8B per input token (prefill), 16B per output token (decode)
Architecture40-layer Causal Encoder-Decoder (20 encoder + 20 decoder)
Experts1 shared + 384 routed, top-6 routing
Context window1,048,576 tokens
Max outputAbout 384K tokens (393,216 on the Qwen platform)
KV cacheAbout 890 bytes per token (FP4 keys and values, Compressed Sparse Attention 2, SWA Bounded Replay)
ModalitiesText and image input (DeepSeek-ViT trained jointly), text output
LicenceMIT, commercial use allowed
API namedeepseek-flash

DeepSeek says the KV cache is about a quarter of V4-Flash and 1/437th of its first-generation model. That single number is the most important thing about this release.

Why 890 Bytes per Token Matters

KV cache is the memory a model keeps for every token already in its context. On long-context workloads, KV cache, not model weights, is often what runs a GPU out of memory.

Our math: 1,048,576 tokens x 890 bytes = about 0.93 GB of KV cache for a full one-million-token context. Many dense models of similar quality need tens of gigabytes per long sequence. That means:

  • More concurrent long sessions per GPU, which cuts serving cost for RAG and agents
  • Cheap cache hits: DeepSeek charges just $0.006 per million cached input tokens at peak
  • Long-horizon agents become affordable, because rereading a huge context costs almost nothing

Pricing: How It Compares

ModelInput (per 1M)Output (per 1M)Cached input
DeepSeek V4.1 Flash (peak)$0.30$1.20$0.006
DeepSeek V4.1 Flash (off-peak)$0.15$0.60$0.003
DeepSeek V4-Pro-0813 (peak)$1.32$3.96$0.044
GPT-6.1 Sol (one-fifth of Astra list price)About $2About $10Varies
GPT-6 Astra$10$50Varies

DeepSeek peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday; everything else, including weekends and Chinese public holidays, is off-peak. On Alibaba's Qwen platform the same model costs CNY 1 to 2 per million input tokens and CNY 4 to 8 per million output tokens, with limits of 15,000 requests and 1 million tokens per minute.

At peak, V4.1 Flash is about 33x cheaper than GPT-6 Astra on input and 42x cheaper on output, and roughly 7x to 8x cheaper than GPT-6.1 Sol.

Benchmarks: Official vs Independent

DeepSeek's model card reports, at maximum reasoning effort:

BenchmarkDeepSeek reported
Terminal-Bench 2.190.6
DeepSWE v1.1 (resolved)74.2
Terminal-Bench 3.030.0
Terminal-Bench 4.031.2
NL2Repo-Bench64.0
GPQA Diamond90.9
Humanity's Last Exam36.8
Codeforces rating3471
SWE-bench Verified (API changelog, base config)66.0

Now the independent view. Vals AI scored V4.1 Flash at 57.86% on its Vals Index (September 10 update): 15th of 56 models overall and first among open-weight models, just ahead of Kimi K3 at 57.81%. Vals says the run cost about $0.30 per test, against $6.47 for Kimi K3.

But Vals measured Terminal-Bench 2.1 at 74.53% across three full trials, well below DeepSeek's 90.6. DeepSeek ran its headline numbers with its own "DeepSeek Harness Minimal" scaffold, and its own table shows how much the harness matters:

Agent harnessDeepSWE v1.1Terminal-Bench 2.1
mini-SWE74.290.3
DeepSeek Harness Minimal72.690.6
Claude Code69.888.0
Pi66.286.1
Codex65.684.1
OpenCode65.585.0

Our Analysis: Strong Model, Inflated Headline

1. The cost-performance story is real. Even at Vals' lower numbers, V4.1 Flash is the best open-weight model it has tested, at a small fraction of frontier prices. For high-volume workloads such as document parsing, support bots, code search and RAG, it is the new default to benchmark against.

2. The harness gap is a warning. A 16-point difference between self-reported and independent Terminal-Bench scores is large. If you use Codex or OpenCode as your agent scaffold, expect DeepSWE results closer to 65% than 74%. Always test on your own harness.

3. It is not a frontier model. Fifteenth on Vals means more than a dozen closed models still score higher. For the hardest reasoning, long autonomous coding runs or high-stakes outputs, frontier models still lead. See our GPT-6 Astra vs Fable 5.1 vs Gemini 3.8 Flash comparison.

4. The China angle. One day after V4.1 Flash weights went live, DeepSeek open-sourced DeepGEMM-Ascend and DeepEP-Ascend, ports of its matrix-multiplication and expert-communication libraries for Huawei Ascend 950 chips (requiring CANN 9.20). That is the pattern: release a model that runs efficiently, then make sure it runs on domestic silicon. Our coverage of Huawei Ascend 950PR and DeepSeek's Ascend inference cluster explains why this matters for export controls.

Self-Hosting: What You Need

With 552B total parameters, the weights alone are roughly 552 GB at 8-bit and about 276 GB at 4-bit precision, before KV cache and activations. Rough guide:

SetupFits?Notes
8x H200 (141 GB each, 1,128 GB total)Yes at 8-bitRoom for many long-context sessions thanks to the small KV cache
8x H100 (80 GB each, 640 GB total)Tight at 8-bit, comfortable at 4-bitQuantisation quality needs testing
4x H2004-bit onlyFewer concurrent sessions
Single consumer GPUNoUse the API or a smaller model

Only 8B to 16B parameters are active per token, so throughput is closer to a mid-size dense model once the weights are loaded. Check your inference engine's release notes for support of the new attention and FP4 KV cache before planning a deployment.

When to Switch

WorkloadRecommendation
RAG over long documentsStrong switch candidate; cheap cache hits and 1M context
Customer support and classificationSwitch candidate; test tone and refusal behaviour
Code search and repository Q&ATest against your current model; harness matters
Autonomous multi-hour coding agentsKeep a frontier model as primary; use V4.1 Flash for subtasks
Regulated data in the US or EUSelf-host the MIT weights rather than calling a China-hosted API

Compare current prices across providers on our LLM API Pricing Tracker, and see earlier coverage of DeepSeek V4's launch and the Chinese open-weights price war.

Key Takeaways

  • DeepSeek V4.1 Flash: 552B MoE, 8B / 16B active, 1M context, text and image input, MIT licence
  • Weights on ModelScope and live on Alibaba Qwen as of Oct 6, 2026; on DeepSeek's API (deepseek-flash) since September
  • Pricing: $0.30 input / $1.20 output per million tokens at peak, half off-peak, $0.006 cached
  • About 890 bytes of KV cache per token: a full 1M-token context needs under 1 GB
  • Vals AI: #1 open-weight model, 15th of 56 overall, about $0.30 per test
  • Caveat: Vals measured Terminal-Bench 2.1 at 74.5% vs DeepSeek's 90.6%
  • Huawei Ascend ports of DeepSeek kernels followed a day later

Sources

  • DeepSeek-V4.1-Flash model card on Hugging Face and ModelScope (October 2026)
  • Alibaba Cloud Qwen platform listing and pricing (Oct 6, 2026)
  • Vals AI index results (Sept 10, 2026 update), via LaunchBoosts
  • DeepSeek API pricing documentation and LLMCost price tracking (Oct 9, 2026)
  • TechRadar on DeepGEMM-Ascend and DeepEP-Ascend (Oct 7, 2026)

FAQ

Frequently Asked Questions

What is DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is an open-weight multimodal Mixture-of-Experts model with 552 billion total parameters, 8 billion active for input and 16 billion for output, a 1,048,576-token context window and an MIT licence. It is served as deepseek-flash on the DeepSeek API.

How much does DeepSeek V4.1 Flash cost?

On the DeepSeek API it costs $0.30 per million input tokens and $1.20 per million output tokens at peak, half that off-peak, and $0.006 per million cached input tokens. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

Is DeepSeek V4.1 Flash better than GPT-6 Astra?

No, not overall. Vals AI ranks it 15th of 56 models and first among open-weight models. It is far cheaper, about 42 times cheaper than GPT-6 Astra on output tokens, which makes it strong for high-volume tasks, but frontier closed models still lead on the hardest reasoning and long autonomous coding.

Are the DeepSeek V4.1 Flash benchmarks accurate?

DeepSeek reports 90.6 on Terminal-Bench 2.1 using its own harness, but Vals AI measured 74.53 percent across three trials. DeepSeek own table also shows scores vary by agent harness, from 65.5 to 74.2 on DeepSWE v1.1, so developers should test on their own setup.

Can I self-host DeepSeek V4.1 Flash?

Yes, the weights are MIT licensed. The 552B model needs roughly 552 GB at 8-bit or 276 GB at 4-bit for weights alone, so an 8x H200 or 8x H100 server is a realistic minimum. Its small KV cache of about 890 bytes per token allows many long-context sessions per server.

Advertisement

Free Weekly Briefing

The AI & Dev Briefing

One honest email a week — what actually matters in AI and software engineering. No noise, no sponsored content. Read by developers across 30+ countries.

No spam. Unsubscribe anytime.

Free Tool

Will AI replace your job?

4 questions. Get a personalised developer risk score based on your stack, role, and what you actually build day to day.

Check Your AI Risk Score →

Written by

Software Engineer based in Delhi, India. Writes about AI models, semiconductor supply chains, and tech geopolitics — covering the intersection of infrastructure and global events. 1054+ posts cited by ChatGPT, Perplexity, and Gemini. Read in 167 countries.