284B total, 13B active
A sparse mixture of experts: only about 13B parameters fire per token, which is why it answers fast and costs so little.
DeepSeek V4 Flash is the cheap half of the V4 family: about 284B total parameters with only 13B active per token, a 1 million token context window and hybrid attention built for long agent runs. The 0731 build, released on July 31, 2026, kept the same body and redid the post-training, and DeepSeek reports it now beats its own V4-Pro preview across nine agent benchmarks. On avots.ai you run it pay per use, with no DeepSeek account and no subscription, and the same balance also covers Claude Opus 5, GPT-5.6, Gemini and 40+ image, video and music models.
A sparse mixture of experts: only about 13B parameters fire per token, which is why it answers fast and costs so little.
A full million tokens in one conversation, with very large outputs, so whole repositories and document sets fit in a single pass.
Per million tokens on DeepSeek's own list, with cached input at 0.0028 USD, so repeated context in agent loops is almost free.
DeepSeek reported for the 0731 build, alongside 70.3 on Toolathlon verified and 76.7 on Cybergym.
On the independent Artificial Analysis Intelligence Index, against a median of 25 for comparable models.
No image input on this one. For screenshots and scans switch to Qwen3.7 or Gemini on the same balance.
| DeepSeek V4 Flash | DeepSeek V4-Pro | Qwen3.7 Flash | |
|---|---|---|---|
| Price per 1M in / out | $0.14 / $0.28 | $0.435 / $0.87 | $0.03 / $0.13 |
| Parameters | 284B total, 13B active | about 1.6T total, 49B active | not published |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Terminal-Bench 2.1 | 82.7 | not published | not published |
| Toolathlon verified | 70.3 | not published | not published |
| Image input | no | no | yes |
| On avots.ai | pay per use | pay per use | pay per use |
Agent scores are vendor reported and very sensitive to the harness, so treat them as a direction, not a verdict. What is not in doubt is the price gap: Flash runs at roughly a third of V4-Pro per token, and Qwen3.7 Flash undercuts both while adding image input. All three sit in the same picker.
Identical architecture and size to the preview build. Everything gained came from redone post-training, not a new base model.
DeepSeek reports 54.4 on DeepSWE, 54.2 on NL2Repo and 25.2 on Agents' Last Exam, its strongest agent set yet.
Cache hit input at 0.0028 USD per million tokens makes tool loops that resend the same prompt genuinely cheap.
High and maximum reasoning effort are supported, so you spend more thinking only on the calls that need it.
Sign up on app.avots.ai with email, Google or Telegram. Takes a minute.
1,000 tokens cost 0.99 EUR. No plan to choose, no renewal, the balance sits there until you use it.
Choose DeepSeek V4 Flash in the model picker, or let Agent Avots route for you. Every reply shows its token cost.
app.avots.ai on desktop and mobile, no install needed.
V4 Flash in @AvotsAIbot, same balance, from any phone.
Use it inside Claude Desktop, Cursor or Cline via the MCP server.
The OpenAI compatible API takes your existing client, one key for every model.
1,000 tokens cost 0.99 EUR and a Flash answer costs a handful of them. The same balance unlocks Claude Opus 5, GPT-5.6, Gemini, Veo video and Nano Banana images, with no subscription anywhere.
Sign up and chat ✦The full family, V4-Pro and V4-Flash, open weights and 1M context.
Every LLM on avots.ai and which one to pick per task.
Alibaba's Max, Plus and Flash tiers, 1M context, vision on the small ones.
Moonshot's frontier open model, number one at coding, 1M context.
Anthropic's newest flagship, 1M context, 96% SWE-bench Verified.
What to use instead of ChatGPT Plus and how one balance replaces it.
avots.ai runs 40+ AI models for chat, image, video and audio on one balance. See the full list on the Tools page or open the app and start creating.