DeepSeek-V4-Flash
PreviewSmaller/cheaper V4 model (~284B total / ~13B active params per authoritative third-party reports). Context length 1M, max output 384K tokens. Input price $0.44/M cache-miss, $0.014/M cache-hit; output $1.32/M (USD). Supports dual modes (Thinking / Non-Thinking), JSON output, tool calls, chat prefix completion; FIM completion is non-thinking-mode only. Concurrency limit 2500. The legacy aliases deepseek-chat and deepseek-reasoner currently route to this model (non-thinking / thinking respectively). Part of the 'DeepSeek V4 Preview' generation (released 2026-04-24), hence status=preview. Knowledge cutoff NOT officially published by DeepSeek -> left null. Prices re-verified to the cent against api-docs.deepseek.com/quick_start/pricing/ on 2026-08-09. VENDOR-DECLARED FORWARD PRICE EVENT, RESOLVED 2026-08-14 — the same footnote now carries a date and a table. Until 2026-08-13 it read only 'We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected... subject to official notice', with no date and no figure. DeepSeek now states: 'DeepSeek API pricing will be updated to peak / off-peak billing, with off-peak rates at half the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC (all other hours are off-peak). The new prices take effect at 16:00 UTC on August 16, 2026'. Announced rates for this model per 1M tokens: OFF-PEAK $0.007 cache-hit / $0.22 cache-miss / $0.66 out; PEAK $0.014 cache-hit / $0.44 cache-miss / $1.32 out. EXECUTED 2026-08-16: the priced fields above were re-pointed at the PEAK column on the cutover day, because DeepSeek presents peak as the rate and off-peak as half of it, not the other way round. The OFF-PEAK figure is exactly half on every axis and is what you actually pay for 17 of every 24 hours, since peak is only 01:00-04:00 and 06:00-10:00 UTC. Re-read first-hand from api-docs.deepseek.com/quick_start/pricing/ on 2026-08-16, when the page still stated the cutover as forthcoming: the fields were moved ~8 hours early, in the run of the cutover day, because this catalog is refreshed once daily at ~08:05 UTC and the alternative was to publish the superseded rates for ~16 hours after they stopped being true. THAT WINDOW IS NOW CLOSED: re-read first-hand 2026-08-17, the page presents the peak/off-peak table as the rates in force and the forward-dated announcement banner is gone; every figure above still matches it to the cent, so the fields are simply current from here on. Responses API supported (footnote (1)). The pricing page is served at the TRAILING-SLASH url (2026-08-11): the extensionless form returns a different document ("Your First API Call") with HTTP 200.
DeepSeek-V4-Flash by DeepSeek costs $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66/1M blended), with a 1M (1.000.000-token) context window. It is in preview. Its weights are publicly downloadable, so it can be self-hosted on your own infrastructure.
Last verified: 19 Aug 2026 · sourced from official provider documentation
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · no spam · unsubscribe anytime.
Compare DeepSeek-V4-Flash head-to-head
$0.66 vs $10 blended /M · cross-provider
$0.66 vs $11.25 blended /M · cross-provider
$0.66 vs $1.50 blended /M · cross-provider
$0.66 vs $3 blended /M · cross-provider
$0.66 vs $0.75 blended /M · cross-provider
$0.66 vs $1.40 blended /M · cross-provider
DeepSeek-V4-Flash — questions & answers
How much does DeepSeek-V4-Flash cost?
DeepSeek-V4-Flash is priced at $0.44 per 1M input tokens and $1.32 per 1M output tokens ($0.66/1M blended at a 3:1 input-to-output ratio), with cached input at $0.014 per 1M tokens.
What is the context window of DeepSeek-V4-Flash?
DeepSeek-V4-Flash has a 1.000.000-token context window (1M), with up to 384.000 output tokens per request.
Is DeepSeek-V4-Flash deprecated?
No — DeepSeek-V4-Flash is in preview and not currently scheduled for deprecation or retirement in our tracker.
Does DeepSeek-V4-Flash support image or vision input?
Our catalog lists DeepSeek-V4-Flash's modalities as text; image/vision input is not among them.
Can I self-host DeepSeek-V4-Flash?
Yes — DeepSeek-V4-Flash is an open-weight model: its weights are publicly downloadable, so you can run it on your own infrastructure (subject to its license) instead of only calling DeepSeek's API.
What is the API model string for DeepSeek-V4-Flash?
The API model identifier for DeepSeek-V4-Flash is "deepseek-v4-flash" when calling DeepSeek's API.
Track DeepSeek-V4-Flash price & status changes
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
The last few changes we caught — this is what lands in your inbox:
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.