Changelog
Every launch, price move, deprecation and retirement we track — newest first.
Executed. DeepSeek's price rise, announced with a date and a table on 2026-08-14, takes effect at 16:00 UTC today, and both V4 rows have been re-pointed from the old flat rate to the new PEAK column. deepseek-v4-flash moves from $0.0028 cache-hit / $0.14 cache-miss / $0.28 output per 1M to $0.014 / $0.44 / $1.32; deepseek-v4-pro moves from $0.003625 / $0.435 / $0.87 to $0.044 / $1.32 / $3.96. That is 4.71x and 4.55x on output respectively, and 12.1x on pro's cache-hit input — the largest single multiple in the table. Two things a consumer of this feed should know, because the schema has one price field and DeepSeek now publishes two. First, we carry PEAK because DeepSeek presents peak as the price and off-peak as half of it, not the reverse — so the fields are the ceiling, and the OFF-PEAK rate is exactly half on every axis. Second, peak is the minority of the day: 01:00-04:00 and 06:00-10:00 UTC, 7 hours of 24, so a workload that does not care when it runs pays the half rate for 17 hours a day. The legacy deepseek-chat and deepseek-reasoner aliases are NOT affected and keep their old figures: they were retired 2026-07-24 15:59 UTC, three weeks before this schedule began, so it never applied to them.
source ↗DeepSeek's pricing page has replaced the undated warning it carried since at least 2026-08-01 — 'We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected' — with a dated schedule. From 16:00 UTC on August 16, 2026 the API moves to peak/off-peak billing, with off-peak rates stated as half the peak rates; peak hours are 01:00-04:00 and 06:00-10:00 UTC, so 7 of every 24 hours are peak and the rest are off-peak. Per 1M tokens, deepseek-v4-flash goes from $0.0028 cache-hit / $0.14 cache-miss / $0.28 output to OFF-PEAK $0.007 / $0.22 / $0.66 and PEAK $0.014 / $0.44 / $1.32 — 4.71x on peak output and 2.36x even off-peak. deepseek-v4-pro goes from $0.003625 / $0.435 / $0.87 to OFF-PEAK $0.022 / $0.66 / $1.98 and PEAK $0.044 / $1.32 / $3.96 — 4.55x on peak output, and 12.1x on peak cache-hit input, the largest single multiple in the table. This entry is the ANNOUNCEMENT; the executed re-point is the separate 2026-08-16 entry, where the catalog fields moved to the PEAK column, since DeepSeek presents peak as the price and off-peak as half of it. Caught by scripts/check-pricing-prose.mjs, which diffs a provider’s own sentences rather than its tables — the change was invisible to every rate diff we run, because on 2026-08-14 not one published number had moved yet.
source ↗Google's release notes for August 13, 2026 make Gemini 3.7 Flash (gemini-3.7-flash) generally available, calling it 'our most intelligent workhorse model yet for coding and agents' with substantial improvements across software engineering, web development and agentic workflows. Specs: 1,048,576-token context, 65,536 max output, text/image/video/audio/PDF in and text out, thinking at low/medium/high ('minimal' returns an error), batch, flex and priority inference all supported. Google publishes no knowledge cutoff, and the deprecations table announces no shutdown date. The pricing is the part worth reading carefully: paid-tier Standard is $0.75 in / $3.75 out per 1M with context caching at $0.075/1M — but only 'through December 31, 2026'. From January 1, 2027 the same page states $1.50 / $7.50 / $0.15, exactly double on all three axes, and Google's own release note calls the current rate an introductory price. That makes 3.7 Flash identical in price to gemini-3.6-flash today and identical again after the cutover. Anyone budgeting off today's number is budgeting off a rate with a published expiry date; our fields carry the in-force column and a dated ticket moves them on the boundary.
source ↗Anthropic's deprecations table now states Current state = Retired for claude-opus-4-1-20250805, with Deprecated June 5, 2026 and Retirement August 5, 2026 — exactly the date announced two months earlier. Calls to claude-opus-4-1-20250805 (and the claude-opus-4-1 alias) no longer serve; Anthropic's recommended replacement is claude-opus-4-8. Opus 4.1 was $15 in / $75 out per 1M tokens, 200K context, 32K max output. Caught by scripts/check-lifecycle-drift.mjs on 2026-08-06 — the provider's declared status flipped while our catalog still read 'deprecated'. Worth noting for anyone building on rate data: this event is invisible to a price diff, because the price never changed — the endpoint simply stopped existing.
source ↗Google's deprecations table now names gemini-3.6-flash — not gemini-3.5-flash — as the recommended replacement for gemini-2.5-flash (shutdown 2026-10-16), gemini-2.5-flash-preview-05-20, gemini-2.5-flash-preview-09-25, and the whole gemini-2.0-flash line (gemini-2.0-flash and gemini-2.0-flash-001, both shut down 2026-06-01). The new target is CHEAPER on output than the old one: Gemini 3.6 Flash is $0.75 in / $3.75 out per 1M (context caching $0.075) against Gemini 3.5 Flash's $1.50 / $9.00, so the two Flash rows are not a straight ladder. Gemini 3.6 Flash is GA/stable, 1,048,576-token context, 65,536 max output, text+image+video+audio+PDF input; Google publishes no knowledge cutoff for it. CORRECTED 2026-08-15: this entry originally quoted Gemini 3.6 Flash at $1.50 / $7.50 / $0.15 and said the two models shared an input price. Those are Google's 'starting January 1, 2027' figures — the page prices this model on a two-date schedule and the rates above are the ones in force through 2026-12-31, so 3.6 Flash is currently cheaper than 3.5 Flash on BOTH axes and the gap closes on 2027-01-01. models.json was corrected on 2026-08-14; this entry, which had copied the same wrong column, was missed for a day. The release date this entry said Google did not publish is in the deprecations table: July 21, 2026.
source ↗Google released two new embodied-reasoning endpoints for robotics in public preview: gemini-robotics-er-2-preview (advanced spatial reasoning, agentic code execution, multi-step tool orchestration, video moment finding, progress classification, multi-robot coordination) and gemini-robotics-er-2-streaming-preview (real-time text streaming over the Live API for low-latency robot agents with bidirectional audio and video input). Both accept text, image, video and audio input, output text, and carry 131,072-token input / 65,536-token output limits. Paid-tier standard rates are $2.00/M in and $10.00/M out for both — 2x the input and 2x the output of Gemini Robotics-ER 1.6 ($1.00/$5.00). Only the non-streaming endpoint publishes a context-caching rate ($0.20/1M) and a batch tier ($1.00/$5.00). Google's own upgrade guidance is to replace model="gemini-robotics-er-1.6-preview" with either ER 2 endpoint; ER 1.6 is shut down 2026-08-31.
source ↗Google shipped stable, production-ready Gemini 3.6 Flash and Gemini 3.5 Flash-Lite on the same day. Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) is billed at $0.30/M in and $2.50/M out on the paid tier, context caching $0.03/M, batch $0.15/$1.25 — against gemini-3.1-flash-lite's $0.25/$1.50, i.e. the recommended migration target is more expensive on both sides, unusually for a Flash-Lite step. One offsetting simplification: 3.5 Flash-Lite publishes a SINGLE input rate covering text, image, video and audio, where 3.1 and 2.5 Flash-Lite both charged a higher audio-input tier. 1,048,576-token context, 65,536 max output, text/image/video/audio/PDF in, text out; computer use in preview. Google's deprecations table names it the recommended replacement for gemini-3.1-flash-lite, which shuts down 2027-05-07. No shutdown date announced for 3.5 Flash-Lite itself.
source ↗OpenAI's GPT-5.6 generation went GA on the API + Codex on 2026-07-09 (preview from 06-26), in three durable capability tiers: Sol (flagship) $5 in / $30 out, Terra (balanced) $2.50 / $15, Luna (fast, high-volume) $1 / $6 — all per 1M tokens, cached input $0.50 / $0.25 / $0.10. Every tier has a 1.05M-token context window, 128K max output, text + image input, and a 2026-02-16 knowledge cutoff. Sol matches GPT-5.5's price but is stronger on coding, knowledge work, cybersecurity and science. New caching model: explicit cache breakpoints, 30-min minimum cache life, cache writes billed 1.25x uncached input.
source ↗xAI's smartest model for chat, coding, agentic and knowledge work. Base pricing $2 input / $6 output per 1M tokens (cached input $0.50), tiered 2x above a 200K-token prompt. 500K context window (down from Grok 4.3's 1M), text + image input. Public API access opened 2026-07-09; aliases grok-4.5-latest / grok-build-latest.
source ↗Anthropic's most agentic Sonnet, with performance close to Opus 4.8 at lower cost. Standard pricing $3 input / $15 output per 1M tokens, with introductory pricing of $2/$10 in effect through 2026-08-31. 1M context window, 128k max output, Jan 2026 knowledge cutoff. Supersedes Claude Sonnet 4.6, now a legacy model.
source ↗Mistral released Leanstral 1.5, a specialist model for Lean 4 formal proof engineering, automated theorem proving and autoformalization (not a general chat model). Sparse MoE 119B total / 6.5B active, 256K context, 128K max output, text + image input. Open weights under Apache 2.0. Offered FREE in Mistral's experimental Labs tier ($0/M in and out), with a stated retirement of 2026-09-30 — the catalog's first $0 model, recorded as preview (Labs/experimental) so it stays out of the GA 'cheapest' rankings.
source ↗Tracking 181 models across 12 providers.
Mistral deprecated its frontier reasoning model (magistral-medium-2509, $2/$5 per MTok, 128k context). Migrate to Mistral Medium 3.5 — Mistral is folding its specialist reasoning line back into its general-purpose models rather than shipping a Magistral successor.
source ↗Mistral deprecated its frontier agentic-coding model (devstral-2512, $0.4/$2 per MTok, 256k context, open weights under a Modified MIT licence). Migrate to Mistral Medium 3.5 — the specialist coding line is being folded back into the general-purpose models.
source ↗Mistral deprecated its small reasoning model (magistral-small-2509, $0.5/$1.5 per MTok, 128k context, Apache-2.0 weights, 24B). Migrate to Mistral Small 4.
source ↗Video generation (Videos API). Priced per second, not per token: sora-2 $0.10/s (720p, standard) / $0.05/s batch. Videos API + sora-2* announced for discontinua
source ↗Higher-quality video model. Per-second pricing: $0.30/s (720p) up to $0.70/s (1080p) standard; batch half. Discontinued with Videos API, shuts down 2026-09-24,
source ↗Mistral deprecated its small agentic-coding model (labs-devstral-small-2512, $0.1/$0.3 per MTok, 256k context, Apache-2.0 weights, 24B). Migrate to Mistral Medium 3.5. Its stated 2026-03-31 retirement has passed, but Mistral still lists the model as deprecated rather than retired and still prices it publicly.
source ↗Get the next change before your code breaks
New models, price cuts, and deprecations — a short email when something actually changes. No spam, unsubscribe anytime.
◎ You're on the watch list. We'll ping you the moment a model launches, changes price, or gets deprecated.
✕ Couldn't subscribe right now — please try again in a moment.
Free forever · powered by the same data on this page.