Rate Card / 2026.09.24

Generative AI API Rate Card

Unit prices, specs, batch discounts and prompt-caching rules for the major text, image, video, voice and music APIs. Click a column header to sort. Price movements are recorded in the change log at the bottom.

Researched 2026-09-24 Previous 2026-09-14 FX 1 USD = 159.33 JPY (default, 2026-08-28) Listed 13 providers / 95 items

Latest moves — on a single day, 22 September, OpenAI and Anthropic both cut prices. ① GPT-6 Sol (9/22) is $2/$10, exactly half of GPT-5.6 Sol's current price ($4/$20), and a permanent price, not a promotion. GPT-6 Luna, released alongside it, is $0.10/$0.50. ② Claude Opus 5.5 (9/22) is $4/$20 (Opus 5 was $5/$25). Cache reads also fell from $0.50 to $0.20. Anthropic reports that it beats Fable 5.1, which costs 2.5x as much, on coding benchmarks. ③ Grok 4.7 (9/21) stays at $2/$6, the same as 4.6, on a larger base model. ④ Gemini 3.8 Live (9/15) is out. Audio in/out is $0.005 / $0.018 per minute, far cheaper than GPT-Live-1 ($0.05/min), though the two are not directly comparable (see the Voice section). ⑤ The Sora 2 API ends today, 9/24. Earlier versions of this page said Sora ended on 3/24. 3/24 was the announcement date; the app closed on 4/26 and the API closes on 9/24. This has been corrected. ⑥ Alibaba announced Qwen 4 Max (9/22), but only as an announcement, with no price or release date, so it is not in the tables.

The top tier has split into a "$10/$50 premium band" and a "$2–$4 practical band." GPT-6 Astra and Fable 5.1 sit together at $10/$50, while GPT-6 Sol, Opus 5.5, Grok 4.7 and Sonnet 5 crowd into the $2–$4 range. Early September was about the top end getting pricier; late September was about mid-tier price cuts. Anthropic has also said Sonnet 5.5 and Haiku 5.5 will arrive "in the coming weeks," so the mid and light tiers are likely to move again.

DeepSeek's peak hours overlap the Asian business day. Peak is weekdays (Mon–Fri) 01:00–04:00 and 06:00–10:00 UTC, which is 10:00–13:00 and 15:00–19:00 in Japan (Chinese public holidays excluded). Outside those 7 hours, prices are half. For users in Europe and the Americas, most working hours already fall off-peak; for batch jobs anywhere, scheduling outside the peak windows halves the bill.

Text per 1M tokens / click headers to sort

Output tokens typically cost 4–6x as much as input. The total is driven mostly by output volume, so estimate "how much will it write" before choosing a model. Sorting by Tier lines up each company's equivalent models side by side.

Provider Model Tier Input Output Notes
OpenAIGPT-6 AstraTop$10.00$50.009/3 newNew flagship. 1M context, cached input $1.00, cache write $12.50. Batch/Flex half price, Fast mode double. Sits above GPT-5.6 Sol: 2x on input, 1.7x on output
OpenAIGPT-6 SolMid$2.00$10.009/22 newPermanent priceThe tier below Astra. Overall intelligence on par with GPT-5.6 Sol at half the price (5.6 Sol is $4/$20). Cached input $0.20, cache write $2.50. 1.05M context, 128K output; above 272K, input 2x and output 1.5x. Batch/Flex half, Fast double. Six effort levels, none to max. On DeepSWE and OSWorld, however, GPT-5.6 Sol still has the higher top score, at a higher cost per task. Its output is cheaper than Terra ($2/$12) at the same input price, so there is little reason left to choose Terra
OpenAIGPT-6 LunaLight$0.10$0.509/22 newPermanent priceInput halved and output down 60% from GPT-5.6 Luna ($0.20/$1.20). Cached input $0.01. Among the cheapest input prices of any generally available text model. 1.05M context. In ChatGPT, Free and Go users can use it in the desktop app
AnthropicClaude Fable 5.1Top$10.00$50.009/1 newBase price unchanged. Cache reads cut 75%, from $1.00 to $0.25 (0.025x input; other Claude models are 0.1x, so this one is an exception). Cache write $12.50 (5 min) / $20 (1 hr). Batch $5/$25
AnthropicClaude Mythos 5.1Top$10.00$50.00RestrictedRestricted-access version of Fable 5.1, via a vetting program
AnthropicClaude Fable 5Top (prev.)$10.00$50.00One generation before 5.1. Cache reads still $1.00
Sakana AIFugu Ultra v2.0Top$5.00 / $10.00$30.00 / $45.009/11 new★A Japanese company (Sakana AI). Right-hand figures apply above 272K tokens. Not a single model but an "orchestrator" that routes work to other companies' models and recombines the answers. For hard, many-step problems. 1M context
Sakana AIFugu MaxMid$2.00$6.009/11 newSame architecture as Ultra, aimed at results per dollar. Sakana says it undercuts Sonnet 5 and Kimi K3 on output price by 40–60%
OpenAIGPT-5.6 Sol / GPT-5.5Upper (prev.)$5.00 / promo $4.00$30.00 / promo $20.001M context. No deprecation announced. Promo price ($4/$20) runs at least until 11/21. GPT-6 Sol matches it at half the price, so there is little reason to pick it for new work
AnthropicClaude Opus 5.5Top$4.00$20.009/22 newFirst of the 5.5 generation. 20% cheaper than Opus 5; cache read $0.50→$0.20, cache write $6.25→$5. Batch $2/$10. 1M context, 128K output (300K in Batch beta), model ID claude-opus-5-5. Anthropic reports it beats Fable 5.1 ($10/$50) on Terminal-Bench 4.0, 66.4 vs 55.8, while trailing GPT-6 Astra on some benchmarks. The "about 40% cheaper" claim compares medium effort with high effort; the model change alone is 20%. MigrationThinking cannot be turned off / forced tool use errors / the old computer-use tool is rejected, so some Opus 5 code will not run unchanged. Sonnet 5.5 and Haiku 5.5 announced for "the coming weeks"
AnthropicClaude Opus 5Top (prev.)$5.00$25.001M context, five-level effort dial. Same price for five generations since Opus 4.5. Inference speed improved on 8/12
MoonshotKimi K3Top$3.00 / $0.30$15.00Right-hand figure is on cache hit. 1M context. No Batch support
GoogleGemini 3.1 ProTop$2.00 / $4.00$12.00 / $18.00Right-hand figures apply above 200K tokens. 3.5 Pro has been delayed three times and is unreleased, so this is the current Pro
xAIGrok 4.7Top$2.00 / $4.00$6.00 / $12.009/21 newSame price as 4.6. A larger base model with longer reinforcement learning aimed at hours-long tasks. Right-hand figures apply above 200K. Cached input $0.50 ($1 above 200K), 500K context, four reasoning levels (low to xhigh), model ID grok-4.7. The double-price Fast variant is only available via Cursor and Grok Build, not the public API. xAI's own CursorBench 4.0 score is 46.3, short of Fable 5.1 (51.8). OpenRouter listed it at $1.60/$4.80 on 9/22, about 20% below list
xAIGrok 4.6Top (prev.)$2.00 / $4.00$6.00 / $12.008/12Same price as 4.5. 500K context, cached input $0.50, knowledge to 2026/2. Four reasoning depths (low/medium/high/xhigh). Artificial Analysis index 61, tied with GPT-5.6 Sol and one point behind Fable 5
xAIGrok 4.5Upper$2.00 / $4.00$6.00 / $12.00Opus class. Doubles above 200K. Cached input $0.50. Not eligible for Batch discount
ZhipuGLM-5.2Top$0.40 – $1.40$1.26 – $4.40Conflicting prices744B MoE, 1M context, cached input $0.26. Z.ai first-party sources say $1.40/$4.40; aggregators such as OpenRouter say $0.40/$1.26, a gap of more than 3x. Check the route you actually use
DeepSeekDeepSeek V4 ProTop$0.66 / $1.32$1.98 / $3.968/16 increaseLeft is off-peak, right is peak. Was $0.435/$0.87 flat. Partly reverses May's 75% cut. Still cheap for a top-tier model. The shutdown planned for 9/14 was cancelled, and pricing continues unchanged
OpenAIGPT-5.6 TerraMid$2.00$12.00Cut 20% on 7/30 (from $2.50/$15.00)
AnthropicClaude Sonnet 5Mid$2.00$10.00Increase cancelledWas due to rise to $3/$15 on 9/1, but on 8/10 $2/$10 became the permanent price
GoogleGemini 3.6 FlashMid$1.50$7.50Released 7/21. 3.7 Flash will match this price when its intro pricing ends on 2027/1/1
MetaMuse Spark 1.3Mid~$0.10 blended~$0.10 blended9/2 newSeparate input/output prices not published; only a blended price. A big cut from 1.2 ($1.25/$4.25)
MetaMuse Spark 1.2Mid (prev.)$1.25$4.251M context. An even cheaper tier exists for customers who allow training on their data
xAIGrok 4.3Mid$1.25$2.50Value option for long text. 500K–1M context. 20% off with Batch
GoogleGemini 3.8 FlashMid$0.75$3.759/2 newDoubles 2027/1/1Same price as 3.7 Flash and better on every benchmark. Intro price until 2026/12/31, then $1.50/$7.50. Batch/Flex half, Priority 1.8x. A vetted Cyber variant also exists
GoogleGemini 3.7 FlashMid (prev.)$0.75$3.75Released 8/13. Same price as 3.8, so little reason to choose it
ZhipuGLM-5.3-FlashLight$0.15$0.50Promo ended 9/9Promo: $0.075 / $0.25The 50%-off launch price ended at 24:00 on 2026-09-09 (Singapore time) and list price returned. Intelligence Index 57
ZhipuGLM-5-TurboMid$1.20$4.00Released 2026/3/15. A speed-focused line separate from GLM-5.2
ZhipuGLM-4.6Mid$0.43$1.74Older generation. 200K context
MiniMaxMiniMax-M3Mid$0.30$1.20After the "Permanent 50% off." Input up to 512K, cache read $0.12
AnthropicClaude Haiku 4.5Light$1.00$5.00On the expensive side for the light tier
GoogleGemini 3.5 Flash-LiteLight$0.30$2.50For high throughput and low latency
OpenAIGPT-5.6 LunaLight (prev.)$0.20$1.20Cut 80% on 7/30 (from $1.00/$6.00). GPT-6 Luna ($0.10/$0.50) arrived on 9/22 and took over the lowest input price. No deprecation announced, but little reason to choose it for new work
DeepSeekDeepSeek V4.1-FlashLight$0.15 / $0.30$0.60 / $1.209/10 newLeft is off-peak, right is peak. Model ID deepseek-flash. 552B MoE (8B active for input, 16B for output), image input, weights released under MIT. Cache hits are extremely cheap at $0.003 ($0.006 peak). Requests to the old names deepseek-v4-flash and deepseek-v4-flash-vision-exp are served by this model and billed at its price (the old models themselves are retired)
DeepSeekDeepSeek V4 FlashLight$0.22 / $0.44$0.66 / $1.328/16 increaseLeft is off-peak, right is peak. Was $0.14/$0.28 flat. Automatic caching brings input down to $0.007 ($0.014 peak)
DeepSeekDeepSeek V4 Flash VisionLight / exp.$0.22 / $0.44$0.66 / $1.328/21Experimental. Image input, same price as V4 Flash. 1M context, 384,000-token output limit
DeepSeekDeepSeek R1Reasoning$0.55$2.19Input $0.14 on cache hit
No matches

Images per image

For text inside images, especially non-Latin scripts such as Japanese, choose GPT Image 2 or Nano Banana Pro; both are established as rarely garbling characters. If you just need volume, Nano Banana 2 costs one third of Pro.

Provider Model Tier Price Notes
OpenAIGPT Image 2.5
(Flare / Sunburst)
Top$0.25 1K
$0.30 2K
$0.60 4K
9/8 newTwo variants: Flare for fast volume work, Sunburst for precise editing. Both have the same price. Official pricing is per token (text input $5.00 / image input $8.00 / image output $30.00 per 1M tokens; cached $1.25 and $2.00); the per-image figures on the left are reseller street prices. Supports both generation and editing of reference images
OpenAIGPT Image 2Top$0.006 low
$0.053 med
$0.211 high
Batch 50%At 1024×1024. High quality at 3840×2160 is $0.401. Landscape/portrait are slightly cheaper at medium–high ($0.041/$0.165). Accurate Japanese text
GoogleNano Banana Pro
(Gemini 3 Pro Image)
Top$0.039 std
$0.134 2K
$0.24 4K
Batch 50%Japanese text rarely breaks. Strong at realistic people and text layering. Image generation does not support caching
ByteDanceSeedream 5.0 ProMid$0.045 ≤2.36MP
$0.09 above
8/12Generation and editing, multi-reference fusion. First reference image free, $0.003 per extra. On BytePlus ModelArk and also fal.ai, Replicate, etc. A Lite version exists
GoogleNano Banana 2Mid$0.067 1KPro-level quality at Flash-level speed. 3–5 seconds per image, native 4K. One third the price of Nano Banana Pro
xAIgrok-imagine-image-qualityMid$0.05 1K / $0.07 2KBatch 50%
xAIgrok-imagine-imageLight$0.02Batch 50%1K. Can go down to $0.01
MiniMaximage-01Light$0.0035Cheapest. When volume matters more than quality
OpenAIGPT Image 1Retiring$0.02 – $0.19Retires 2026/10/23Avoid for new projects
No matches

Video per second

Even within one company, prices differ by more than 10x across tiers. Sort by Tier to line up equivalent models across companies, instead of seeing Veo Lite / Fast / Standard scattered. Turnover is fast; model names from six months ago are often gone.

Provider Model Tier Resolution Price Notes
ByteDanceSeedance 2.5High1080p (reseller)~$0.577/31, API 8/75 seconds at 1080p is $2.843 on BytePlus. Native output of the official API is 480p/720p only; resellers' 1080p and 4K are upscaled. Up to 30 seconds with audio in one pass, references up to 30 images / 10 videos / 10 audio clips, timestamp-level editing
GoogleVeo 3.1 StandardHighwith audio$0.40Some sources report $0.75/s via Vertex AI; prices vary widely by route
AlibabaWan 3.0High1080p$0.208/24$6 for 30 seconds at 1080p. Failed generations are not charged
MiniMaxH3 (Hailuo 3.0)High2K (1440p) 24fps$0.137/31Omni-modal. 5–15 seconds with stereo audio generated together. Accepts up to 9 images, 3 videos and 3 audio clips as references. There is no 1080p tier. $7.80 per minute
ByteDanceSeedance 2.0Highofficial~$0.14#1 qualityBilled at CNY 46 per 1M tokens. 15 seconds ≈ 308,880 tokens. Video input (editing) is CNY 28 per 1M
Luma AIRay 3 / Ray 3.2High1080p~$0.21 – $0.2480 credits/second. About $12.60 for 60 seconds
KuaishouKling 3.0Standard720p / 1080p$0.075 – $0.144K, 60fps, 15 seconds, multilingual lip-sync. Official pricing uses credits (6–8/s without audio, 9–12/s with audio). Turbo version added in mid-June
RunwayGen-4.5Standardstandard$0.1212 credits/second. About $1.20 for 10 seconds
MiniMaxH3 (third-party)Standard2K / 768p$0.124 / $0.077AIHubMix and others. Slightly cheaper than official. Segmind is more expensive at $0.1625 for 2K
AlibabaWan 3.0Standard720p$0.108/24★Omni-Reference — besides text, images, audio and video, it accepts documents, spreadsheets, slides, PDFs and web pages as input and builds a video from them. 2–30 seconds; t2v, i2v and reference generation. Public beta 8/6, GA 8/24. Used in production for short dramas, ads, tourism promotion and music videos
AlibabaWan 2.7Standardtext-to-video$0.10Older generation. Available on Together AI and others
GoogleVeo 3.1 FastStandardstandard$0.10 – $0.15
ByteDanceSeedance 2.0 FastStandard1080p$0.09Rated "the cheapest at production quality"
MiniMaxH3 (Hailuo 3.0)Standard768p$0.08 – $0.09Effectively unusableOn the price list, but invitation-only closed beta. The API only accepts 2K as a resolution, and you are billed at the 2K price even if you specify 768p
ByteDanceSeedance 2.0Lightthird-party$0.04 – $0.067OpenRouter $0.067, PoYo from $0.04. Official API access is limited, so most real use goes through these
Luma AIRay 3 / Ray 3.2Light720p~$0.0620 credits/second (at Plus plan rates)
xAIGrok Imagine VideoLightstandard$0.05Batch 50%$0.025/s with Batch. The only video model with an explicitly stated Batch discount
AlibabaWan 3.0Light480p$0.058/24Alibaba Cloud Model Studio. CNY 0.30/s in mainland China (Beijing), CNY 0.37471/s in Singapore
GoogleVeo 3.1 LiteLight720p no audio$0.03 – $0.08Among the cheapest video options. First choice if you do not need audio
GoogleGemini Omni / Omni FlashUnknown—Not publishedAwaiting priceAnnounced at I/O on 5/19; Omni Flash released 6/30. As of 8/25 the official pricing page lists only Veo 3.1 under video, with no pay-as-you-go price
OpenAISora 2Ended—DiscontinuedAPI ends 9/24App and web closed on 4/26. The API closes on 2026-09-24 (3/24 was the announcement to developers). OpenAI has not named a successor
No matches

Voice / TTS

Real-time conversation (you talk, it talks back)

This area mixes "per minute" and "per token" pricing, so the numbers cannot be compared as they are. On top of that, GPT-Live-1 is priced for the voice layer only; the thinking is handed to a separate model and billed separately. Estimating from the listed price alone will leave you short.

Provider Model Price Notes
OpenAIGPT-Live-1$0.05 / minAPI since 9/10Its selling point is listening while talking (full duplex). Billed per second, and silence and background thinking time are billed at the same rate. Extra chargesThis covers the voice layer only; reasoning is delegated to a separate model such as GPT-6 Astra or GPT-5.6 Luna. A 10-minute call costs about $0.51 with Luna, about $0.86 with Astra. Endpoint v1/live/sessions, 128K context, knowledge to 2025/7. 12 voices, but the published list covers English accents plus Brazilian Portuguese and Filipino; Japanese is not listed. There is no mini in the API (mini is for ChatGPT's free plan)
OpenAIgpt-realtime-2.1audio in $32 / out $64 per 1M tokensStill sold at the same price after GPT-Live-1. Cached input $0.40. Measured at roughly $0.05/min. Input accumulates as the conversation grows, so beyond 10–15 minutes the effective per-minute cost rises
OpenAIgpt-realtime-2.1-miniaudio in $10 / out $20 per 1M tokensMeasured at roughly $0.016/min, close to a third for the same use
GoogleGemini 3.8 Liveaudio in $0.005 / min
audio out $0.018 / min
9/15 newA native speech-to-speech model (no transcription step). Two model IDs: gemini-3.8-live and gemini-3.8-live-extended-thinking for multi-step reasoning. Reported to support 97 languages, including Japanese (secondary source; not confirmed in Google's own docs). About 1.18 s to first audio, near-real-time video input, SynthID watermark on all generated audio. UnverifiedPublished prices are for the standard model; the extended-thinking price has not been confirmed. Whether the identical numbers to 3.1 Flash Live are a coincidence or a copy is also unverified
GoogleGemini 3.1 Flash Liveaudio in $0.005 / min
audio out $0.018 / min
Among the cheapest for real-time conversation. For the same "1M audio output tokens," OpenAI charges $64 and Google $12. However, GPT-Live-1's $0.05/min covers only the listening-and-speaking layer, with thinking billed separately, whereas Google's model does the reasoning inside the voice model, so the numbers in this table cannot simply be ranked against each other

Text-to-speech (TTS)

Provider Model Per 1,000 characters Notes
ElevenLabsFlash / Turbo$0.05At 0.5 credits per character, effectively even cheaper
MiniMaxSpeech Turbo$0.06 ($60 / 1M chars)
ElevenLabsMultilingual v2 / v3$0.10High quality
MiniMaxSpeech HD$0.10 ($100 / 1M chars)
OpenAIgpt-4o-mini-tts~$0.015 / minConfirmedBilled per token, not per character (text input $0.60 / audio output $12.00 per 1M tokens), so the per-1,000-character cost varies with the text. Per minute, it is among the cheapest in this table
GoogleCloud TTS (Standard / WaveNet)$0.004Confirmed$4.00 / 1M characters. 4M characters free per month. Voice quality matches the price
GoogleCloud TTS (Neural2)$0.016$16.00 / 1M characters
GoogleCloud TTS (Chirp 3: HD)$0.03$30.00 / 1M characters. Google's high-quality tier
GoogleCloud TTS (Studio)$0.16$160.00 / 1M characters. 40x Standard. Top tier for narration
GoogleGemini 3.1 Flash TTSper tokenText input $1.00 / audio output $20.00 per 1M tokens. Free tier available, half price with Batch. A separate line from the per-character Cloud TTS, so not directly comparable
No matches

Music

In this area many companies offer no official pay-as-you-go API. Suno has no official API and is used via third parties; Udio and Stable Audio are subscription only.

Provider Model Price Notes
Sunov4 family$0.014 – $0.111 / songNo official API; third-party street prices. Direct subscription from Pro $10/month
GoogleLyria 3 Pro$0.05 – $0.15 / songVia MusicAPI, 10 credits per task. The RealTime version streams over WebSocket
MiniMaxMusic 3.0 / 2.6$0.15 / up to 5 minNo new sign-ups since 8/20New users cannot sign up. Existing paying users only. The free tier (music-3.0-free etc.) is gone. New users are directed to the MiniMax Audio app or the open weights on Hugging Face. The lyrics API is the same. Music 3.0 weights were released on 8/13
ElevenLabsMusic$0.80 / min
UdioStandard / Pro$10 – $30 / monthNo pay-as-you-go API
Stability AIStable Audio$5.99 – $89.99 / monthPay-as-you-go API not yet researched
No matches

Batch API and caching discounts

A Batch API lets you submit many requests at once asynchronously and collect results within 24 hours, in exchange for a lower price. It suits bulk jobs that do not need an immediate response. Prompt caching reuses the same long prefix to cut input cost. The two can be combined.

Provider Batch Batch discount Cache Notes
AnthropicYes50% off90% offAll active models. Batch and caching can be combined
OpenAIYes50% offUp to 90%GPT-4o and later. From GPT-5.6, cache writes carry a 1.25x surcharge. o3/o4-mini etc. do not support caching in Batch
GoogleYes50% offImplicit by defaultImplicit caching applies automatically from Gemini 2.5. Image generation does not support caching. Cache hits work inside Batch too
ZhipuYes50% offYesBoth across the GLM lineup
DeepSeekYes~50% offAutomaticBatch is mostly via hosts such as SiliconCloud. Caching is automatic, no opt-out needed
xAISome modelsText 20% / image & video 50%YesText Batch discount applies only to Grok 4.3 and the 4.20 family; Grok 4.5 is excluded. We could not confirm a discount for 4.6 or 4.7. Images and video are a separate category at a flat 50% off. Max 25MB per request
MoonshotOlder models only60% of list priceYesBatch covers K2.7 Code / K2.6 / K2.5. The latest K3 does not support Batch
MiniMaxLimited infoNot disclosedAutomaticCaching is automatic with no setup. Batch is hinted at, but the official docs give no discount rate
MetaTo researchUnconfirmedTo researchThe Model API only entered public preview on 7/9, and discount schemes appear not to be in place yet
No matches

Cache lifetime and when the clock starts

Almost every provider measures idle time since the cache was last used. Each hit resets the timer, so as long as an agent keeps sending requests, even a 5-minute TTL will not expire. The risk is when the gap between requests exceeds the TTL: a single tool run takes too long, the agent waits on a human decision, or it only calls once every few tens of minutes.

Gemini's explicit cache is the one exception. It is a fixed 60 minutes from creation and is not extended by use. For long sessions, set a longer TTL or extend it explicitly via the update API.

Provider Type Default TTL Clock starts Extended on hit Notes
AnthropicPrompt caching5 minLast useResetsCan be extended to 1 hour with ttl:"1h", but writes then cost 2x (normally 1.25x), so you need 3+ reads to break even. The default changed from 1 hour to 5 minutes on 2026/3/6
OpenAIGPT-5.5 and earlier5–10 min idleLast useResetsTwo limits: "cleared after 5–10 minutes idle" and "at most 1 hour after last use." GPT-5 family / 4.1 support extended retention up to 24 hours
OpenAIGPT-5.6 and later30 minLast useResetsKept for at least 30 minutes. Much more relaxed than earlier generations
GoogleImplicitNot disclosedAutomaticAutomaticOn by default from Gemini 2.5. No setup
GoogleExplicit60 minCreationNot extendedThe only case that differs from other providers. Set freely with ttl/expire_time, no upper or lower bound. An API exists to update the TTL after creation
MiniMaxAutomatic caching5 minLast useExtended freeExplicitly stated as "no setup required"
DeepSeekContext cachingNo TTL——Cached on disk; unused entries are deleted after hours to days. You need not think about expiry, but there is also no way to guarantee retention
MoonshotContext cachingManaged——Optimized automatically by the system. There is no TTL parameter at all
ZhipuCachingUnconfirmed——To researchThe cached input price ($0.26/1M) is published, but the expiry rules could not be found
No matches

Change log newest first

A record of "what it used to cost" and "since when this price applies." Prices mostly go down, but introductory prices can be extended or doubled, so dated changes are noted ahead of time.

2027-01-01
Google | Gemini 3.7 Flash intro price ends (scheduled)
$0.75 / $3.75 → $1.50 / $7.50. The price doubles, matching 3.6 Flash.
2026-11-21
OpenAI | GPT-5.6 Sol promo ends (scheduled)
$4.00 / $20.00 → $5.00 / $30.00. Returns to list price.
2026-10-23
OpenAI | GPT Image 1 retired (scheduled)
Avoid for new projects. Move to GPT Image 2.
2026-09-24
OpenAI | Sora 2 API ends — earlier "ended 3/24" was wrong
The Sora 2 API closes today, 2026-09-24. The app and web closed on 4/26, and 3/24 was the date developers were notified. Earlier versions of this page listed 3/24 as the end date; this has been corrected.
OpenAI has not named a successor. Suggested alternatives include Veo 3.1 (with audio), Kling, Hailuo and Seedance. The data-export deadline is described as "conditional," so do not count on a grace period.
2026-09-22
OpenAI | GPT-6 Sol / GPT-6 Luna released — half price, permanently
GPT-6 Sol (GPT-5.6 Sol $4 / $20) → $2.00 / $10.00 (cached input $0.20, cache write $2.50).
GPT-6 Luna (GPT-5.6 Luna $0.20 / $1.20) → $0.10 / $0.50 (cached input $0.01).
Both are permanent prices, not promotions. Model IDs gpt-6-sol and gpt-6-luna; 1.05M context, 128K output; above 272K, input 2x and output 1.5x. Batch/Flex half, Fast double. Six effort levels: none / low / medium / high / xhigh / max.
GPT-6 Astra stays at $10 / $50. The GPT-5.6 family (Sol / Terra / Luna) has no deprecation date and remains on sale.
Positioned as "same intelligence, half the cost." On DeepSWE and OSWorld, GPT-5.6 Sol still has the higher top score. On AutomationBench, GPT-6 Sol (xhigh) scores 33.2% at about $0.27 per task. In ChatGPT it rolls out to Plus / Pro / Business / Enterprise / Education; Free and Go users get Luna in the desktop app.
On the release date: most reports say 9/22, one says 9/23. We use 9/22, following several primary-leaning reports (VentureBeat, MarkTechPost, digitalapplied).
2026-09-22
Anthropic | Claude Opus 5.5 released — 20% cheaper, cache reads down 60%
Opus 5 $5 / $25 → $4.00 / $20.00 (−20%). Cache read $0.50 → $0.20 (−60%), cache write $6.25 → $5 (−20%). Batch $2 / $10.
Model ID claude-opus-5-5, 1M context, 128K output (300K in the Message Batches beta), knowledge to 2026/6. Fast mode (Claude API only, research preview) gives about 2.5x output speed at 2x the price. Also on AWS, Google Cloud and Microsoft.
Anthropic describes it as "about 40% cheaper on typical work," but that figure compares Opus 5 at high effort with Opus 5.5 at medium (the new default); switching models alone is a 20% unit-price cut.
Terminal-Bench 4.0: 66.4 (Fable 5.1: 55.8). FrontierCode v1.1 Main: 54.4 (vs 50.3). It trails GPT-6 Astra on AutomationBench and Terminal-Bench-Science.
Migration notes (where Opus 5 code will error): thinking cannot be disabled (adaptive thinking always on, tuned via effort) / forced tool use errors / thinking blocks are bound to model and conversation state / the old computer-use tool is rejected.
Sonnet 5.5 and Haiku 5.5 are announced for "the coming weeks."
2026-09-22
Alibaba | Qwen 4 announced (Yunqi 2026) — no price or date yet
A four-tier family was previewed on stage: Qwen 4 Max (flagship), Qwen 4 Plus (mid), Qwen 4 Flash (high throughput) and Qwen 4 27B (open weights for local use).
No model card, API identifier, price, release date or scores have been published, and the models are said to be still in training. Only the 27B weights are promised to be released. Not added to the tables. Currently usable are Qwen3.8-Max (8/3, $2/$6) and Qwen3.8 Flash (8/26, $0.14/$0.42).
2026-09-21
xAI (SpaceXAI) | Grok 4.7 released — same price as 4.6
$2 / $6 ($4 / $12 above 200K; cached input $0.50 / $1). Unchanged from 4.6. 500K context, text and image in / text out, model ID grok-4.7, four reasoning levels (low to xhigh).
A new, larger base model with longer reinforcement learning for hours-long tasks. xAI's own measurements: DeepSWE v1.1 71.0%, CursorBench 4.0 46.3% (short of Fable 5.1's 51.8%).
The double-price Fast variant (4.7 Fast) is only available via Cursor and Grok Build, not the public API. Crossing 200K doubles the rate, so long coding sessions can swing the bill considerably.
2026-09-15
Google | Gemini 3.8 Live released — voice conversation gets cheaper again
Two models: gemini-3.8-live (conversation at scale) and gemini-3.8-live-extended-thinking (with multi-step reasoning). Audio in $0.005 / min, audio out $0.018 / min, about 2.3 cents per conversation minute.
Reported to support 97 languages with mid-conversation switching, including Japanese (secondary source). About 1.18 s to first audio, near-real-time visual input, and background tool calls while the conversation continues. SynthID watermark on all output.
Unverified: the extended-thinking price. The numbers are also identical to 3.1 Flash Live, and we have not been able to tell whether this is a generation change with no price change or secondary sources repeating the older model's figures.
2026-09-14
DeepSeek | V4 Pro shutdown cancelled — correcting our earlier entry
The plan announced on 9/10 — "at noon on 9/14 (Beijing time) V4 Pro ends, and requests to deepseek-v4-pro are handled by V4.1-Flash" — was withdrawn before it took effect.
DeepSeek's changelog says that "in response to user demand, the V4 Pro API will continue after 2026-09-14, with billing unchanged." V4 Pro stays at $0.66 / $1.98 (off-peak $0.66/$1.32, peak $1.98/$3.96).
Our 9/13 update reflected the pre-withdrawal announcement. This corrects it.
Something else was retired separately: the old names deepseek-v4-flash and deepseek-v4-flash-vision-exp are discontinued, and requests to them are served by V4.1-Flash at Flash prices. Do not confuse this with V4 Pro.
2026-09-11
Sakana AI | Fugu Max / Fugu Ultra v2.0 released — a Japanese company joins the rate card
Fugu Max $2.00 / $6.00, Fugu Ultra v2.0 $5.00 / $30.00 ($10 / $45 above 272K tokens).
These are not single models but "orchestrators" that route work to other companies' models and recombine the answers. Max aims for results per dollar; Ultra v2.0 aims for correct answers on hard, many-step problems.
Sakana says Fugu Max undercuts Sonnet 5 and Kimi K3 on output price by 40–60%.
2026-09-10
OpenAI | GPT-Live-1 available in the API — listed at $0.05/min, but that is not the whole bill
$0.05 / min (billed per second, no rounding). Released first in the ChatGPT app on 7/8, it reached the API on 9/10.
It is full duplex (listens while speaking), so you can interrupt and the conversation continues. Versus GPT-Realtime-2.1: +30 points on Full Duplex Bench, turn-taking latency 1.41 s → 0.798 s, Tau3 Voice Intelligence task success 45.7% → 86.2%.
Watch the scope of billing. ① $0.05/min covers only the listening-and-speaking layer; the thinking is passed to a separate model such as GPT-6 Astra or GPT-5.6 Luna, whose tokens are billed separately. A 10-minute call is about $0.51 with Luna, about $0.86 with Astra. ② Time when the user is silent and time spent processing in the background are billed at the same rate.
Endpoint v1/live/sessions, 128K context (at 90% it hands off to an 8,192-token summary), instructions up to 16,384 tokens, knowledge to 2025-07-31. 12 voices, but the published list covers English accents (Australian, British, Irish, North American, Southern US) plus Brazilian Portuguese and Filipino; Japanese is not listed.
mini is for ChatGPT's free plan; there is no mini in the API. The existing gpt-realtime-2.1 remains on sale at the same price (audio input $32 / cached $0.40 / output $64 per 1M tokens).
2026-09-10
DeepSeek | V4.1-Flash released
Off-peak $0.15 / $0.60, peak double at $0.30 / $1.20. Cache hit $0.003 ($0.006 peak).
552B MoE, 8B active for input and 16B for output, image input, weights released under MIT.
At launch it was announced that "from 9/14 this model will also handle V4 Pro requests," but that plan was withdrawn before it took effect (see the 9/14 entry).
2026-09-09
Zhipu | GLM-5.3-Flash 50% promo ends — back to list price
$0.075 / $0.25 → $0.15 / $0.50. Ended at 24:00 on 9/9, Singapore time. Z.ai's pricing page also shows list price again.
2026-09-08
OpenAI | GPT Image 2.5 released — Flare and Sunburst
Model IDs gpt-image-2.5-flare and gpt-image-2.5-sunburst. Flare is for fast volume work, Sunburst for precise editing. Both have the same price.
Official pricing is per token: text input $5.00 / image input $8.00 / image output $30.00 (per 1M tokens). Cached: text $1.25, image $2.00.
Per image, reseller street prices are around 1K $0.25 / 2K $0.30 / 4K $0.60. Supports both generation and editing of reference images.
2026-09-03
OpenAI | GPT-6 Astra released — a new top tier
$10.00 / $50.00 (cached input $1.00, cache write $12.50). 1M context, model ID gpt-6-astra. Batch and Flex half price, Fast mode double. Sits above the previous top model GPT-5.6 Sol ($5 / $30): 2x on input, 1.7x on output. Rolled out first through Daybreak, a program for cybersecurity companies.
2026-09-02
Google | Gemini 3.8 Flash released — better, at the same price
$0.75 / $3.75. Same price as 3.7 Flash and better on every published benchmark. Intro price until 2026-12-31, doubling on 2027-01-01 ($1.50/$7.50). Batch and Flex half, Priority 1.8x. A vetted Cyber variant (Fairwind Program) launched the same day.
2026-09-02
Meta | Muse Spark 1.3 released
About $0.10 blended. A big cut from 1.2's $1.25 / $4.25. Separate input/output prices not published.
2026-09-01
Anthropic | Fable 5.1 / Mythos 5.1 — cache reads cut 75%
Base price unchanged at $10 / $50. What moved was caching:
read $1.00 → $0.25 (−75%). Write $12.50 (5 min) / $20 (1 hr). Batch $5/$25.
Fable 5.1's cache read is 0.025x the input price. Other Claude models are 0.1x, so this model is an exception.
Terminal-Bench-Science 52.6 (Fable 5: 24.7). Mythos 5.1 is restricted, via a vetting program.
2026-08-26
Zhipu | GLM-5.3-Flash released
List $0.15 / $0.50. 50% off at $0.075 / $0.25 until 2026-09-09. Intelligence Index 57; about $0.045 per task at the promo price.
2026-08-20
MiniMax | Music and lyrics APIs stop taking new sign-ups
New users can no longer sign up. Existing paying users only. The free API (music-3.0-free and others) is gone. New users are directed to the MiniMax Audio app or the open weights on Hugging Face.
This came right after Music 3.0 was released with open weights on 8/13.
2026-08-24
Alibaba | Wan 3.0 released — video from documents
$0.05 / $0.10 / $0.20 per second (480p / 720p / 1080p). 2–30 seconds. $6 for 30 seconds at 1080p; failed generations are not charged.
The headline feature is Omni-Reference: besides text, images, audio and video, it takes documents, spreadsheets, slides, PDFs and web pages as input and builds a video from them. Since the public beta on 8/6 it has been used in production for short dramas, ads, tourism promotion and music videos.
2026-08-21
DeepSeek | V4 Flash Vision (Experimental) released
Image input. Same price as V4 Flash (off-peak $0.22 / $0.66). 1M context, 384,000-token output limit.
2026-08-16
DeepSeek | Price increase and move to time-of-day pricing
From 16:00 UTC, flat pricing was replaced with peak and off-peak rates.
V4-Pro $0.435 / $0.87 → off-peak $0.66 / $1.98, peak $1.32 / $3.96
V4-Flash $0.14 / $0.28 → off-peak $0.22 / $0.66, peak $0.44 / $1.32
Peak is 01:00–04:00 and 06:00–10:00 UTC (10:00–13:00 and 15:00–19:00 in Japan), 7 hours in total. The remaining 17 hours are off-peak at half price.
This partly reverses May's 75% cut. As a result, the lowest input price moved to GPT-5.6 Luna ($0.20).
2026-08-12
xAI | Grok 4.6 released
$2 / $6 ($4/$12 above 200K, cached input $0.50). Same price as 4.5. 500K context, knowledge to 2026/2. Reasoning depth now selectable in four levels (low / medium / high / xhigh). Artificial Analysis index 61, tied with GPT-5.6 Sol and one point behind Claude Fable 5.
2026-08-13
Google | Gemini 3.7 Flash released
Intro price $0.75 / $3.75 (until 2026-12-31). 1M context, 64,000-token output limit, cached input $0.075. Launched only 23 days after 3.6 Flash.
2026-08-12
Anthropic | Claude Opus 5 update
Improved inference speed and scientific research ability. No price change ($5/$25).
2026-08-12
ByteDance | Seedream 5.0 Pro released
Image generation and editing. $0.045 (up to 2.36MP) / $0.09 (above). First reference image free, $0.003 per extra. Multi-reference fusion. A Lite version also exists.
2026-08-07
ByteDance | Seedance 2.5 developer API opens
The model itself was released on 7/31. Up to 30 seconds with audio in one pass, references up to 30 images / 10 videos / 10 audio clips, timestamp-level editing. 5 seconds at 1080p is $2.843 on BytePlus (~$0.57/s). Native output of the official API is 480p/720p only; resellers' 1080p and 4K are upscaled.
2026-08-10
Anthropic | Claude Sonnet 5 price increase cancelled
A 50% increase to $3 / $15 from 9/1 was withdrawn, and $2 / $10 became the permanent price. Terra, Sonnet 5 and Gemini 3.1 Pro now line up at $2 input in the mid tier.
2026-07-31
MiniMax | H3 (Hailuo 3.0) released
$0.13/second (2K, 1440p, 24fps). Omni-modal, generating audio together. A 768p tier is on the price list but in closed beta.
2026-07-30
OpenAI | Lower two tiers cut
Luna $1.00 / $6.00 → $0.20 / $1.20 (−80%)
Terra $2.50 / $15.00 → $2.00 / $12.00 (−20%)
The flagship Sol was unchanged. This set $0.20 as the floor for mainstream LLM input prices.
2026-07-24
Anthropic | Claude Opus 5 released
$5 / $25. Unchanged for five generations since Opus 4.5. 1M context, five-level effort dial.
2026-07-21
Google | Gemini 3.6 Flash and three other models released
3.6 Flash $1.50 / $7.50 (output down from 3.5 Flash's $9.00, with 17% fewer output tokens); 3.5 Flash-Lite $0.30 / $2.50. 3.5 Pro did not appear this time either.
2026-07-09
Meta | Muse Spark 1.1 released — first paid developer API
$1.25 / $4.25, 1M context, $20 in free credit. Meta Model API opens as a public preview for US developers. Described as "about 25% of OpenAI/Anthropic."
2026-07-09
xAI | Grok 4.5 released
$2 / $6 ($4/$12 above 200K). 500K context. Opus class, but not eligible for Batch discount.
2026-06-30
Google | Gemini Omni Flash released
A world-model type that unifies text, images and video. Pay-as-you-go price still unpublished as of 2026-08-25.
2026-06-16
Zhipu | GLM-5.2 released
$1.40 / $4.40. 744B MoE, 1M context, cached input $0.26.
2026-06-09
Anthropic | Claude Mythos 5 released — less than half the price
Mythos Preview $25 / $125 → Mythos 5 $10 / $50. The public Fable 5 is the same price.
2026-05-22
DeepSeek | V4 Pro cut by about 75%
$1.74 / $3.48 → $0.435 / $0.87, with cached input down to one tenth. Called "permanent," but partly reversed three months later on 8/16 with the move to time-of-day pricing. (Some sources give the cut date as 5/31.)
2026-03-15
Zhipu | GLM-5-Turbo released
$1.20 / $4.00. A speed-focused line separate from GLM-5.2.
2026-05
xAI | Grok 4 / Grok 3 families retired together
Requests are redirected to current models. Grok 4 ($3/$15) and Grok 4.1 Fast ($0.20/$0.50) disappeared at this point. Current models from then on were 4.5 and 4.3.
2026-05-19
Google | Gemini Omni and 3.5 Pro announced at I/O
Omni's Flash version was released on 6/30, but 3.5 Pro was still unreleased as of 8/25. It has missed three targets (June, mid-July, early August), and its model ID, price and date are all undecided.
2026-04-21
OpenAI | GPT Image 2 released
Renders Japanese text accurately, and is well regarded for summary images and infographics.
2026-03-24
OpenAI | Sora 2 shutdown announced (in two stages)
Developers were notified. The app and web close on 4/26, the API on 9/24. This marked the shift of video leadership to Seedance, Veo and Kling.
Correction: this entry previously read "Sora 2 discontinued 3/24." 3/24 was the announcement date, not the end date.
2026-03-06
Anthropic | Default prompt-cache TTL shortened
1 hour → 5 minutes. The announcement was low-key, so articles from before this date say 1 hour.
2026-02-26
Google | Nano Banana 2 released
$0.067 per 1K image. Pro-level quality at one third of Nano Banana Pro's price; 3–5 seconds per image, native 4K.
2026-02-04
Kuaishou | Kling 3.0 released
Native 4K, 60fps, 15 seconds, multilingual lip-sync. Turbo version added in mid-June.
2026-02
ByteDance | Seedance 2.0 released
Price published on 3/5, ~$0.14/second. Took first place in the Artificial Analysis video ranking.