Unit prices, specs, batch discounts and prompt-caching rules for the major text, image, video, voice and music APIs. Click a column header to sort. Price movements are recorded in the change log at the bottom.
Latest moves — on a single day, 22 September, OpenAI and Anthropic both cut prices. ① GPT-6 Sol (9/22) is $2/$10, exactly half of GPT-5.6 Sol's current price ($4/$20), and a permanent price, not a promotion. GPT-6 Luna, released alongside it, is $0.10/$0.50. ② Claude Opus 5.5 (9/22) is $4/$20 (Opus 5 was $5/$25). Cache reads also fell from $0.50 to $0.20. Anthropic reports that it beats Fable 5.1, which costs 2.5x as much, on coding benchmarks. ③ Grok 4.7 (9/21) stays at $2/$6, the same as 4.6, on a larger base model. ④ Gemini 3.8 Live (9/15) is out. Audio in/out is $0.005 / $0.018 per minute, far cheaper than GPT-Live-1 ($0.05/min), though the two are not directly comparable (see the Voice section). ⑤ The Sora 2 API ends today, 9/24. Earlier versions of this page said Sora ended on 3/24. 3/24 was the announcement date; the app closed on 4/26 and the API closes on 9/24. This has been corrected. ⑥ Alibaba announced Qwen 4 Max (9/22), but only as an announcement, with no price or release date, so it is not in the tables.
The top tier has split into a "$10/$50 premium band" and a "$2–$4 practical band." GPT-6 Astra and Fable 5.1 sit together at $10/$50, while GPT-6 Sol, Opus 5.5, Grok 4.7 and Sonnet 5 crowd into the $2–$4 range. Early September was about the top end getting pricier; late September was about mid-tier price cuts. Anthropic has also said Sonnet 5.5 and Haiku 5.5 will arrive "in the coming weeks," so the mid and light tiers are likely to move again.
DeepSeek's peak hours overlap the Asian business day. Peak is weekdays (Mon–Fri) 01:00–04:00 and 06:00–10:00 UTC, which is 10:00–13:00 and 15:00–19:00 in Japan (Chinese public holidays excluded). Outside those 7 hours, prices are half. For users in Europe and the Americas, most working hours already fall off-peak; for batch jobs anywhere, scheduling outside the peak windows halves the bill.
Output tokens typically cost 4–6x as much as input. The total is driven mostly by output volume, so estimate "how much will it write" before choosing a model. Sorting by Tier lines up each company's equivalent models side by side.
| Provider | Model | Tier | Input | Output | Notes |
|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | Top | $10.00 | $50.00 | 9/3 newNew flagship. 1M context, cached input $1.00, cache write $12.50. Batch/Flex half price, Fast mode double. Sits above GPT-5.6 Sol: 2x on input, 1.7x on output |
| OpenAI | GPT-6 Sol | Mid | $2.00 | $10.00 | 9/22 newPermanent priceThe tier below Astra. Overall intelligence on par with GPT-5.6 Sol at half the price (5.6 Sol is $4/$20). Cached input $0.20, cache write $2.50. 1.05M context, 128K output; above 272K, input 2x and output 1.5x. Batch/Flex half, Fast double. Six effort levels, none to max. On DeepSWE and OSWorld, however, GPT-5.6 Sol still has the higher top score, at a higher cost per task. Its output is cheaper than Terra ($2/$12) at the same input price, so there is little reason left to choose Terra |
| OpenAI | GPT-6 Luna | Light | $0.10 | $0.50 | 9/22 newPermanent priceInput halved and output down 60% from GPT-5.6 Luna ($0.20/$1.20). Cached input $0.01. Among the cheapest input prices of any generally available text model. 1.05M context. In ChatGPT, Free and Go users can use it in the desktop app |
| Anthropic | Claude Fable 5.1 | Top | $10.00 | $50.00 | 9/1 newBase price unchanged. Cache reads cut 75%, from $1.00 to $0.25 (0.025x input; other Claude models are 0.1x, so this one is an exception). Cache write $12.50 (5 min) / $20 (1 hr). Batch $5/$25 |
| Anthropic | Claude Mythos 5.1 | Top | $10.00 | $50.00 | RestrictedRestricted-access version of Fable 5.1, via a vetting program |
| Anthropic | Claude Fable 5 | Top (prev.) | $10.00 | $50.00 | One generation before 5.1. Cache reads still $1.00 |
| Sakana AI | Fugu Ultra v2.0 | Top | $5.00 / $10.00 | $30.00 / $45.00 | 9/11 new★A Japanese company (Sakana AI). Right-hand figures apply above 272K tokens. Not a single model but an "orchestrator" that routes work to other companies' models and recombines the answers. For hard, many-step problems. 1M context |
| Sakana AI | Fugu Max | Mid | $2.00 | $6.00 | 9/11 newSame architecture as Ultra, aimed at results per dollar. Sakana says it undercuts Sonnet 5 and Kimi K3 on output price by 40–60% |
| OpenAI | GPT-5.6 Sol / GPT-5.5 | Upper (prev.) | $5.00 / promo $4.00 | $30.00 / promo $20.00 | 1M context. No deprecation announced. Promo price ($4/$20) runs at least until 11/21. GPT-6 Sol matches it at half the price, so there is little reason to pick it for new work |
| Anthropic | Claude Opus 5.5 | Top | $4.00 | $20.00 | 9/22 newFirst of the 5.5 generation. 20% cheaper than Opus 5; cache read $0.50→$0.20, cache write $6.25→$5. Batch $2/$10. 1M context, 128K output (300K in Batch beta), model ID claude-opus-5-5. Anthropic reports it beats Fable 5.1 ($10/$50) on Terminal-Bench 4.0, 66.4 vs 55.8, while trailing GPT-6 Astra on some benchmarks. The "about 40% cheaper" claim compares medium effort with high effort; the model change alone is 20%. MigrationThinking cannot be turned off / forced tool use errors / the old computer-use tool is rejected, so some Opus 5 code will not run unchanged. Sonnet 5.5 and Haiku 5.5 announced for "the coming weeks" |
| Anthropic | Claude Opus 5 | Top (prev.) | $5.00 | $25.00 | 1M context, five-level effort dial. Same price for five generations since Opus 4.5. Inference speed improved on 8/12 |
| Moonshot | Kimi K3 | Top | $3.00 / $0.30 | $15.00 | Right-hand figure is on cache hit. 1M context. No Batch support |
| Gemini 3.1 Pro | Top | $2.00 / $4.00 | $12.00 / $18.00 | Right-hand figures apply above 200K tokens. 3.5 Pro has been delayed three times and is unreleased, so this is the current Pro | |
| xAI | Grok 4.7 | Top | $2.00 / $4.00 | $6.00 / $12.00 | 9/21 newSame price as 4.6. A larger base model with longer reinforcement learning aimed at hours-long tasks. Right-hand figures apply above 200K. Cached input $0.50 ($1 above 200K), 500K context, four reasoning levels (low to xhigh), model ID grok-4.7. The double-price Fast variant is only available via Cursor and Grok Build, not the public API. xAI's own CursorBench 4.0 score is 46.3, short of Fable 5.1 (51.8). OpenRouter listed it at $1.60/$4.80 on 9/22, about 20% below list |
| xAI | Grok 4.6 | Top (prev.) | $2.00 / $4.00 | $6.00 / $12.00 | 8/12Same price as 4.5. 500K context, cached input $0.50, knowledge to 2026/2. Four reasoning depths (low/medium/high/xhigh). Artificial Analysis index 61, tied with GPT-5.6 Sol and one point behind Fable 5 |
| xAI | Grok 4.5 | Upper | $2.00 / $4.00 | $6.00 / $12.00 | Opus class. Doubles above 200K. Cached input $0.50. Not eligible for Batch discount |
| Zhipu | GLM-5.2 | Top | $0.40 – $1.40 | $1.26 – $4.40 | Conflicting prices744B MoE, 1M context, cached input $0.26. Z.ai first-party sources say $1.40/$4.40; aggregators such as OpenRouter say $0.40/$1.26, a gap of more than 3x. Check the route you actually use |
| DeepSeek | DeepSeek V4 Pro | Top | $0.66 / $1.32 | $1.98 / $3.96 | 8/16 increaseLeft is off-peak, right is peak. Was $0.435/$0.87 flat. Partly reverses May's 75% cut. Still cheap for a top-tier model. The shutdown planned for 9/14 was cancelled, and pricing continues unchanged |
| OpenAI | GPT-5.6 Terra | Mid | $2.00 | $12.00 | Cut 20% on 7/30 (from $2.50/$15.00) |
| Anthropic | Claude Sonnet 5 | Mid | $2.00 | $10.00 | Increase cancelledWas due to rise to $3/$15 on 9/1, but on 8/10 $2/$10 became the permanent price |
| Gemini 3.6 Flash | Mid | $1.50 | $7.50 | Released 7/21. 3.7 Flash will match this price when its intro pricing ends on 2027/1/1 | |
| Meta | Muse Spark 1.3 | Mid | ~$0.10 blended | ~$0.10 blended | 9/2 newSeparate input/output prices not published; only a blended price. A big cut from 1.2 ($1.25/$4.25) |
| Meta | Muse Spark 1.2 | Mid (prev.) | $1.25 | $4.25 | 1M context. An even cheaper tier exists for customers who allow training on their data |
| xAI | Grok 4.3 | Mid | $1.25 | $2.50 | Value option for long text. 500K–1M context. 20% off with Batch |
| Gemini 3.8 Flash | Mid | $0.75 | $3.75 | 9/2 newDoubles 2027/1/1Same price as 3.7 Flash and better on every benchmark. Intro price until 2026/12/31, then $1.50/$7.50. Batch/Flex half, Priority 1.8x. A vetted Cyber variant also exists | |
| Gemini 3.7 Flash | Mid (prev.) | $0.75 | $3.75 | Released 8/13. Same price as 3.8, so little reason to choose it | |
| Zhipu | GLM-5.3-Flash | Light | $0.15 | $0.50 | Promo ended 9/9Promo: $0.075 / $0.25The 50%-off launch price ended at 24:00 on 2026-09-09 (Singapore time) and list price returned. Intelligence Index 57 |
| Zhipu | GLM-5-Turbo | Mid | $1.20 | $4.00 | Released 2026/3/15. A speed-focused line separate from GLM-5.2 |
| Zhipu | GLM-4.6 | Mid | $0.43 | $1.74 | Older generation. 200K context |
| MiniMax | MiniMax-M3 | Mid | $0.30 | $1.20 | After the "Permanent 50% off." Input up to 512K, cache read $0.12 |
| Anthropic | Claude Haiku 4.5 | Light | $1.00 | $5.00 | On the expensive side for the light tier |
| Gemini 3.5 Flash-Lite | Light | $0.30 | $2.50 | For high throughput and low latency | |
| OpenAI | GPT-5.6 Luna | Light (prev.) | $0.20 | $1.20 | Cut 80% on 7/30 (from $1.00/$6.00). GPT-6 Luna ($0.10/$0.50) arrived on 9/22 and took over the lowest input price. No deprecation announced, but little reason to choose it for new work |
| DeepSeek | DeepSeek V4.1-Flash | Light | $0.15 / $0.30 | $0.60 / $1.20 | 9/10 newLeft is off-peak, right is peak. Model ID deepseek-flash. 552B MoE (8B active for input, 16B for output), image input, weights released under MIT. Cache hits are extremely cheap at $0.003 ($0.006 peak). Requests to the old names deepseek-v4-flash and deepseek-v4-flash-vision-exp are served by this model and billed at its price (the old models themselves are retired) |
| DeepSeek | DeepSeek V4 Flash | Light | $0.22 / $0.44 | $0.66 / $1.32 | 8/16 increaseLeft is off-peak, right is peak. Was $0.14/$0.28 flat. Automatic caching brings input down to $0.007 ($0.014 peak) |
| DeepSeek | DeepSeek V4 Flash Vision | Light / exp. | $0.22 / $0.44 | $0.66 / $1.32 | 8/21Experimental. Image input, same price as V4 Flash. 1M context, 384,000-token output limit |
| DeepSeek | DeepSeek R1 | Reasoning | $0.55 | $2.19 | Input $0.14 on cache hit |
For text inside images, especially non-Latin scripts such as Japanese, choose GPT Image 2 or Nano Banana Pro; both are established as rarely garbling characters. If you just need volume, Nano Banana 2 costs one third of Pro.
| Provider | Model | Tier | Price | Notes |
|---|---|---|---|---|
| OpenAI | GPT Image 2.5 (Flare / Sunburst) | Top | $0.25 1K $0.30 2K $0.60 4K | 9/8 newTwo variants: Flare for fast volume work, Sunburst for precise editing. Both have the same price. Official pricing is per token (text input $5.00 / image input $8.00 / image output $30.00 per 1M tokens; cached $1.25 and $2.00); the per-image figures on the left are reseller street prices. Supports both generation and editing of reference images |
| OpenAI | GPT Image 2 | Top | $0.006 low $0.053 med $0.211 high | Batch 50%At 1024×1024. High quality at 3840×2160 is $0.401. Landscape/portrait are slightly cheaper at medium–high ($0.041/$0.165). Accurate Japanese text |
| Nano Banana Pro (Gemini 3 Pro Image) | Top | $0.039 std $0.134 2K $0.24 4K | Batch 50%Japanese text rarely breaks. Strong at realistic people and text layering. Image generation does not support caching | |
| ByteDance | Seedream 5.0 Pro | Mid | $0.045 ≤2.36MP $0.09 above | 8/12Generation and editing, multi-reference fusion. First reference image free, $0.003 per extra. On BytePlus ModelArk and also fal.ai, Replicate, etc. A Lite version exists |
| Nano Banana 2 | Mid | $0.067 1K | Pro-level quality at Flash-level speed. 3–5 seconds per image, native 4K. One third the price of Nano Banana Pro | |
| xAI | grok-imagine-image-quality | Mid | $0.05 1K / $0.07 2K | Batch 50% |
| xAI | grok-imagine-image | Light | $0.02 | Batch 50%1K. Can go down to $0.01 |
| MiniMax | image-01 | Light | $0.0035 | Cheapest. When volume matters more than quality |
| OpenAI | GPT Image 1 | Retiring | $0.02 – $0.19 | Retires 2026/10/23Avoid for new projects |
Even within one company, prices differ by more than 10x across tiers. Sort by Tier to line up equivalent models across companies, instead of seeing Veo Lite / Fast / Standard scattered. Turnover is fast; model names from six months ago are often gone.
| Provider | Model | Tier | Resolution | Price | Notes |
|---|---|---|---|---|---|
| ByteDance | Seedance 2.5 | High | 1080p (reseller) | ~$0.57 | 7/31, API 8/75 seconds at 1080p is $2.843 on BytePlus. Native output of the official API is 480p/720p only; resellers' 1080p and 4K are upscaled. Up to 30 seconds with audio in one pass, references up to 30 images / 10 videos / 10 audio clips, timestamp-level editing |
| Veo 3.1 Standard | High | with audio | $0.40 | Some sources report $0.75/s via Vertex AI; prices vary widely by route | |
| Alibaba | Wan 3.0 | High | 1080p | $0.20 | 8/24$6 for 30 seconds at 1080p. Failed generations are not charged |
| MiniMax | H3 (Hailuo 3.0) | High | 2K (1440p) 24fps | $0.13 | 7/31Omni-modal. 5–15 seconds with stereo audio generated together. Accepts up to 9 images, 3 videos and 3 audio clips as references. There is no 1080p tier. $7.80 per minute |
| ByteDance | Seedance 2.0 | High | official | ~$0.14 | #1 qualityBilled at CNY 46 per 1M tokens. 15 seconds ≈ 308,880 tokens. Video input (editing) is CNY 28 per 1M |
| Luma AI | Ray 3 / Ray 3.2 | High | 1080p | ~$0.21 – $0.24 | 80 credits/second. About $12.60 for 60 seconds |
| Kuaishou | Kling 3.0 | Standard | 720p / 1080p | $0.075 – $0.14 | 4K, 60fps, 15 seconds, multilingual lip-sync. Official pricing uses credits (6–8/s without audio, 9–12/s with audio). Turbo version added in mid-June |
| Runway | Gen-4.5 | Standard | standard | $0.12 | 12 credits/second. About $1.20 for 10 seconds |
| MiniMax | H3 (third-party) | Standard | 2K / 768p | $0.124 / $0.077 | AIHubMix and others. Slightly cheaper than official. Segmind is more expensive at $0.1625 for 2K |
| Alibaba | Wan 3.0 | Standard | 720p | $0.10 | 8/24★Omni-Reference — besides text, images, audio and video, it accepts documents, spreadsheets, slides, PDFs and web pages as input and builds a video from them. 2–30 seconds; t2v, i2v and reference generation. Public beta 8/6, GA 8/24. Used in production for short dramas, ads, tourism promotion and music videos |
| Alibaba | Wan 2.7 | Standard | text-to-video | $0.10 | Older generation. Available on Together AI and others |
| Veo 3.1 Fast | Standard | standard | $0.10 – $0.15 | ||
| ByteDance | Seedance 2.0 Fast | Standard | 1080p | $0.09 | Rated "the cheapest at production quality" |
| MiniMax | H3 (Hailuo 3.0) | Standard | 768p | $0.08 – $0.09 | Effectively unusableOn the price list, but invitation-only closed beta. The API only accepts 2K as a resolution, and you are billed at the 2K price even if you specify 768p |
| ByteDance | Seedance 2.0 | Light | third-party | $0.04 – $0.067 | OpenRouter $0.067, PoYo from $0.04. Official API access is limited, so most real use goes through these |
| Luma AI | Ray 3 / Ray 3.2 | Light | 720p | ~$0.06 | 20 credits/second (at Plus plan rates) |
| xAI | Grok Imagine Video | Light | standard | $0.05 | Batch 50%$0.025/s with Batch. The only video model with an explicitly stated Batch discount |
| Alibaba | Wan 3.0 | Light | 480p | $0.05 | 8/24Alibaba Cloud Model Studio. CNY 0.30/s in mainland China (Beijing), CNY 0.37471/s in Singapore |
| Veo 3.1 Lite | Light | 720p no audio | $0.03 – $0.08 | Among the cheapest video options. First choice if you do not need audio | |
| Gemini Omni / Omni Flash | Unknown | — | Not published | Awaiting priceAnnounced at I/O on 5/19; Omni Flash released 6/30. As of 8/25 the official pricing page lists only Veo 3.1 under video, with no pay-as-you-go price | |
| OpenAI | Sora 2 | Ended | — | Discontinued | API ends 9/24App and web closed on 4/26. The API closes on 2026-09-24 (3/24 was the announcement to developers). OpenAI has not named a successor |
This area mixes "per minute" and "per token" pricing, so the numbers cannot be compared as they are. On top of that, GPT-Live-1 is priced for the voice layer only; the thinking is handed to a separate model and billed separately. Estimating from the listed price alone will leave you short.
| Provider | Model | Price | Notes |
|---|---|---|---|
| OpenAI | GPT-Live-1 | $0.05 / min | API since 9/10Its selling point is listening while talking (full duplex). Billed per second, and silence and background thinking time are billed at the same rate. Extra chargesThis covers the voice layer only; reasoning is delegated to a separate model such as GPT-6 Astra or GPT-5.6 Luna. A 10-minute call costs about $0.51 with Luna, about $0.86 with Astra. Endpoint v1/live/sessions, 128K context, knowledge to 2025/7. 12 voices, but the published list covers English accents plus Brazilian Portuguese and Filipino; Japanese is not listed. There is no mini in the API (mini is for ChatGPT's free plan) |
| OpenAI | gpt-realtime-2.1 | audio in $32 / out $64 per 1M tokens | Still sold at the same price after GPT-Live-1. Cached input $0.40. Measured at roughly $0.05/min. Input accumulates as the conversation grows, so beyond 10–15 minutes the effective per-minute cost rises |
| OpenAI | gpt-realtime-2.1-mini | audio in $10 / out $20 per 1M tokens | Measured at roughly $0.016/min, close to a third for the same use |
| Gemini 3.8 Live | audio in $0.005 / min audio out $0.018 / min | 9/15 newA native speech-to-speech model (no transcription step). Two model IDs: gemini-3.8-live and gemini-3.8-live-extended-thinking for multi-step reasoning. Reported to support 97 languages, including Japanese (secondary source; not confirmed in Google's own docs). About 1.18 s to first audio, near-real-time video input, SynthID watermark on all generated audio. UnverifiedPublished prices are for the standard model; the extended-thinking price has not been confirmed. Whether the identical numbers to 3.1 Flash Live are a coincidence or a copy is also unverified | |
| Gemini 3.1 Flash Live | audio in $0.005 / min audio out $0.018 / min | Among the cheapest for real-time conversation. For the same "1M audio output tokens," OpenAI charges $64 and Google $12. However, GPT-Live-1's $0.05/min covers only the listening-and-speaking layer, with thinking billed separately, whereas Google's model does the reasoning inside the voice model, so the numbers in this table cannot simply be ranked against each other |
| Provider | Model | Per 1,000 characters | Notes |
|---|---|---|---|
| ElevenLabs | Flash / Turbo | $0.05 | At 0.5 credits per character, effectively even cheaper |
| MiniMax | Speech Turbo | $0.06 ($60 / 1M chars) | |
| ElevenLabs | Multilingual v2 / v3 | $0.10 | High quality |
| MiniMax | Speech HD | $0.10 ($100 / 1M chars) | |
| OpenAI | gpt-4o-mini-tts | ~$0.015 / min | ConfirmedBilled per token, not per character (text input $0.60 / audio output $12.00 per 1M tokens), so the per-1,000-character cost varies with the text. Per minute, it is among the cheapest in this table |
| Cloud TTS (Standard / WaveNet) | $0.004 | Confirmed$4.00 / 1M characters. 4M characters free per month. Voice quality matches the price | |
| Cloud TTS (Neural2) | $0.016 | $16.00 / 1M characters | |
| Cloud TTS (Chirp 3: HD) | $0.03 | $30.00 / 1M characters. Google's high-quality tier | |
| Cloud TTS (Studio) | $0.16 | $160.00 / 1M characters. 40x Standard. Top tier for narration | |
| Gemini 3.1 Flash TTS | per token | Text input $1.00 / audio output $20.00 per 1M tokens. Free tier available, half price with Batch. A separate line from the per-character Cloud TTS, so not directly comparable |
In this area many companies offer no official pay-as-you-go API. Suno has no official API and is used via third parties; Udio and Stable Audio are subscription only.
| Provider | Model | Price | Notes |
|---|---|---|---|
| Suno | v4 family | $0.014 – $0.111 / song | No official API; third-party street prices. Direct subscription from Pro $10/month |
| Lyria 3 Pro | $0.05 – $0.15 / song | Via MusicAPI, 10 credits per task. The RealTime version streams over WebSocket | |
| MiniMax | Music 3.0 / 2.6 | $0.15 / up to 5 min | No new sign-ups since 8/20New users cannot sign up. Existing paying users only. The free tier (music-3.0-free etc.) is gone. New users are directed to the MiniMax Audio app or the open weights on Hugging Face. The lyrics API is the same. Music 3.0 weights were released on 8/13 |
| ElevenLabs | Music | $0.80 / min | |
| Udio | Standard / Pro | $10 – $30 / month | No pay-as-you-go API |
| Stability AI | Stable Audio | $5.99 – $89.99 / month | Pay-as-you-go API not yet researched |
A Batch API lets you submit many requests at once asynchronously and collect results within 24 hours, in exchange for a lower price. It suits bulk jobs that do not need an immediate response. Prompt caching reuses the same long prefix to cut input cost. The two can be combined.
| Provider | Batch | Batch discount | Cache | Notes |
|---|---|---|---|---|
| Anthropic | Yes | 50% off | 90% off | All active models. Batch and caching can be combined |
| OpenAI | Yes | 50% off | Up to 90% | GPT-4o and later. From GPT-5.6, cache writes carry a 1.25x surcharge. o3/o4-mini etc. do not support caching in Batch |
| Yes | 50% off | Implicit by default | Implicit caching applies automatically from Gemini 2.5. Image generation does not support caching. Cache hits work inside Batch too | |
| Zhipu | Yes | 50% off | Yes | Both across the GLM lineup |
| DeepSeek | Yes | ~50% off | Automatic | Batch is mostly via hosts such as SiliconCloud. Caching is automatic, no opt-out needed |
| xAI | Some models | Text 20% / image & video 50% | Yes | Text Batch discount applies only to Grok 4.3 and the 4.20 family; Grok 4.5 is excluded. We could not confirm a discount for 4.6 or 4.7. Images and video are a separate category at a flat 50% off. Max 25MB per request |
| Moonshot | Older models only | 60% of list price | Yes | Batch covers K2.7 Code / K2.6 / K2.5. The latest K3 does not support Batch |
| MiniMax | Limited info | Not disclosed | Automatic | Caching is automatic with no setup. Batch is hinted at, but the official docs give no discount rate |
| Meta | To research | Unconfirmed | To research | The Model API only entered public preview on 7/9, and discount schemes appear not to be in place yet |
Almost every provider measures idle time since the cache was last used. Each hit resets the timer, so as long as an agent keeps sending requests, even a 5-minute TTL will not expire. The risk is when the gap between requests exceeds the TTL: a single tool run takes too long, the agent waits on a human decision, or it only calls once every few tens of minutes.
Gemini's explicit cache is the one exception. It is a fixed 60 minutes from creation and is not extended by use. For long sessions, set a longer TTL or extend it explicitly via the update API.
| Provider | Type | Default TTL | Clock starts | Extended on hit | Notes |
|---|---|---|---|---|---|
| Anthropic | Prompt caching | 5 min | Last use | Resets | Can be extended to 1 hour with ttl:"1h", but writes then cost 2x (normally 1.25x), so you need 3+ reads to break even. The default changed from 1 hour to 5 minutes on 2026/3/6 |
| OpenAI | GPT-5.5 and earlier | 5–10 min idle | Last use | Resets | Two limits: "cleared after 5–10 minutes idle" and "at most 1 hour after last use." GPT-5 family / 4.1 support extended retention up to 24 hours |
| OpenAI | GPT-5.6 and later | 30 min | Last use | Resets | Kept for at least 30 minutes. Much more relaxed than earlier generations |
| Implicit | Not disclosed | Automatic | Automatic | On by default from Gemini 2.5. No setup | |
| Explicit | 60 min | Creation | Not extended | The only case that differs from other providers. Set freely with ttl/expire_time, no upper or lower bound. An API exists to update the TTL after creation | |
| MiniMax | Automatic caching | 5 min | Last use | Extended free | Explicitly stated as "no setup required" |
| DeepSeek | Context caching | No TTL | — | — | Cached on disk; unused entries are deleted after hours to days. You need not think about expiry, but there is also no way to guarantee retention |
| Moonshot | Context caching | Managed | — | — | Optimized automatically by the system. There is no TTL parameter at all |
| Zhipu | Caching | Unconfirmed | — | — | To researchThe cached input price ($0.26/1M) is published, but the expiry rules could not be found |
A record of "what it used to cost" and "since when this price applies." Prices mostly go down, but introductory prices can be extended or doubled, so dated changes are noted ahead of time.
gpt-6-sol and gpt-6-luna; 1.05M context, 128K output; above 272K, input 2x and output 1.5x. Batch/Flex half, Fast double. Six effort levels: none / low / medium / high / xhigh / max.claude-opus-5-5, 1M context, 128K output (300K in the Message Batches beta), knowledge to 2026/6. Fast mode (Claude API only, research preview) gives about 2.5x output speed at 2x the price. Also on AWS, Google Cloud and Microsoft.grok-4.7, four reasoning levels (low to xhigh).gemini-3.8-live (conversation at scale) and gemini-3.8-live-extended-thinking (with multi-step reasoning). Audio in $0.005 / min, audio out $0.018 / min, about 2.3 cents per conversation minute.deepseek-v4-pro are handled by V4.1-Flash" — was withdrawn before it took effect.deepseek-v4-flash and deepseek-v4-flash-vision-exp are discontinued, and requests to them are served by V4.1-Flash at Flash prices. Do not confuse this with V4 Pro.v1/live/sessions, 128K context (at 90% it hands off to an 8,192-token summary), instructions up to 16,384 tokens, knowledge to 2025-07-31. 12 voices, but the published list covers English accents (Australian, British, Irish, North American, Southern US) plus Brazilian Portuguese and Filipino; Japanese is not listed.gpt-image-2.5-flare and gpt-image-2.5-sunburst. Flare is for fast volume work, Sunburst for precise editing. Both have the same price.gpt-6-astra. Batch and Flex half price, Fast mode double. Sits above the previous top model GPT-5.6 Sol ($5 / $30): 2x on input, 1.7x on output. Rolled out first through Daybreak, a program for cybersecurity companies.music-3.0-free and others) is gone. New users are directed to the MiniMax Audio app or the open weights on Hugging Face.