Same model, 43 prices: this week in competitive pricing
GLM-5.3 costs between $0.83 and $4.30 a million tokens depending only on which endpoint you pick. The weekly price, spread and outage ledger for the platforms developers actually buy from.

Updated 8 Oct 2026 · next update 13 Oct 2026
Every price on this page is a list rate, in US dollars per million tokens, read from the platform's own endpoint listing on the date above. Where one model is served by several companies, every one of those endpoints is priced separately here. Nothing on this page is sponsored and no platform paid for placement.
The short version
- GLM-5.3 is the most contested model on the internet right now. 43 companies serve it, and the cheapest of them charges 5.2 times less than the dearest for the same job: OpenInference at $0.832 against Alibaba at $4.30, per million tokens at a 3:1 input-to-output mix.
- Widest spread anywhere in the basket: DeepSeek V4 Pro, 9.5x between StreamLake and Reka — but only 15 endpoints serve it, so treat that as a thin market rather than a saving.
- Cheapest on the whole basket: DigitalOcean, typically 26% under the median endpoint for the same weights, across 6 of the 8 models tracked.
- Then: StreamLake at 25% under the median and DeepInfra at 16% under.
- Most expensive name brand: Fireworks runs 50% above the median endpoint across 4 models. You are paying for throughput, support and a dashboard, not for weights.
- 18 service incidents were posted by these platforms in the last seven days, plus 21 scheduled maintenance windows.
- Degraded as of 2026-10-08: Cloudflare AI Gateway (Partial System Outage), Novita AI (degraded).
- 162 price points are tracked on this page, across 49 providers and 8 reference models. The full tables are below the charts.
First edition. This is the baseline. Every number on the page is dated and every price was read from the provider's own listing, so from next Tuesday the week-over-week moves start being measured against today. Come back then for the arrows.
The same model, dozens of prices
This is the part of the market that no pricing page tells you. The model is fixed — same weights, same quantization where we can see it, same day — and the bill is not. The chart sorts every endpoint by what 1M input plus 1M output actually costs at a 3:1 chat mix.

The gap is not a rounding difference. OpenInference and Alibaba are selling the same weights 5.2x apart, and the spread is widest exactly where competition is fiercest — GLM-5.3 has more endpoints than anything else in the basket, and the most endpoints means the most disagreement about what the weights are worth. Note what the extremes have in common: the cheapest endpoints are the ones with the least to lose, and the dearest are often the ones with a brand to defend.

Input price is not the bill
A cheap input rate with a dear output rate is not a cheap model, and several of the endpoints above are priced exactly that way — a headline number that wins the comparison table and an output rate that wins it back. Rank on the blend, and pin the endpoint you actually want, because the gap between endpoints on the same model is larger than any platform fee in this market.

The reliability ledger
Price is half of it. The other half is whether the endpoint answers. Availability below is the worst monitor on each platform over the period its status page reports — not the average, the worst, because the worst is what your users see.

Incidents in the last seven days (18)
| Date | Platform | Impact | Incident |
|---|---|---|---|
| 2026-10-08 | Cloudflare AI Gateway | minor | Network Performance Issues in Sofia |
| 2026-10-08 | Fireworks AI | minor | Service Degradation for GLM 5.3 Flash |
| 2026-10-08 | Fireworks AI | minor | Service Degradation for GLM 5.3 Flash |
| 2026-10-08 | Cloudflare AI Gateway | minor | Network Performance Issues in Kochi |
| 2026-10-08 | Cloudflare AI Gateway | none | Network Performance Issues in Kochi |
| 2026-10-07 | Fireworks AI | minor | Service Degradation for GLM 5.3 |
| 2026-10-07 | Fireworks AI | minor | Service Degradation for Qwen 3.8 Max |
| 2026-10-07 | Fireworks AI | minor | Service Degradation for GLM 5.3 Flash |
| 2026-10-06 | Vercel AI Gateway | major | Vercel Auth login failures for pages with Deployment Protection enabled |
| 2026-10-06 | Vercel AI Gateway | minor | Elevated Sandbox error rates in Paris (cdg1) |
| 2026-10-05 | Baseten | major | DeepSeek v4.1 outage on Model APIs. Dedicated Inference is not affected |
| 2026-10-05 | Fireworks AI | minor | Service Degradation for GLM 5.3 |
| 2026-10-05 | Fireworks AI | minor | Service Degradation for Kimi K3 US |
| 2026-10-05 | Fireworks AI | minor | Service Degradation for Qwen 3.8 Max |
| 2026-10-05 | Fireworks AI | minor | Service Degradation for Deepseek V4.1 Flash |
| 2026-10-04 | Fireworks AI | minor | Service Degradation for Deepseek V4.1 Flash |
| 2026-10-02 | Fireworks AI | minor | Service Degradation for GLM 5.3 Fast |
| 2026-10-01 | Fireworks AI | minor | Service Degradation for Deepseek V4.1 Flash |
Scheduled maintenance posted in the same window (21)
Maintenance is not an outage, so it is counted separately. It is still the thing that wakes you up if you scheduled a training run.
| Date | Platform | Window |
|---|---|---|
| 2026-10-08 | Cloudflare AI Gateway | DUB (Dublin) on 2026-10-09 |
| 2026-10-08 | Cloudflare AI Gateway | DUB (Dublin) on 2026-10-08 |
| 2026-10-08 | Cloudflare AI Gateway | ZRH (Zurich) on 2026-10-12 |
| 2026-10-08 | Cloudflare AI Gateway | ATL (Atlanta) on 2026-10-12 |
| 2026-10-08 | Cloudflare AI Gateway | BOS (Boston) on 2026-10-12 |
| 2026-10-07 | Cloudflare AI Gateway | HKG (Hong Kong) on 2026-10-14 |
| 2026-10-07 | Cloudflare AI Gateway | FCO (Rome) on 2026-10-09 |
| 2026-10-07 | Cloudflare AI Gateway | KEF (Reykjavík) on 2026-10-09 |
| 2026-10-07 | Cloudflare AI Gateway | LUX (Luxembourg City) on 2026-10-12 |
| 2026-10-07 | Cloudflare AI Gateway | PMO (Palermo) on 2026-10-09 |
What a picture, a second of video and an hour of GPU cost
A token table cannot hold this half of the market. fal, Replicate and Segmind bill per second of video, per image, per thousand characters, or per hour of GPU — so here they are on their own, read from their own pricing pages on the date at the top of this page. Every row below carries the price the vendor publishes; where a vendor has moved a number, the row is marked unverified rather than guessed at.
The same GPU, three platforms, per hour
| GPU | fal | Replicate | Segmind |
|---|---|---|---|
| B300 | $12.99 | — | — |
| H100 | $4.50 | $5.49 | $9.00 |
| B200 | $7.99 | — | — |
| H200 | $6.00 | $5.49 | — |
| A100 | — | $5.04 | $4.32 |
| L40S | — | — | $2.66 |
| T4 | — | $0.81 | — |
| A5000 | — | — | $0.68 |
That table is the whole argument of this page in four rows: the identical H100 sells for $0 an hour on one platform and more than double that on another, and you are buying the same silicon by the same clock. What you are actually choosing between is the queue, the tooling and the support you get behind it.
Video, by the second (fal)
Prices are per second of generated video.
| Model | Unit | Price | What $1 buys |
|---|---|---|---|
| MiniMax H3 Max (text, image or reference to video) | per second | $0.05 | 20 second |
| MiniMax H3 Max Turbo (image to video) | per second | $0.025 | 40 second |
| Kling Video v3 Pro (text to video) | per second | $0.14 | 7 second |
Images, by the image (fal)
Prices are per generated image.
| Model | Unit | Price | What $1 buys |
|---|---|---|---|
| Nano Banana 2 | per image | $0.08 | 12 image |
| Nano Banana Pro | per image | $0.15 | 7 image |
| Nano Banana | per image | $0.04 | 25 image |
Speech and sound, by the character or the second (fal)
| Model | Unit | Price | What $1 buys |
|---|---|---|---|
| ElevenLabs TTS Multilingual v2 | per 1,000 characters | $0.1 | 10 1,000 characters |
| ElevenLabs TTS Turbo v2.5 | per 1,000 characters | $0.05 | 20 1,000 characters |
| ElevenLabs Sound Effects V2 | per second | $0.002 | 500 second |
Segmind's plans
Segmind is the one platform in this group that publishes a subscription ladder beside its pay-as-you-go rates, which matters if your volume is steady: the credits included in each tier are worth exactly the fee, so you are buying rate limits, storage and support rather than a discount.
| Plan | Price | Monthly credits | Limits |
|---|---|---|---|
| Flexible (pay as you go) | $10 credit to start | 60 requests/min | 1 GB storage |
| Pro | $39/mo, $50 monthly credits | 120 requests/min | 10 GB storage |
| Business | $99/mo, $99 monthly credits | 500 requests/min | 100 GB storage |
| Scale | $599/mo, $599 monthly credits | 1 | 000 requests/min pooled, 1 TB storage |
Free tiers, credits and standing offers
Only offers stated on the platform's own page are listed here, because a discount code that expired last quarter is worse than no code at all. Check the linked page before you plan around any of it.
- OpenRouter — free tier: 25+ free models from 4 free providers at 50 requests a day, no card. The Standard plan carries a 5.5% platform fee on credit purchases, Business 8%, Enterprise negotiable — so buying credit in larger blocks beats topping up in dribs.
- Together AI — its pricing page carries a batch column beside the serverless rate, and a cached-input rate on most models, both well under the headline number.
- Fireworks AI, DeepInfra and Novita AI — postpaid per-token with no minimum; DeepInfra publishes no free tier at all, which is why its rates are the ones the others get measured against.
- Cloudflare AI Gateway — the gateway layer itself is free to run in front of your existing provider keys, which is the cheapest way to get fallback and logging without changing your bill.
- Replicate — Wan 3.0 is 30% off this week, per the platform's own pages.
- The real discount is the endpoint. The spread inside a single model on this page is larger than any platform fee in the market: pinning the cheap endpoint and falling back when it breaks is worth more than any credit promo.
The platforms, and what each one is
- OpenRouter — the router most teams start on: 500+ models, 80+ providers, pass-through rates plus a 5.5% platform fee on credit purchases on the Standard plan (8% on Business), and a free tier of 25+ free models at 50 requests a day. Its public model endpoint is where the prices on this page come from.
- Together AI — serverless plus dedicated capacity, a batch column on its pricing page, and a cached-input rate on most models. Premium list rates, deep stack.
- Fireworks AI — production inference with flat rate ceilings.
- DeepInfra — the price leader on open weights for most of this year, no free tier, postpaid.
- Novita AI, Baseten and Modal — GPU-second and dedicated-endpoint shapes alongside per-token.
- Replicate and fal — per-second compute for image and video models, where per-token pricing does not apply.
- Cloudflare AI Gateway and Vercel AI Gateway — gateway layers over the same providers, billed at pass-through plus their own fee.
- Groq and Cerebras — custom silicon, priced for speed rather than for the cheapest token.
How this page is made
- Prices come from each model's own endpoint listing, which publishes a rate per provider for the same weights. Where a platform also publishes its own price list, we check one against the other: Together's published rate for GLM-5.3 and the rate its endpoint reports agree, which is the check that keeps this page honest.
- The blend is (3 x input + output) / 4. Input tokens outnumber output tokens roughly 3:1 in ordinary chat and agent traffic; change that ratio and the ranking moves, which is why both columns are printed.
- Uptime and incidents are read from each platform's own status page, and BetterStack-hosted pages report per-day monitor history while Statuspage-hosted pages report incident names and impact.
- Prices move without notice. Every figure here carries the date it was read; the page is rebuilt every Tuesday morning and the tables are regenerated, not edited.
The full tables
Every endpoint on every reference model, cheapest first. in and out are US dollars per million tokens; blend is the 3:1 mix. Quantization is as published, and - means the provider does not say.
DeepSeek V4 Pro — 15 endpoints, 10x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| StreamLake | $0.287 | $0.575 | $0.359 | 1024K | fp8 |
| GMICloud | $0.957 | $1.91 | $1.20 | 1048K | fp8 |
| Parasail | $0.45 | $3.48 | $1.21 | 1048K | fp8 |
| Relace | $0.284 | $4.20 | $1.26 | 1048K | fp4 |
| DigitalOcean | $1.04 | $2.09 | $1.31 | 1048K | unknown |
| Cloudflare | $1.15 | $2.55 | $1.50 | 1048K | unknown |
| DeepInfra | $1.30 | $2.60 | $1.62 | 1048K | fp8 |
| Alibaba | $1.42 | $2.83 | $1.77 | 1000K | fp8 |
| SiliconFlow | $1.50 | $3.13 | $1.91 | 1048K | fp8 |
| Novita | $1.60 | $3.20 | $2.00 | 1048K | fp8 |
| Venice | $1.65 | $3.30 | $2.06 | 1000K | unknown |
| AtlasCloud | $1.68 | $3.38 | $2.10 | 1048K | fp4 |
| Baidu | $1.69 | $3.38 | $2.11 | 1048K | fp8 |
| Azure | $1.91 | $3.83 | $2.39 | 1048K | unknown |
| Reka | $1.05 | $10.50 | $3.41 | 1048K | unknown |
DeepSeek V4.1 Flash — 31 endpoints, 7x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| Decart | $0.09 | $0.18 | $0.113 | 1048K | fp4 |
| Wafer | $0.03 | $0.4 | $0.122 | 1048K | unknown |
| Sail Research | $0.08 | $0.4 | $0.16 | 1048K | fp4 |
| Morph | $0.02 | $0.6 | $0.165 | 1048K | fp8 |
| DeepInfra | $0.14 | $0.42 | $0.21 | 1048K | fp8 |
| InferenceNet | $0.105 | $0.6 | $0.229 | 1040K | fp8 |
| StreamLake | $0.15 | $0.6 | $0.262 | 1024K | fp8 |
| Alibaba | $0.15 | $0.6 | $0.262 | 1000K | unknown |
| DeepSeek | $0.15 | $0.6 | $0.262 | 1048K | unknown |
| OpenInference | $0.02 | $1.00 | $0.265 | 1048K | fp4 |
| Relace | $0.016 | $1.20 | $0.312 | 1048K | unknown |
| CoreWeave | $0.2 | $0.65 | $0.312 | 1048K | fp8 |
| GMICloud | $0.18 | $0.72 | $0.315 | 1048K | fp8 |
| Ionstream | $0.05 | $1.15 | $0.325 | 1048K | unknown |
| Novita | $0.195 | $0.78 | $0.341 | 1048K | fp8 |
| Phala | $0.21 | $0.84 | $0.367 | 1048K | unknown |
| DekaLLM | $0.12 | $1.20 | $0.39 | 1048K | unknown |
| DigitalOcean | $0.225 | $0.9 | $0.394 | 1048K | unknown |
| Makora | $0.27 | $1.15 | $0.49 | 1048K | fp8 |
| Crusoe | $0.29 | $1.20 | $0.517 | 1048K | fp8 |
| Baidu | $0.3 | $1.20 | $0.525 | 1048K | fp8 |
| AtlasCloud | $0.3 | $1.20 | $0.525 | 1048K | fp8 |
| BaseTen | $0.3 | $1.20 | $0.525 | 1048K | fp8 |
| Together | $0.3 | $1.20 | $0.525 | 1048K | unknown |
| SiliconFlow | $0.3 | $1.20 | $0.525 | 1048K | fp8 |
| Modal | $0.3 | $1.20 | $0.525 | 1048K | unknown |
| BaseTen | $0.3 | $1.20 | $0.525 | 1048K | fp8 |
| Parasail | $0.3 | $1.20 | $0.525 | 1048K | fp8 |
| Fireworks | $0.3 | $1.20 | $0.525 | 1048K | unknown |
| Venice | $0.3 | $1.20 | $0.525 | 1000K | fp8 |
| Fireworks | $0.45 | $1.80 | $0.788 | 1048K | unknown |
gpt-oss-120b — 23 endpoints, 7x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| CoreWeave | $0.03 | $0.17 | $0.065 | 131K | fp4 |
| DekaLLM | $0.03 | $0.18 | $0.068 | 131K | bf16 |
| DeepInfra | $0.037 | $0.17 | $0.07 | 131K | bf16 |
| AkashML | $0.037 | $0.187 | $0.074 | 131K | bf16 |
| Crusoe | $0.05 | $0.25 | $0.1 | 131K | bf16 |
| Mancer 2 | $0.05 | $0.25 | $0.1 | 131K | fp8 |
| Novita | $0.05 | $0.25 | $0.1 | 131K | fp4 |
| DigitalOcean | $0.06 | $0.42 | $0.15 | 128K | unknown |
| $0.09 | $0.36 | $0.158 | 131K | unknown | |
| BaseTen | $0.1 | $0.5 | $0.2 | 128K | fp4 |
| BaseTen | $0.1 | $0.5 | $0.2 | 128K | fp4 |
| Amazon Bedrock | $0.15 | $0.6 | $0.262 | 131K | unknown |
| Nebius | $0.15 | $0.6 | $0.262 | 131K | fp4 |
| Amazon Bedrock | $0.15 | $0.6 | $0.262 | 131K | unknown |
| DeepInfra | $0.15 | $0.6 | $0.262 | 131K | bf16 |
| SiliconFlow | $0.15 | $0.6 | $0.262 | 131K | fp8 |
| Phala | $0.15 | $0.6 | $0.262 | 131K | unknown |
| Together | $0.15 | $0.6 | $0.262 | 131K | unknown |
| Groq | $0.15 | $0.6 | $0.262 | 131K | unknown |
| Parasail | $0.1 | $0.75 | $0.263 | 131K | fp4 |
| Mara | $0.15 | $0.75 | $0.3 | 131K | unknown |
| SambaNova | $0.14 | $0.95 | $0.343 | 131K | unknown |
| Cerebras | $0.35 | $0.75 | $0.45 | 131K | fp16 |
Llama 3.3 70B — 11 endpoints, 7x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| DeepInfra | $0.1 | $0.32 | $0.155 | 131K | fp8 |
| Novita | $0.135 | $0.4 | $0.201 | 12K | bf16 |
| AkashML | $0.2 | $0.52 | $0.28 | 131K | fp8 |
| Parasail | $0.22 | $0.5 | $0.29 | 131K | fp8 |
| SambaNova | $0.45 | $0.9 | $0.562 | 131K | unknown |
| Groq | $0.59 | $0.79 | $0.64 | 131K | unknown |
| CoreWeave | $0.71 | $0.71 | $0.71 | 128K | fp16 |
| $0.72 | $0.72 | $0.72 | 128K | unknown | |
| $0.72 | $0.72 | $0.72 | 128K | unknown | |
| Cloudflare | $0.293 | $2.25 | $0.783 | 24K | fp8 |
| Together | $1.04 | $1.04 | $1.04 | 131K | unknown |
GLM-5.3 — 43 endpoints, 5x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| OpenInference | $0.04 | $3.21 | $0.832 | 1048K | unknown |
| Wafer | $0.039 | $3.39 | $0.877 | 1048K | unknown |
| Wafer | $0.17 | $3.39 | $0.975 | 1048K | unknown |
| Sail Research | $0.2 | $3.40 | $1.00 | 1048K | fp8 |
| Sail Research | $0.2 | $3.40 | $1.00 | 1048K | fp8 |
| DeepInfra | $0.562 | $2.50 | $1.05 | 1048K | fp4 |
| Novita | $0.7 | $2.20 | $1.07 | 1048K | fp8 |
| Reka | $0.14 | $4.20 | $1.16 | 1048K | unknown |
| InferenceNet | $0.14 | $4.40 | $1.21 | 1048K | fp4 |
| Makora | $0.14 | $4.40 | $1.21 | 1048K | fp4 |
| AkashML | $0.19 | $4.40 | $1.24 | 1048K | fp8 |
| Phala | $0.84 | $2.64 | $1.29 | 1048K | unknown |
| Morph | $0.476 | $3.74 | $1.29 | 1048K | fp8 |
| Decart | $0.842 | $2.65 | $1.29 | 1048K | fp4 |
| Inceptron | $0.6 | $3.39 | $1.30 | 1048K | fp4 |
| DigitalOcean | $0.91 | $2.86 | $1.40 | 1048K | unknown |
| GMICloud | $0.98 | $3.08 | $1.50 | 1048K | fp8 |
| SiliconFlow | $1.12 | $3.52 | $1.72 | 1048K | fp8 |
| Alibaba | $1.19 | $3.74 | $1.83 | 1000K | unknown |
| Friendli | $1.26 | $3.96 | $1.94 | 1048K | unknown |
| Io Net | $1.25 | $4.40 | $2.04 | 262K | fp8 |
| Mistral | $1.40 | $4.40 | $2.15 | 1048K | nvfp4 |
| Baidu | $1.40 | $4.40 | $2.15 | 1048K | fp8 |
| BaseTen | $1.40 | $4.40 | $2.15 | 1048K | fp4 |
| Mistral | $1.40 | $4.40 | $2.15 | 1048K | nvfp4 |
| Nebius | $1.40 | $4.40 | $2.15 | 1024K | fp4 |
| Crusoe | $1.40 | $4.40 | $2.15 | 1048K | fp4 |
| PrimeIntellect | $1.40 | $4.40 | $2.15 | 1048K | unknown |
| Venice | $1.40 | $4.40 | $2.15 | 1000K | unknown |
| Together | $1.40 | $4.40 | $2.15 | 1048K | unknown |
| Parasail | $1.40 | $4.40 | $2.15 | 1048K | fp8 |
| Modal | $1.40 | $4.40 | $2.15 | 1048K | unknown |
| BaseTen | $1.40 | $4.40 | $2.15 | 1048K | fp4 |
| Fireworks | $1.40 | $4.40 | $2.15 | 1048K | unknown |
| Cloudflare | $1.40 | $4.40 | $2.15 | 1048K | unknown |
| AtlasCloud | $1.40 | $4.40 | $2.15 | 1048K | fp8 |
| Z.AI | $1.40 | $4.40 | $2.15 | 1048K | fp8 |
| Mistral | $1.54 | $4.84 | $2.37 | 1048K | nvfp4 |
| Relace | $0.031 | $12.00 | $3.02 | 1048K | unknown |
| Fireworks | $2.10 | $6.60 | $3.23 | 1048K | unknown |
| BaseTen | $2.10 | $6.60 | $3.23 | 1048K | fp8 |
| BaseTen | $2.10 | $6.60 | $3.23 | 1048K | fp8 |
| Alibaba | $2.80 | $8.80 | $4.30 | 1000K | unknown |
Kimi K2.7 Code — 12 endpoints, 3x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| StreamLake | $0.713 | $3.00 | $1.28 | 256K | unknown |
| Inceptron | $0.671 | $3.35 | $1.34 | 262K | int4 |
| CoreWeave | $0.71 | $3.50 | $1.41 | 262K | int4 |
| Venice | $0.75 | $3.50 | $1.44 | 256K | int4 |
| SiliconFlow | $0.859 | $3.80 | $1.59 | 262K | fp8 |
| Novita | $0.912 | $3.84 | $1.64 | 262K | int4 |
| Nebius | $0.95 | $4.00 | $1.71 | 262K | fp4 |
| GMICloud | $0.95 | $4.00 | $1.71 | 262K | fp8 |
| Alibaba | $0.95 | $4.00 | $1.71 | 262K | fp8 |
| Cloudflare | $0.95 | $4.00 | $1.71 | 262K | unknown |
| Moonshot AI | $0.95 | $4.00 | $1.71 | 262K | int4 |
| Moonshot AI | $1.90 | $8.00 | $3.42 | 262K | int4 |
Kimi K2.6 — 17 endpoints, 2x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| Inceptron | $0.465 | $2.45 | $0.961 | 262K | int4 |
| DigitalOcean | $0.57 | $2.40 | $1.03 | 262K | unknown |
| StreamLake | $0.599 | $2.52 | $1.08 | 256K | fp8 |
| Chutes | $0.5 | $2.85 | $1.09 | 262K | int4 |
| CoreWeave | $0.65 | $3.41 | $1.34 | 262K | fp4 |
| Crusoe | $0.7 | $3.50 | $1.40 | 262K | bf16 |
| SiliconFlow | $0.77 | $3.40 | $1.43 | 262K | fp8 |
| DeepInfra | $0.75 | $3.50 | $1.44 | 262K | fp4 |
| Venice | $0.75 | $3.50 | $1.44 | 256K | int4 |
| Parasail | $0.75 | $3.50 | $1.44 | 262K | int4 |
| Novita | $0.8 | $3.40 | $1.45 | 262K | unknown |
| GMICloud | $0.855 | $3.60 | $1.54 | 262K | fp8 |
| Baidu | $0.95 | $4.00 | $1.71 | 262K | fp4 |
| AtlasCloud | $0.95 | $4.00 | $1.71 | 262K | int4 |
| Cloudflare | $0.95 | $4.00 | $1.71 | 262K | unknown |
| Moonshot AI | $0.95 | $4.00 | $1.71 | 262K | int4 |
| Phala | $1.09 | $4.60 | $1.97 | 262K | unknown |
Qwen3.5 397B A17B — 10 endpoints, 2x spread
| Endpoint | in | out | blend | context | quant |
|---|---|---|---|---|---|
| Alibaba | $0.39 | $2.34 | $0.877 | 262K | unknown |
| DeepInfra | $0.45 | $3.00 | $1.09 | 262K | fp8 |
| Parasail | $0.5 | $3.60 | $1.27 | 262K | fp8 |
| DigitalOcean | $0.55 | $3.50 | $1.29 | 131K | unknown |
| Phala | $0.55 | $3.50 | $1.29 | 262K | unknown |
| AtlasCloud | $0.55 | $3.50 | $1.29 | 262K | fp8 |
| StreamLake | $0.6 | $3.60 | $1.35 | 256K | unknown |
| GMICloud | $0.6 | $3.60 | $1.35 | 262K | fp8 |
| Novita | $0.6 | $3.60 | $1.35 | 262K | unknown |
| Venice | $0.75 | $4.50 | $1.69 | 128K | unknown |
Sources
- OpenRouter model and endpoint catalogue: https://openrouter.ai/api/v1/models and
/api/v1/models/{author}/{slug}/endpoints(public, no key). - OpenRouter pricing: https://openrouter.ai/pricing
- Together AI pricing: https://www.together.ai/pricing
- Fireworks AI pricing: https://fireworks.ai/pricing
- DeepInfra pricing: https://deepinfra.com/pricing
- Novita AI pricing: https://novita.ai/pricing
- Groq pricing: https://groq.com/pricing
- Cerebras pricing: https://www.cerebras.ai/pricing
- Replicate pricing: https://replicate.com/pricing
- fal pricing: https://fal.ai/pricing
- Baseten pricing: https://www.baseten.co/pricing/
- Modal pricing: https://modal.com/pricing
- Vercel AI Gateway pricing: https://vercel.com/docs/ai-gateway/pricing
- Baseten status: https://status.baseten.co/index.json
- Cerebras status: https://status.cerebras.ai/index.json
- Cloudflare AI Gateway status: https://www.cloudflarestatus.com/index.json
- Fireworks AI status: https://status.fireworks.ai/api/v2/incidents.json
- Groq status: https://groqstatus.com/api/v2/incidents.json
- Modal status: https://status.modal.com/index.json
- NanoGPT status: https://status.nano-gpt.com/index.json
- Novita AI status: https://status.novita.ai/index.json
- Portkey status: https://status.portkey.ai/index.json
- Requesty status: https://status.requesty.ai/index.json
- Segmind status: https://status.segmind.com/index.json
- Together AI status: https://status.together.ai/index.json
- Vercel AI Gateway status: https://www.vercel-status.com/index.json
More from SLOP TV News
- What Runway's Project Continuum actually is
- Everything you need to know about Runway
- Everything you need to know about Higgsfield
- Everything you need to know about Seedance 2.5
- Everything you need to know about MiniMax H3
- Everything you need to know about world models
- China's AI video industry and the subsidies behind cheap generation
- Half of Tavus's testers mistook its AI for a real person
- ElevenLabs v4 performs dialogue, and the price of that is the story
- Ex-OpenAI researcher tells Ben Horowitz: AI is smart and useless at work
- Netflix paid $587m for Ben Affleck's AI: 'you just unfreeze the weights'
- The NO FAKES Act and the right of every creator to own a face
- Newsom signs SB 1050 on synthetic performer disclosure
Sources
- openrouter.ai
- openrouter.ai
- together.ai
- fireworks.ai
- deepinfra.com
- novita.ai
- groq.com
- cerebras.ai
- replicate.com
- fal.ai
- baseten.co
- modal.com
- status.fireworks.ai
- groqstatus.com
- status.cerebras.ai
- vercel-status.com
- cloudflarestatus.com
- status.baseten.co
- status.together.ai
- status.novita.ai
- status.segmind.com
- status.requesty.ai
- status.nano-gpt.com
- status.portkey.ai
- status.modal.com