| model | context | native video † | ZDR § $ per M — in / cached-in / out |
standard § $ per M — in / cached-in / out |
flex ‖ $ per M — in / cached-in / out |
|---|---|---|---|---|---|
| Qwen3.8-27B fp8 | 262,144 | ✓ full-rate, 780 s | $0.40 / $0.05 / $2.75 | $0.30 / $0.05 / $2.0625 | ZDR $0.28 / $0.035 / $1.925 standard $0.21 / $0.035 / $1.44375 |
| Qwen3.8-Flash-Next-NVFP4 nvfp4 | 262,144 | ✓ full-rate, 780 s | $0.15 / $0.016 / $0.47 | $0.1125 / $0.016 / $0.3525 | ZDR $0.105 / $0.0112 / $0.329 standard $0.07875 / $0.0112 / $0.24675 |
| Kimi-K3-NVFP4 fp4 | 262,144 | ✓ native 8 fps, 780 s | $0.88 / $0.33 / $10.53 | $0.66 / $0.33 / $7.8975 | ZDR $0.616 / $0.231 / $7.371 standard $0.462 / $0.231 / $5.52825 |
† Most other endpoints of this model cap how much of a video they ingest and do not tell you: measured, a ten-minute clip reaches them as 3–10% of its frames, returned as a normal 200. We tokenize up to 1,560 frames (780 s @ 2hz), so prompt tokens grow with clip duration — see the measured plot, where their curves flatten and ours does not.
§ Zero-Data-Retention (ZDR): prompts and media are processed in memory and never stored or trained on.
‖ Flex: send "service_tier": "flex" and the request is served only from
capacity standard requests are not using, at these rates; when there is none it is refused immediately and not billed.
How flex works.