← CostPlusIQ

Qwen3.6-27B bf16 · native video · ZDR

Dedicated B200 serving · 262,144 context · 32k max output · video up to 685 s / 1,370 frames at native 2 fps · OpenAI-compatible API

Input

$0.28/M

Cached input

$0.07/M

Output

$2.00/M

Throughput

64 tok/s

TTFT (4 s clip)

0.58 s

Prefill

6.7k tok/s/stream

Throughput vs other providers of this model

Their p50s are text-dominated traffic (OpenRouter published stats); ours is measured on full-rate video with thinking enabled — the hardest case we serve.

CostPlusIQ (measured, video+thinking)64 tok/sCoreWeave80 tok/sDeepInfra74 tok/sVenice74 tok/sIo Net49 tok/sPhala37 tok/sMorph36 tok/sAlibaba33 tok/sSiliconFlow26 tok/sChutes16 tok/s

Latency

CoreWeave303 msDeepInfra369 msCostPlusIQ (4 s clip TTFT)580 msVenice639 msIo Net718.5 msMorph1510 msAlibaba1574 msSiliconFlow1612 msPhala2095.5 msChutes3144 ms

Time to first token vs clip length (measured)

inputvideo tokensTTFT
4 s clip309 tokens0.58 s
60 s clip4,391 tokens1.03 s
600 s clip44,311 tokens6.64 s

What full-rate video means

A 10-minute clip is ~38,000 tokens here (native 2 fps sampling). Every other ZDR endpoint of this model caps it at ~1,200 tokens (~30 frames) — measurably unusable for temporal tasks. We are the only endpoint combining full-rate video with zero data retention.

Metrics regenerate from committed benchmarks; market stats snapshotted 2026-08-11T06:19:04Z.