All systems operational
Live status and speed of the Umans Code gateway and its models, refreshed every 30 seconds. For each model we show its output speed and median time to first token.
API endpoint
Models in production
tok/s = output tokens per second · TTFT = time to first token · p50 = median over the last 5 minutes. On each gauge the midpoint is that model's target; the marker sits further right when it's beating target (faster TTFT, higher throughput).
Operational48.5tok/sthroughput · p50 · last 5 min1.95sTTFT · p50 · last 5 min99.93%uptime · 24h
GLM 5.2 is our best model for coding right now, with a 400K context window for large codebases. Vision is available on the Anthropic Messages API (`/v1/messages`) only, through a server-side handoff (GLM 5.2 generates the text, Kimi preprocesses the image); that handoff will be retired soon in favour of more efficient client-side image handling.
90-day speed trends & events →Operational87.4tok/sthroughput · p50 · last 5 min2.34sTTFT · p50 · last 5 min99.38%uptime · 24h
Kimi K2.7-Code via Umans Code - Moonshot's strongest coding model and the successor to Kimi K2.6. Built for complex, tool-heavy agentic coding; it reasons more efficiently than K2.6, so agent sessions run faster at the same depth.
90-day speed trends & events →Umans Flash Fastestalso served as umans-qwen3.6-35b-a3bOperational293.6tok/sthroughput · p50 · last 5 min1.89sTTFT · p50 · last 5 min99.01%uptime · 24h
Our fastest model: a light workflow complement, not a standalone coder. Think Haiku next to Opus: not everything needs a frontier model, and Flash's speed (200+ tokens per second) compounds on the roles around umans-coder: gathering context, scout subagents, research, summaries, documentation, and quick edits.
90-day speed trends & events →In the playground
Umans Kimi K3 Experimentalprerelease, seat-gated via LabsIn testing55.9tok/sthroughput · p50 · last 5 min1.45sTTFT · p50 · last 5 min82.22%in testing
Kimi K3 in prerelease: Moonshot's largest open-weight release, a 2.8T-parameter mixture-of-experts with a 1M-token context window and native vision, built for repository-scale code understanding and multi-step agentic work. It thinks by default at maximum reasoning effort; select none, low, high, or max to trade depth for speed. While we scale capacity, access is seat-gated through the Labs page, and availability is limited: expect occasional errors during the ramp. For production work today we recommend umans-coder or umans-glm-5.2.
90-day speed trends & events →