An inference performance lab
Frontier open models, served at speed.
The best open-weight models on dedicated Blackwell capacity, engineered for very high throughput, steady uptime, and a fair price per token. Frontier models, your own checkpoints, and the performance work to make them fast.
The lineup
Models we serve today
GLM-5.2
Frontier coding and long-horizon agentic workflows. Project-scale reasoning that holds context across a full build.
DSV4 Pro
Frontier general model for hard reasoning, tool use, and agent stacks that cannot afford wrong answers.
Nemotron Ultra
Deep multi-step reasoning at full weight. The slow-and-right lane for verification, review, and planning.
Nemotron Lightning
The low-latency lane on B200: interactive agents, high-QPS serving, and realtime products.
Qwen 3.8 27B
Fast, capable generalist on B200. The price-performance default for extraction, routing, and bulk work.
Kimi K2.6
Long-horizon coding, visual interface generation, and agent swarms that run a task end to end.
Kimi K2.7 Code
Coding-focused agentic model. Stronger end-to-end task completion with ~30% fewer thinking tokens than K2.6.
Why Numinous
A lab, not a reseller
Fast and standard lanes on Blackwell. Speculative decoding and continuous batching push the highest tokens per second in production.
Dedicated capacity, graceful failover, and a base that never cold starts. Uptime you can route to.
Dedicated hardware run well costs less than marketplaces run poorly. Fair per-token pricing, published, no gates.
Bring your own fine-tunes and LoRAs. The same performance work applies to your weights, not just ours.
Run your workload on it
Access is granted in batches while we scale capacity. Tell us your models and your throughput.
Request access