An inference performance lab

Frontier open models, served at speed.

The best open-weight models on dedicated Blackwell capacity, engineered for very high throughput, steady uptime, and a fair price per token. Frontier models, your own checkpoints, and the performance work to make them fast.

Request accessView models

The lineup

Models we serve today

GLM-5.2

Frontier coding and long-horizon agentic workflows. Project-scale reasoning that holds context across a full build.

744B MoE40B active1M contextFP8

DSV4 Pro

Frontier general model for hard reasoning, tool use, and agent stacks that cannot afford wrong answers.

flagship lane

Nemotron Ultra

Deep multi-step reasoning at full weight. The slow-and-right lane for verification, review, and planning.

reasoning lane

Nemotron Lightning

The low-latency lane on B200: interactive agents, high-QPS serving, and realtime products.

B200low latency

Qwen 3.8 27B

Fast, capable generalist on B200. The price-performance default for extraction, routing, and bulk work.

27BB200

Kimi K2.6

Long-horizon coding, visual interface generation, and agent swarms that run a task end to end.

~1T MoE32B active256K contextFP8

Kimi K2.7 Code

Coding-focused agentic model. Stronger end-to-end task completion with ~30% fewer thinking tokens than K2.6.

~1T MoE32B active256K contextINT4

Why Numinous

A lab, not a reseller

01 · Speed

Fast and standard lanes on Blackwell. Speculative decoding and continuous batching push the highest tokens per second in production.

02 · Reliability

Dedicated capacity, graceful failover, and a base that never cold starts. Uptime you can route to.

03 · Price

Dedicated hardware run well costs less than marketplaces run poorly. Fair per-token pricing, published, no gates.

04 · Your checkpoints

Bring your own fine-tunes and LoRAs. The same performance work applies to your weights, not just ours.

Run your workload on it

Access is granted in batches while we scale capacity. Tell us your models and your throughput.

Request access