Most inference engines publish tok/s only. HELIX adds a trust surface: a 0–100% confidence score at under 1% compute (near-zero added cost on GPU). Serving numbers below are from the official vLLM GuideLLM harness on AMD EPYC 9254 — 640/640 requests, zero errors M.
HELIX runs across GPU, CPU, edge, and IoT. These CPU serving rows are one measured lane — not the whole product identity. Prefer residual risk at a review budget when quoting trust results.
| System | Conc. | TTFT p50 | TTFT p95 | Valid JSON | Tokens / req |
|---|---|---|---|---|---|
| HELIX v1.7 #58 | c=1 | 940 ms | 1,089 ms | 100% | ~85 complete |
| HELIX v1.7 #58 | c=2 | 1,687 ms | 1,946 ms | 100% | ~87 complete |
| HELIX v1.7 #58 | c=4 | 2,469 ms | 3,053 ms | 100% | ~87 complete |
| Stock Ollama | c=1 | 8,416 ms | 10,622 ms | ~84% | ~16 fragment |
| Stock Ollama | c=2 | 16,262 ms | 17,088 ms | ~84% | ~16 fragment |
| Stock Ollama | c=4 | 31,816 ms | 34,531 ms | ~84% | ~16 fragment |
Ollama returns ~16-token fragments at 32K (full Mamba2 SSM state rebuild); HELIX returns complete schema-valid extractions. M
| Conc. | HELIX | Ollama | Advantage |
|---|---|---|---|
| c=1 | 940 ms | 8,416 ms | 8.95× |
| c=2 | 1,687 ms | 16,262 ms | 9.64× |
| c=4 | 2,469 ms | 31,816 ms | 12.89× |
Book a call for the full technical report and a sovereign HELIX pod on your hardware.