THE MCP SCHEMA BOTTLENECK, REMOVED
Fewer schemas.
Faster agents.
An OpenAI-compatible gateway that retrieves the next dependency-ready MCP tool and turns stable tool bundles into reusable llama.cpp prefixes.
01Same model. Same request.
Same model. Same request.
61 schemas versus one.
“Find the launch PDF, create a share link, and email it to Mina.”
- Prompt
- 4,903 tokens
- p95 TTFT
- 717 ms
- Throughput
- 1.42 req/s
→
- Prompt
- 207 tokens
- p95 TTFT
- 100 ms
- Throughput
- 12.50 req/s
02Every workload clears
Every workload clears
the 3× gate.
03Speed without
Speed without
throwing quality away.
complete held-out workflows
complete held-out workflows
0% wrong first calls
04Public, reproducible
Public, reproducible
Arm evidence.
- CPU4-core Arm Neoverse-N2
- Runtimellama.cpp b9623 + KleidiAI
- ModelQwen2.5 1.5B Q4_K_M
- Protocolalternating order · 2 warmups · 7 runs