++++
UNIFIED INFERENCE GATEWAY

Inference Infrastructure
for Cost-Efficient AI

A unified inference gateway built for individual developers, startups, and growing teams — the same frontier model endpoints at one transparent rate, with nothing trained on your prompts.

No credit card required
PRICING · PER 1M TOKENS

Every model, a fraction of the price

FerrixOpenRouter list price
26 of 26 models
Anthropic
claude-fable-5
Anthropic
INPUT / 1M−80%
$2.00
$10.00
OUTPUT / 1M−80%
$10.00
$50.00
CACHE READ / 1M−80%
$0.20
$1.00
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$2.50
$12.50
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$4.00
$20.00
VS OPENROUTER−80%
Anthropic
claude-haiku-4-5
Anthropic
INPUT / 1M−80%
$0.20
$1.00
OUTPUT / 1M−80%
$1.00
$5.00
CACHE READ / 1M−80%
$0.02
$0.10
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$0.25
$1.25
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$0.40
$2.00
VS OPENROUTER−80%
Anthropic
claude-opus-4-6
Anthropic
INPUT / 1M−80%
$1.00
$5.00
OUTPUT / 1M−80%
$5.00
$25.00
CACHE READ / 1M−80%
$0.10
$0.50
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$1.25
$6.25
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$2.00
$10.00
VS OPENROUTER−80%
Anthropic
claude-opus-4-7
Anthropic
INPUT / 1M−80%
$1.00
$5.00
OUTPUT / 1M−80%
$5.00
$25.00
CACHE READ / 1M−80%
$0.10
$0.50
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$1.25
$6.25
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$2.00
$10.00
VS OPENROUTER−80%
Anthropic
claude-opus-4-8
Anthropic
INPUT / 1M−80%
$1.00
$5.00
OUTPUT / 1M−80%
$5.00
$25.00
CACHE READ / 1M−80%
$0.10
$0.50
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$1.25
$6.25
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$2.00
$10.00
VS OPENROUTER−80%
Anthropic
claude-sonnet-4-6
Anthropic
INPUT / 1M−80%
$0.60
$3.00
OUTPUT / 1M−80%
$3.00
$15.00
CACHE READ / 1M−80%
$0.06
$0.30
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$0.75
$3.75
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$1.20
$6.00
VS OPENROUTER−80%
Anthropic
claude-sonnet-5
Anthropic
INPUT / 1M−80%
$0.40
$2.00
OUTPUT / 1M−80%
$2.00
$10.00
CACHE READ / 1M−80%
$0.04
$0.20
CACHE WRITE (5-MINUTE CACHE) / 1M−80%
$0.50
$2.50
CACHE WRITE (1-HOUR CACHE) / 1M−80%
$0.80
$4.00
VS OPENROUTER−80%
DeepSeek
deepseek-v3.2
DeepSeek
INPUT / 1M−64%
$0.0957
$0.269
OUTPUT / 1M−62%
$0.15
$0.40

Cache writes: not charged

VS OPENROUTER−62%
Google
gemini-2.5-flash
Google
INPUT / 1M−80%
$0.06
$0.30
OUTPUT / 1M−80%
$0.50
$2.50
CACHE READ / 1M−80%
$0.006
$0.03
VS OPENROUTER−80%
Google
gemini-2.5-pro
Google
INPUT / 1M−80%
$0.25
$1.25
OUTPUT / 1M−80%
$2.00
$10.00
CACHE READ / 1M−80%
$0.025
$0.125
VS OPENROUTER−80%
Google
gemini-3.1-pro
Google
INPUT / 1M−80%
$0.40
$2.00
OUTPUT / 1M−80%
$2.40
$12.00
CACHE READ / 1M−80%
$0.04
$0.20
VS OPENROUTER−80%
Google
gemini-3.5-flash
Google
INPUT / 1M−80%
$0.30
$1.50
OUTPUT / 1M−80%
$1.80
$9.00
CACHE READ / 1M−80%
$0.03
$0.15
VS OPENROUTER−80%
Page 1 of 3
Average saving across shown models: −77% · No minimums · No platform fee · Volume discounts above $2k/mo
WHAT PEOPLE SAY

Switched for the price. Stayed for everything else.

[ PRICE ]

"Same GPT-4o and Claude endpoints we were already calling, cut roughly in half. It paid for the migration in the first week."

DR
Dani Rowe
Staff Engineer
[ RELIABILITY ]

"We've never seen a blip — even during our biggest traffic spikes, latency stays flat and requests just go through. It's the most boring part of our stack, which is exactly what you want."

MK
Marcus Kang
Head of Platform
[ SWITCHING ]

"One base URL change and our SDK just worked. No rewrites, no new client — we were live in an afternoon."

PS
Priya Sharma
CTO
DROP-IN MIGRATION

Change one line.
Keep your stack.

Point your OpenAI-compatible client at Ferrix's base URL. Same request shape, same SDKs, same tooling — every model behind one endpoint.

WORKS WITH
openai · anthropic · langchain · llamaindex · vercel ai sdk
CLIENT.PY
# before
client = OpenAI(
base_url="https://api.openai.com/v1",
api_key=KEY,
)
# after — that's the whole migration
client = OpenAI(
base_url="https://api.ferrix.sh/v1",
api_key=KEY,
)
$ curl api.ferrix.sh/v1/models
→ 142 models · 99.99% uptime
++++
SPEND CONTROLS

Spend controls that work before the invoice

[ CAP ]
Budget caps
Hard monthly ceilings per org, project, or key — requests stop before you overspend.
[ ALERT ]
Spend alerts
Real-time notifications at the thresholds you set, before the bill surprises you.
[ USAGE ]
Usage analytics
Per-model, per-key, per-project cost breakdowns, exportable anytime.
[ PRIVATE ]
No data training
We never train on, or retain, your prompts or completions. Ever.
RELIABILITY
99.99%
UPTIME, SLA-BACKED
180ms
P50 TIME-TO-FIRST-TOKEN
SOC 2
TYPE II · ZERO PROMPT RETENTION

Change one line tonight. Watch the bill drop.

Start with $5 in free credit — no card, no minimums.

Ferrix
© 2026TermsPrivacySLAtokens routed today: 2,147,483,001
++