Safi صافي Free savings estimate
LLM cost layer · built for MENA

The cost layer between your app and your LLM.

Safi sits in front of your LLM calls and cuts what you pay — automatically, on every request that's safe to touch, with extra care for Arabic. Swap one API endpoint. Your models, prompts, and answers don't change.

Get your free savings estimate See how it works
No code rewrite Works with your current provider Pay only when we save you money
What actually changes
Your app LLM provider

Today: every request, full price, straight through.

Your app Safi LLM provider

With Safi: routine requests cost less, automatically. Anything that shouldn't be touched goes through exactly as before.

Integration One endpoint change
Answer quality Checked on every change
30–60%
typical LLM overspend for MENA AI products, before any optimization
3–6×
what one Arabic word costs vs. one English word, on every major model
$0
cost to find out your own number, before you change anything
What we do

We work out what's safe to change, then change it.

Nobody overspends on purpose — most traffic just never gets sorted or double-checked before it's sent, and it all pays top price by default. Safi profiles your traffic, decides — request by request — what can be made cheaper without changing the answer, and applies exactly that. Nothing more.

Priced right for the request

Routine traffic costs less. Anything that needs the best you have still gets it, no exceptions.

Nothing personal ever gets shortcut

If an answer depends on who's asking, it goes through in full, every time — no near-enough answers on anything that's actually yours.

Arabic doesn't pay extra by accident

Arabic-heavy traffic gets handled with that tax in mind, on top of everything else we do.

Every change ships behind a quality gate. If it would change the answer, we don't apply it. How we make that call is not something we publish — what matters to you is that it's checked before it ships, and reversible if you disagree with it.

How we test it

We built a blind optimizer to prove ours isn't one.

Before an optimization touches real traffic, we run it against two alternatives on a controlled internal benchmark: doing nothing, and a "blind" optimizer that applies the same trick to every request regardless of risk — which is what most generic gateways do.

The blind optimizer lost on both counts: it saved less, and a quarter of its answers came back wrong and had to be redone at full price. Knowing where not to cut is what makes the rest safe to cut.

Path Cost vs. baseline Wrong answers
Do nothing baseline 0 of 40
Blind optimizer −55% 10 of 40
Safi −81% 0 of 40
Internal benchmark: 40 seeded test requests across 4 workloads, synthetic data seeded with waste and the mistakes a careless tool would make — not customer traffic. We also run a separate 100-request suite across 10 industries; Safi scores zero wrong answers there too, every time we run it. Your own number comes from your own traffic, measured free during the audit below.
Who it's for

Any MENA AI team paying an LLM bill.

If you have a product that calls an LLM API and the bill is real and growing, this is for you — English product, Arabic product, or both.

Consumer · e-commerce · telco — Arabic/bilingual support chat B2B SaaS — helpdesk / in-product assistant Sales tech · CRM — qualification, outreach, call summaries Ops · vertical SaaS — agents & workflow automation Fintech · legal · logistics — KYC, invoices, contracts Marketing · media — content & Arabic localization Any team with internal docs — knowledge assistant Ed-tech — tutoring, grading, content
Not every workload has a safe saving — on two we've tested, clinical triage and code review, the honest plan is to leave almost everything untouched, and we say so rather than force a number.
And one thing more

Arabic is where the waste is worse.

Everything above works in any language. Arabic gets extra treatment on top, because it's priced worse to begin with — and the exact multiple shifts depending on the model, so we track it rather than assume it.

خليجي · مصري · شامي · فصحى
One measured example
3.59 tokens
per Arabic word, one tokenizer family
Same word, other family
1.51 tokens
per Arabic word, the other family
2.4× the spread we measured on this one pair — inside the wider 3–6× range across all models
Cheap to answer
ما هي ساعات العمل؟
Asked constantly. One answer serves everyone.
Never cut corners here
ما هو رصيد حسابي رقم ١٠٠٧؟
One customer's balance. A near-enough answer is somebody else's money.
Illustrative examples, not live customer queries. A blind tool can't tell these two apart — telling them apart is the whole product.
How it works

You don't have to become a cost engineer.

That's our job. You send a sample of traffic, we come back with a number and the risk attached to it. If you like the number, we run it.

Your models, prompts, and provider contracts stay exactly as they are.
Switch us off and traffic goes straight back to your provider. Nothing of ours is left behind.
01
Send a sample

A slice of your traffic. Read-only, no integration, nothing switched on.

02
We do the hard part

We profile the workload and build a plan sized to it — including the cuts we refuse to make on your behalf. You get the number and the risk, not a homework assignment.

03
Approve, then save

One endpoint change and it's live for your traffic. Quality is checked continuously; anything that fails the gate is reversed.

Pricing

If you don't save, you don't pay.

You keep paying your provider directly. We invoice a share of the savings we can prove — agreed with you before anything goes live, and sized to your volume and workload.

See your number first
Setup & integration$0
Monthly platform fee$0
Seats & minimumsNone
Our feeA share of proven savings
A month with no savingsNo invoice

Every invoice arrives with the usage it was calculated from, so you can re-derive the number yourself.

Quality is part of the deal

Savings attached to a response that fails your quality gate are struck off the bill. We don't get paid for a worse answer.

Leave whenever

No term, no notice period, no exit fee. The audit alone is worth having.

FAQ
Are the numbers on this page from real customer traffic?

The benchmark table above is not — it's a synthetic, seeded test corpus we built to prove the mechanism, not a savings promise for any specific customer. The market figures (30–60% overspend, the Arabic tax) come from our own research across MENA AI products. Your own number is measured on your own traffic, for free, before anything changes.

Will this change what my users see?

No — nothing you haven't approved. You see the projected saving and the risk on each change before it goes live, and can reverse any of it.

Do my logs leave my infrastructure?

They don't have to. Analysis can run where your data already lives, with sensitive fields removed before anything is read.

Which providers do you work with?

The major international providers, the big clouds, self-hosted models — and we're building specific support for the Arabic-native models most tools ignore.

Find out what you'd save. It costs nothing.

Send a traffic sample, get your number, the risk behind it, and a straight answer on whether we're worth it. Private beta — a few design partners at a time.

No sales sequence. One reply from a founder, with your number in it.