GEKI.AI
← All case studies Case Study · AI document processing

A document-processing SaaS cuts its AI bill by ~70%.

Moving from a hyperscaler API to European GPUs, hosted and managed by GEKI. An open model behind an OpenAI-compatible API; unused capacity is sold at night.

~70%
Lower net AI cost
EU
Tokens & data residency
~100%
Hardware utilisation
0
Own GPU ops team

The challenge

1.5M documents/month, €20,000/month for a US hyperscaler API. Goal: predictable cost and less dependence on a single provider.

The GEKI approach

An open model on GPUs GEKI provides and operates in the EU. OpenAI-compatible API, no hardware purchase. The hyperscaler stays for peaks and fallback.

The outcome

Net AI cost from €20,000 to ~€6,000/month. Unused capacity is sold at night; tokens run in the EU.

Less dependence, a lower bill

A SaaS platform for document processing pays €20,000 per month to a US hyperscaler. Goal: predictable cost, European tokens, without a GPU ops team.

GEKI moved the workload to dedicated GPUs in an EU data center — infrastructure and operations by GEKI, an open model, an OpenAI-compatible API. The application code stayed unchanged.

High volume, one API bill

The platform processes around 1.5 million documents per month: summarising support conversations, extracting structured data. Lots of input, JSON output — all through a hyperscaler AI API.

Why switch

  • Dependence: one US provider for models, capacity and pricing
  • Data: customer data and inference should stay in the EU
  • Cost: the API bill grows with every customer and squeezes the margin
Current workloadValue
Documents / month1,500,000
Avg. input / output2,500 / 500 tokens
Tokens per month (in / out)4B / 800M
Monthly AI cost (hyperscaler)€20,000

Open model, European GPUs, managed by GEKI

1. Validate the model. GEKI measured open models against the platform’s real data. DeepSeek V4 Flash was at or above production level on summarisation, recall, hallucinations and JSON conformance — and was chosen.

2. Provide & operate the infrastructure. Instead of buying GPUs, GEKI provides and operates the capacity in an EU data center. No hardware purchase, no infrastructure team. GEKI handles deployment, monitoring, security and upgrades.

3. Keep the code. The model runs behind an OpenAI-compatible API. The application stays; only the endpoint changes.

  • Tokens on EU infrastructure
  • An open model, no single provider as a bottleneck
  • OpenAI-compatible API — the endpoint is swappable
  • Operated by GEKI — no GPU team at the customer

Cost before and after

The workload fits on the dedicated capacity. The managed-service fee (GEKI infrastructure + operations) replaces the variable API bill:

Before (hyperscaler API)After (GEKI managed)
Monthly AI cost€20,000~€6,300
Idle capacity resale (at night)− revenue credit
Monthly net cost€20,000~€6,000

Unused capacity sold at night

Document processing runs mostly during the day; at night the GPUs would sit idle. GEKI sells unused capacity as tokens on the spot market. The customer workload has priority; external traffic only on free capacity. That brings the hardware to nearly 100% utilisation, the proceeds reduce the service fee — net saving ~70%.

The hyperscaler stays for peaks & fallback

The hyperscaler API stays as a peak and fallback path: for load peaks above the dedicated capacity or during maintenance. It is no longer the main supplier — the provider mix is broader, not just swapped.

What changed

  • Net AI cost: €20,000 → ~€6,000/month (~70% less)
  • Quality held with a validated open model
  • Tokens in the EU, data residency in the EU
  • Hyperscaler only for peaks & fallback
  • Unused capacity sold; hardware near ~100% utilisation
  • No GPU ops team at the customer
Your workload

Your setup next?

See pricing or briefly describe what you want to run — we'll work it out.