Moving from a hyperscaler API to European GPUs, hosted and managed by GEKI. An open model behind an OpenAI-compatible API; unused capacity is sold at night.
1.5M documents/month, €20,000/month for a US hyperscaler API. Goal: predictable cost and less dependence on a single provider.
An open model on GPUs GEKI provides and operates in the EU. OpenAI-compatible API, no hardware purchase. The hyperscaler stays for peaks and fallback.
Net AI cost from €20,000 to ~€6,000/month. Unused capacity is sold at night; tokens run in the EU.
A SaaS platform for document processing pays €20,000 per month to a US hyperscaler. Goal: predictable cost, European tokens, without a GPU ops team.
GEKI moved the workload to dedicated GPUs in an EU data center — infrastructure and operations by GEKI, an open model, an OpenAI-compatible API. The application code stayed unchanged.
The platform processes around 1.5 million documents per month: summarising support conversations, extracting structured data. Lots of input, JSON output — all through a hyperscaler AI API.
| Current workload | Value |
|---|---|
| Documents / month | 1,500,000 |
| Avg. input / output | 2,500 / 500 tokens |
| Tokens per month (in / out) | 4B / 800M |
| Monthly AI cost (hyperscaler) | €20,000 |
1. Validate the model. GEKI measured open models against the platform’s real data. DeepSeek V4 Flash was at or above production level on summarisation, recall, hallucinations and JSON conformance — and was chosen.
2. Provide & operate the infrastructure. Instead of buying GPUs, GEKI provides and operates the capacity in an EU data center. No hardware purchase, no infrastructure team. GEKI handles deployment, monitoring, security and upgrades.
3. Keep the code. The model runs behind an OpenAI-compatible API. The application stays; only the endpoint changes.
The workload fits on the dedicated capacity. The managed-service fee (GEKI infrastructure + operations) replaces the variable API bill:
| Before (hyperscaler API) | After (GEKI managed) | |
|---|---|---|
| Monthly AI cost | €20,000 | ~€6,300 |
| Idle capacity resale (at night) | — | − revenue credit |
| Monthly net cost | €20,000 | ~€6,000 |
Document processing runs mostly during the day; at night the GPUs would sit idle. GEKI sells unused capacity as tokens on the spot market. The customer workload has priority; external traffic only on free capacity. That brings the hardware to nearly 100% utilisation, the proceeds reduce the service fee — net saving ~70%.
The hyperscaler API stays as a peak and fallback path: for load peaks above the dedicated capacity or during maintenance. It is no longer the main supplier — the provider mix is broader, not just swapped.
See pricing or briefly describe what you want to run — we'll work it out.