Start on infrastructure from GEKI in a European data center: tokens, models and AI agents at the click of a button. One monthly fee. GEKI runs operations.
The monthly fee covers inference, model hosting, agents, monitoring and support - no user licences, no token markup. The bill doesn't grow with every user or request. Optional: sell unused capacity to reduce the fee.
Start with no investment - GEKI provides hardware & data center.
GEKI procures, hosts and operates the GPUs and the full AI stack - you just use tokens and agents, without buying or running hardware.
Dedicated GPU capacity in a European data center, fully managed. Two hardware lines, three sizes each - a pro-rata 8 CPU cores and 96 GB RAM per GPU.
| Package | Hardware | VRAM | Price / month |
|---|---|---|---|
| Mini | 1× RTX 6000 Blackwell | 96 GB | €1,990 |
| Standard | 2× RTX 6000 Blackwell | 192 GB | €3,980 |
| Large | 8× RTX 6000 Blackwell | 768 GB | €15,920 |
| Package | Hardware | VRAM | Price / month |
|---|---|---|---|
| Mini | 1× AMD Instinct MI350P | 144 GB | €2,200 |
| Standard | 2× AMD Instinct MI350P | 288 GB | €4,400 |
| Large | 8× AMD Instinct MI350P | 1,152 GB | €17,600 |
All prices excl. VAT. 36-month minimum term. A pro-rata 8 CPU cores and 96 GB RAM per GPU. Optional: peak capacity & token resale.
For each model the right hardware, the expected token volume per month and the effective price per 1M tokens - incl. GEKI operations. Guide figures at high utilisation; actual output and cost depend on workload, context and utilisation.
| Model | Recommended hardware | Tokens / month | Effective price / 1M tokens |
|---|---|---|---|
| Gemma 4 | 1× RTX PRO 6000 | approx. 20B | approx. €0.10 |
| Qwen 3.6 | 1× RTX PRO 6000 | approx. 15B | approx. €0.13 |
| DeepSeek V4 Flash | 2× AMD MI350P | on request | on request |
| GLM 5.2 | 8× H200 | approx. 9B | approx. €0.70 |
* Guide figures on recommended hardware at high utilisation - actual token output and cost depend on workload and setup.
GEKI bills for the infrastructure used, not for token markups.
It gets cheaper when setup and workload fit: the right GPU, the right model, routing by cost and quality - and no frontier model where an open model is enough.
When your dedicated GPU capacity isn't fully used, GEKI can optionally sell unused capacity as tokens on the spot market. Your own workloads always take priority; spot-market traffic only runs on free capacity and is displaced as soon as your demand rises. The proceeds from sold capacity are shared between GEKI and the customer.
If your workloads temporarily exceed the dedicated capacity, GEKI can optionally route peaks to vetted European token providers - configurable by policy. Peak capacity is bought at market price.
Share your expected usage. GEKI recommends the right infrastructure, AI model and operating model.