DeepSeek V4 Flash: strong agents via the API, open and runnable on two GPUs. Why that makes sovereign AI in Europe even more attractive now.
Matthias Allitsch-Wutte · GEKI founder
31 July 2026 · 4 min read
A capable model that runs on two GPUs and is open: until recently, that didn’t exist. Sovereign AI is slowly arriving, driven by developments in open source, first Kimi K3, now DeepSeek V4 Flash.
On 31 July 2026, DeepSeek significantly improved the API version of V4 Flash. In parallel, the open V4 Flash architecture can already run on your own hardware, under an MIT license. The model is built for agents — AI that carries out tasks on its own rather than just chatting. The new benchmark figures apply to the DeepSeek API (V4-Flash-0731); the open weights are still the preview — same architecture, the re-post-training not yet released as weights.
What’s special is the combination: strong performance on agent tasks, open weights, and hardware a company can control itself. The open question is how far the new figures, measured via the API, reproduce on your own hardware.
A good yardstick is GPT-5.6 Terra, OpenAI’s strong mid-range model for cost-conscious enterprise use. The two benchmarks measure how well a model solves real tasks — working in the terminal or operating tools, for example. Higher is better.
| Vendor figures (31 July 2026, DeepSeek API) | DeepSeek V4 Flash | GPT-5.6 Terra |
|---|---|---|
| Terminal Bench 2.1 | 82.7 | 78.4 |
| Toolathlon Verified | 70.3 | 53.1 |
Source: DeepSeek release data (not independently reproduced).
In these vendor test runs, V4 Flash sits above a strong proprietary mid-range model on agent tasks, at a fraction of the cost. For many practical purposes it is genuinely comparable — Terminal Bench makes that plain. (Even on Hacker News, where model releases are usually met with a shrug, one comment read: “This is more exciting than K3.”)
From the model itself (per DeepSeek):
On our own infrastructure. All GEKI figures in this article, throughput and prices alike, are measured on GEKI’s own GPUs, in a European data center and with the GEKI software stack. We already run the preview heavily in our own agents and workflows — not just as a demo.
A strong open model is half the story. The other half is the price. Here are the effective prices per one million tokens, mixed at 8 input : 1 output — a common split for agent work:
| Effective / 1M tokens (8 in : 1 out) | GPT-5.6 Terra (US hyperscaler, example) | GEKI Managed (DS V4 Flash) |
|---|---|---|
| Blended | ~$3.58 | approx. €0.20 (≈ $0.23) |
Terra from API list prices (input ~$2.30 / output ~$13.80), converted to 8:1. GEKI: guide figure at high utilisation on recommended hardware, see pricing. All USD comparisons at the 31 July 2026 rate (€1 = $1.15).
Here’s how to read the table: GEKI does not sell tokens individually — you get an instance including operations. The effective price per million tokens is a guide figure; depending on utilisation it is often lower still. Against Terra it sits roughly an order of magnitude below. At the volumes automation consumes, that alone decides the business case.
These benchmarks don’t measure theoretical ability. They measure real tasks: building software, using the terminal, calling APIs, driving tools, automating security workflows.
Translated: agents that solve tasks reliably, at a cost where automation adds up even for mid-sized companies. Examples from practice:
Sovereignty doesn’t come from where a model was made. It comes from control: over data, infrastructure, model versions, access and day-to-day operation.
Concretely, that means:
An open model like V4 Flash is the prerequisite for this. But it only becomes truly operable through an operating layer that delivers exactly this control. That’s where GEKI comes in, not as plain on-premise hosting, but as managed, controllable operation.
GEKI's entry configuration on 2× RTX PRO 6000, operated with SLAs in the right data center.