GEKI.AI
← All posts Blog · Open Models

Strong AI agents on two GPUs: DeepSeek V4 Flash makes sovereign AI affordable

DeepSeek V4 Flash: strong agents via the API, open and runnable on two GPUs. Why that makes sovereign AI in Europe even more attractive now.

Matthias Allitsch-Wutte

Matthias Allitsch-Wutte · GEKI founder

31 July 2026 · 4 min read

A capable model that runs on two GPUs and is open: until recently, that didn’t exist. Sovereign AI is slowly arriving, driven by developments in open source, first Kimi K3, now DeepSeek V4 Flash.

On 31 July 2026, DeepSeek significantly improved the API version of V4 Flash. In parallel, the open V4 Flash architecture can already run on your own hardware, under an MIT license. The model is built for agents — AI that carries out tasks on its own rather than just chatting. The new benchmark figures apply to the DeepSeek API (V4-Flash-0731); the open weights are still the preview — same architecture, the re-post-training not yet released as weights.

What’s special is the combination: strong performance on agent tasks, open weights, and hardware a company can control itself. The open question is how far the new figures, measured via the API, reproduce on your own hardware.

How strong is it?

A good yardstick is GPT-5.6 Terra, OpenAI’s strong mid-range model for cost-conscious enterprise use. The two benchmarks measure how well a model solves real tasks — working in the terminal or operating tools, for example. Higher is better.

Vendor figures (31 July 2026, DeepSeek API)DeepSeek V4 FlashGPT-5.6 Terra
Terminal Bench 2.182.778.4
Toolathlon Verified70.353.1

Source: DeepSeek release data (not independently reproduced).

In these vendor test runs, V4 Flash sits above a strong proprietary mid-range model on agent tasks, at a fraction of the cost. For many practical purposes it is genuinely comparable — Terminal Bench makes that plain. (Even on Hacker News, where model releases are usually met with a shrug, one comment read: “This is more exciting than K3.”)

The specs, briefly

From the model itself (per DeepSeek):

  • 284B parameters, 13B active (Mixture-of-Experts)
  • Up to 1M token context
  • Open weights, MIT license

On our own infrastructure. All GEKI figures in this article, throughput and prices alike, are measured on GEKI’s own GPUs, in a European data center and with the GEKI software stack. We already run the preview heavily in our own agents and workflows — not just as a demo.

  • Our entry configuration is 2× NVIDIA RTX PRO 6000 with 192 GB of GPU memory combined.
  • In our own first measurements we reach over 200 tokens per second. We share the exact test conditions (quantization, inference engine, context length, single-stream vs. aggregate, parallel requests) on request.
  • As an alternative, we’re evaluating configurations based on the AMD MI350P, which can be attractive on memory bandwidth and total cost.

Why the economics decide it

A strong open model is half the story. The other half is the price. Here are the effective prices per one million tokens, mixed at 8 input : 1 output — a common split for agent work:

Effective / 1M tokens (8 in : 1 out)GPT-5.6 Terra (US hyperscaler, example)GEKI Managed (DS V4 Flash)
Blended~$3.58approx. €0.20 (≈ $0.23)

Terra from API list prices (input ~$2.30 / output ~$13.80), converted to 8:1. GEKI: guide figure at high utilisation on recommended hardware, see pricing. All USD comparisons at the 31 July 2026 rate (€1 = $1.15).

Here’s how to read the table: GEKI does not sell tokens individually — you get an instance including operations. The effective price per million tokens is a guide figure; depending on utilisation it is often lower still. Against Terra it sits roughly an order of magnitude below. At the volumes automation consumes, that alone decides the business case.

Where the model really shines

These benchmarks don’t measure theoretical ability. They measure real tasks: building software, using the terminal, calling APIs, driving tools, automating security workflows.

Translated: agents that solve tasks reliably, at a cost where automation adds up even for mid-sized companies. Examples from practice:

  • document processing
  • sorting and answering tickets
  • making company wikis searchable and answerable
  • research and analysis
  • coding and software automation

What sovereignty really means here

Sovereignty doesn’t come from where a model was made. It comes from control: over data, infrastructure, model versions, access and day-to-day operation.

Concretely, that means:

  • control over data and where it is stored
  • free choice of operator
  • no forced dependence on a single model API
  • controllable access and permission layers
  • reproducible model versions
  • observability and auditability
  • the option to switch models

An open model like V4 Flash is the prerequisite for this. But it only becomes truly operable through an operating layer that delivers exactly this control. That’s where GEKI comes in, not as plain on-premise hosting, but as managed, controllable operation.

Further reading

Sources

  • DeepSeek API Updates: V4-Flash-0731 (31 July 2026), plus the open weights from the original V4 release (model specs, MIT license).
  • Benchmark data from the DeepSeek release; discussion e.g. in the Hacker News thread.
Sovereign AI, today

DeepSeek V4 Flash, managed on your hardware

GEKI's entry configuration on 2× RTX PRO 6000, operated with SLAs in the right data center.