BlogEnterprise

Sovereign AI Needs Sovereign Memory

Sovereign AI is more than controlling where the model runs. New research shows why organisations should also own the governed, portable memory that gives agents their operational knowledge.

Valentin TablanCo-founder & CTO · Memco6 min readEssay #03
FeaturedSovereign AIOpen WeightsResearch
SOVEREIGN CONTROL PLANE SELF-HOSTED MODEL GOVERNED EXTERNAL MEMORY INSPECT EDIT PROVENANCE PRUNE OPAQUE PROVIDER MEMORY AVOID LOCK-IN
Fig. 01 A self-hosted model and governed external memory keep both reasoning and operational knowledge inside the organisation's control.

Many organisations are choosing to run their AI on infrastructure they control, for sound business reasons: AI is becoming a business-critical part of the stack, and critical systems deserve direct control; regulators increasingly expect it, particularly in financial services and the public sector; self-hosting puts costs under your own management rather than a per-token meter; and keeping inference close to your systems keeps latency predictable. Sovereignty is a strategic position, and the organisations taking it are taking ownership of their most important new capability.

The conversation about sovereignty, though, is almost entirely about where the model runs. That covers half the stack. An agent's competence comes from two things: the model that does the thinking, and the knowledge it thinks with. Our new research, just published in Learning on the Job, shows how much that second half is worth, and why sovereignty-minded organisations should own it too.

What self-hosting actually costs

The models an organisation can realistically self-host are smaller than the largest frontier models, and the common reading is that self-hosting means accepting a capability penalty.

That reading deserves a deeper look. What a smaller model gives up is mostly breadth of world knowledge. By contrast, the cognitive machinery, the ability to reason over information, follow procedures, and use tools, remains strong in the current generation of open-weights models. And the knowledge that decides success in your deployment is your know-how and policies, your systems, your products, and the exceptions your business has accumulated over decades. No model has that knowledge, at any size. The gap between a self-hosted model and a frontier model is smaller than the gap between any model and your operation.

This changes the challenge: if the missing ingredient is knowledge rather than intelligence, the efficient solution is to supply the knowledge.

Filling the gap with learning on the job

Our study tested the most economical source of that knowledge: the feedback your operation already produces. Every task an agent runs generates signal, and corrections from routine QA describe precisely what should have been done. In our study, we paired a model with the Spark external memory: after each task, the agent distils the feedback into a lesson it can reuse the next time a similar situation arises.

For the model, we selected Mistral Large: it is an open-weights model of precisely the class a sovereignty-minded organisation would deploy. The tasks came from the banking domain of τ-bench, customer-service work in a simulated retail bank, and a hard test: even the latest frontier models, searching the bank's policy documents, succeed on only around a quarter of attempts.

Learning from corrections lifted the self-hosted model's success rate to 2.6 times its RAG baseline. Put in sovereignty terms: learning on the job closed most of the gap between the self-hosted deployment and what frontier models achieve on this benchmark out of the box. The learning did not need to change any model weights, and operates on infrastructure you would control, learning from signals you already generate.

Knowledge transfers between models

The study's second experiment speaks directly to a worry many organisations share: lock-in to a particular model. We took the memory accumulated by one model and handed it to a model from a different provider, and it worked in both directions. Knowledge captured as external memory is portable across models.

For a sovereign deployment this means the knowledge you accumulate is an appreciating asset rather than a bet on one model. When a better open-weights model ships, it inherits everything its predecessor learned. In our transfer experiment, the self-hosted Mistral Large, reading a mature memory store, reached a success rate on par with what the frontier model achieved out of the box. A smaller model you control, equipped with the right knowledge, caught up with a larger model without it.

Contrast this with fine-tuning, the traditional method for getting models to learn your business. Fine-tuning is expensive, requiring specialised skills and infrastructure. A fine-tuned model comes with significant sunk costs, creating high inertia when organisations need to adopt new upstream model versions. Spark's memory is external and not coupled to the model it was created with. It travels unchanged to the next model, and continues to accumulate learning.

Govern the knowledge layer too

This is why we believe sovereignty has a second half. If you self-host the model but your agents' accumulated knowledge lives somewhere you cannot inspect or control, the sovereignty programme is incomplete. The knowledge layer deserves the same treatment as the model: hosted on your infrastructure, governed by your rules.

Sovereign memory has concrete properties. The knowledge is stored in natural language, so your teams can read every lesson the agents have learned. Each insight carries its provenance, so you can trace where a behaviour came from. Access is governed, so knowledge flows only where you decide it should. And individual items can be corrected or removed, so the knowledge base stays under active editorial control.

Fine-tuning has a second cost beyond inertia: knowledge absorbed into weights cannot be individually inspected, attributed, or removed; it is distributed through billions of parameters. Knowledge held in a governed memory can be all three. When a supervisor asks why an agent made a decision, a governed memory gives you an answer with a paper trail: the lesson applied, where it came from, and the feedback that established its trust. For organisations working under the EU AI Act, DORA, or similar regimes, that auditability is the difference between a system you can explain and one you can only describe.

The full sovereign stack

Spark is the memory layer for this stack: a shared, governed memory that sits outside the model, connects over MCP, and turns operational feedback into knowledge your agents keep using. It runs where you need it to run, and everything it holds is yours to read, audit, and carry to the next model you choose.

The full study, with the protocol, the statistics, and the released data, is at arxiv.org/abs/2607.22157. If your organisation is building a sovereign AI capability, the model choice is well served by the current open-weights ecosystem. The knowledge layer deserves the same deliberate ownership, and the returns arrive quickly, as shown in our study.


Valentin Tablan

Co-founder & CTO · Memco

Former Lead Scientist for Amazon Alexa, with 20+ years at the cutting edge of natural-language and knowledge-based AI. Chief AI Officer at Ieso Digital Health, where he created Velora — the world's first clinically validated generative AI therapy agent, with outcomes on par with human-delivered care.

Follow on X →

Running coding agents on real repos?

Try Memco free

the loop

benchmarks · product · research

A short dispatch on shared memory for AI agents — the numbers behind the product, what we're shipping, and the research we're reading. No filler.

Unsubscribe anytime · no spam