Frozen-weights agents, external memory, one-bit verdicts and corrections. Cross-model transfer.
arXiv:2607.22157 · Jul 2026
Research · journal · field reports
What we’re learning as we build shared memory for AI agents: what to store, what compounds, and where the edge of today’s tools really is.
Agents paired with a natural-language memory improve without a single weight update — learning from outcome verdicts alone, and more still from corrections. Measured on Mistral Large, replicated on Claude Sonnet 5, with cross-model memory transfer. Harness, protocol and data released.
1.6×
Single-trial success vs static RAG
learning from outcome verdicts
2.6×
When learning from corrections
same weights · both arms
22 of 84
Never-solved tasks converted
baseline never solved them in any trial
01Papers & benchmarks
Frozen-weights agents, external memory, one-bit verdicts and corrections. Cross-model transfer.
arXiv:2607.22157 · Jul 2026
SWE-bench variant and DS-1000: token cost, completion time and recommendation quality with shared memory.
arXiv:2511.08301 · Nov 2025
The field guide
A book-length guide for engineering teams making agent work compound — fifteen chapters, two plates, one diagnostic. First edition, May 2026.
“Most coding-agent demos show the first run. The second run is the signal.”
— Scott Taylor · founder’s note
02Journal
03Videos
Recorded walkthroughs and conference sessions live in the video archive.
04Methodology
Every item on this site carries one of these labels. A number without its class and source is a bug.
Numbers from evaluations Memco designed and ran. Always published with the harness and denominator.
Claims that appear in a public arXiv record. Cited by identifier; never borrowed for another evaluation’s numbers.
Someone else’s research, quoted as context. Their result, their credit — never restyled as ours.
What we’re learning while building — product updates, essays, working notes. No benchmark authority claimed.
Evidence from a named team, published only with written permission. None shown until permissioned.
the loop
benchmarks · product · research
A short dispatch on shared memory for AI agents — the numbers behind the product, what we're shipping, and the research we're reading. No filler.