Collective learning · Multimodal memory

EpiConCollective Agent Learning through Co-Evolving Multimodal Memory

* Work done while at MIT-IBM Computing Research Lab.   † Corresponding Author.

One agent’s experience.
A shared resource for many.

Refine text and visual evidence together. Consolidate lessons into a shared bank.
Reuse experience across harnesses and backbones—without changing host model parameters.

EpiCon framework: a memory controller co-evolves textual and visual memory; a tree self-organizer consolidates lessons; retrieved rules support a single inference.
THE FRAMEWORK Two independently trained 2B models connect question-level memory evolution with persistent, shared experience.
11benchmarks across 4 domains
+1.7–4.9macro-average points vs. No Memory¹
67–74%less memory-operation time²

¹ EpiCon 2B across four host configurations. ² Relative to backbone-sized EpiCon memory models; refers to memory operations, not total solving time.

Abstract

Experience that
agents can build on.

Agents can learn from past executions, but enabling different agents to reuse and build on one another’s experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hierarchically and retrieves experience and rules for new problems. We evaluate EpiCon on eleven benchmarks spanning four multimodal task domains, using two harnesses and multiple backbones. A frozen bank improves other systems even with a single solving attempt. Experience contributed by another harness raises the original system's macro-average score by 2.1 to 3.2 points across eleven benchmarks. Across four host configurations, the 2B variant improves macro-average scores by 1.7 to 4.9 points over No Memory and reduces memory-operation time by 67% to 74% relative to backbone-sized memory models.

01 / Method

Refine. Consolidate. Reuse.

01

Co-evolve memory

The Memory Controller jointly revises textual guidance and supporting image regions using the question, answer, and execution trace. An adaptive gate decides when visual memory should be included.

02

Organize experience

The Tree Self-Organizer places, merges, splits, and consolidates lessons into a hierarchy of experience and reusable rules. The memory harness validates and executes these operations.

03

Share what was learned

A new system retrieves relevant experience from the shared bank. The host keeps its own harness and backbone, with no host parameter updates.

Illustrative transit-map example: refine a Rome route, consolidate lessons from multiple cities, and retrieve a routing rule for a new Singapore map.
A WORKED EXAMPLE Text and image crops evolve on a Rome transit problem. Related lessons become a routing rule that supports solving on a new map. Trajectories and tree operations are reconstructed for illustration.

02 / Video overview

From one lesson to collective learning.

03 / Evaluation

Shared memory. Measurable gains.

4,538 questions across document understanding, visual-to-code generation, vision-grounded math, and general visual-language reasoning.

One solving attempt per question. All main-table methods use a single attempt. Model parameters and historical banks stay frozen during evaluation; question-level memory updates are disabled. Construction and evaluation questions are disjoint.
Codex · Qwen3.8-27B — main results
Memory settingMean ↑DocVQA2026MP-DocVQAParseBenchVision2CodeOmni-I2CChartMimicMATH-VisionWeMathWorldBenchReasonMapBabyVisionTime · MASTime · memoryTokens · MASTokens · memory
No Memory49.315.779.557.153.066.266.420.481.049.240.114.21.000.001.000.00
Mem050.91.481.661.961.370.678.024.480.648.639.612.21.330.311.230.11
Cognee52.614.373.760.161.569.978.144.278.251.235.411.51.161.661.110.75
A-Mem52.512.980.266.360.072.378.625.481.845.440.613.51.230.241.270.81
Agent-KB45.711.478.659.525.867.469.825.280.845.024.115.61.540.311.120.14
EpiCon (ours 2B)54.318.666.167.359.971.475.245.284.252.441.515.31.160.301.040.67
EpiCon (backbone)57.922.983.568.561.773.079.348.685.453.843.416.31.150.921.030.64

Scores: higher is better. Mean is computed as an equal-weight average of the 11 displayed benchmark scores. Time and token costs are normalized to the corresponding No Memory MAS cost; MAS and memory costs are separate. Baseline memory models match the selected host backbone and operate zero-shot. EpiCon 2B uses two independently trained 2B models. Scroll horizontally to view all columns.

Download main results as CSV ↓
Paper figure comparing benchmark profiles and score versus normalized total time for the Qwen3.8-27B configurations.
PERFORMANCE & COST Qwen3.8-27B benchmark profiles and performance–time trade-offs. The radar is normalized by the best method mean per benchmark and starts at 0.25; total time combines normalized MAS and memory costs.

04 / Collective learning

Experience travels across systems.

+3.2

Across backbones

Codex switches from Qwen3.8-27B to Gemma4-31B while reusing the frozen bank.

+4.2

Across harnesses

DeepSeek-Harness reuses experience built by Codex, with Qwen3.8-27B.

+2.1–3.2

Back to the contributor

Experience added by another harness improves the original contributor’s macro-average.

All gains are macro-average score points over 11 benchmarks, as reported in the paper. Transfer compares EpiCon 2B with No Memory; the return-to-contributor comparison uses original vs. evolved banks.

Cross-backbone & cross-harness transfer

Bank source: Codex with Qwen3.8-27B. Each target makes one solving attempt with a frozen bank and no question-level updates.

Cross-backbone: Codex · Gemma4-31B
Memory settingMean ↑DocVQA2026MP-DocVQAParseBenchVision2CodeOmni-I2CChartMimicMATH-VisionWeMathWorldBenchReasonMapBabyVision
No Memory44.05.746.243.351.765.965.625.885.045.431.617.7
EpiCon (ours 2B)47.25.760.445.353.867.868.233.085.047.834.418.1
Cross-harness: DeepSeek-Harness · Qwen3.8-27B
Memory settingMean ↑DocVQA2026MP-DocVQAParseBenchVision2CodeOmni-I2CChartMimicMATH-VisionWeMathWorldBenchReasonMapBabyVision
No Memory50.414.379.552.453.761.465.342.288.854.022.620.1
EpiCon (ours 2B)54.618.679.165.759.263.368.756.090.652.226.420.5
Cross-harness memory evolution

The original and contributing harnesses each use 1,000 disjoint construction questions. Evaluation uses 4,388 questions disjoint from both construction sets. Both bank conditions use one solve. This comparison measures bank evolution and expansion together; individual benchmarks do not improve uniformly.

Codex receives experience from DeepSeek-Harness
Memory settingMean ↑DocVQA2026MP-DocVQAParseBenchVision2CodeOmni-I2CChartMimicMATH-VisionWeMathWorldBenchReasonMapBabyVision
Original bank53.518.666.163.459.971.475.245.284.252.438.314.3
Evolved bank55.718.680.358.961.370.976.745.488.455.240.116.8
DeepSeek-Harness receives experience from Codex
Memory settingMean ↑DocVQA2026MP-DocVQAParseBenchVision2CodeOmni-I2CChartMimicMATH-VisionWeMathWorldBenchReasonMapBabyVision
Original bank53.418.681.852.356.963.565.256.692.257.423.519.3
Evolved bank56.620.082.765.958.769.272.153.491.257.629.622.3

05 / Analysis

What makes the memory useful?

Structure matters

Tree memory organizes and consolidates lessons from the same source questions, improving all four reported scores over a flat bank.

Flat vs. tree memory · Codex / Qwen3.8-27B
Memory settingParseBenchVision2CodeMATH-VisionBabyVision
Flat52.257.543.613.2
Tree67.359.945.215.3

Useful to a stronger solver

With Codex and GPT-5.6-Luna, the fixed 2B components and frozen bank improve all four reported benchmarks, without target-specific retraining.

Codex · GPT-5.6-Luna · single solving attempt
Memory settingParseBenchVision2CodeMATH-VisionBabyVision
No Memory63.865.851.211.1
EpiCon (ours 2B)66.567.557.214.6
Question-level control & visual-memory ablations

These experiments disable historical retrieval. No Memory uses one attempt; memory-enabled settings allow up to five attempts. Original question images remain available in every setting. Both harnesses use Qwen3.8-27B.

Codex · question-level memory control
Memory settingParseBenchVision2CodeMATH-VisionBabyVisionTime · MASTime · memoryTokens · MASTokens · memory
No Memory (1 attempt)57.153.020.414.21.000.001.000.00
MC (ours 2B; ≤5 attempts)59.855.941.217.72.160.232.400.96
MC (Qwen3.8-27B; ≤5 attempts)63.156.649.613.92.570.233.070.95
Codex · visual-memory ablation
Memory settingParseBenchVision2CodeMATH-VisionBabyVisionTime · MASTime · memoryTokens · MASTokens · memory
No Memory (1 attempt)57.153.020.414.21.000.001.000.00
Text only58.755.727.216.01.430.281.230.61
Frozen visual · always55.255.040.616.02.820.433.000.94
Frozen visual · adaptive57.155.941.016.32.020.252.420.94
Evolving visual · always55.355.241.016.32.890.443.020.93
Evolving visual · adaptive59.855.941.217.72.160.232.400.96
DeepSeek-Harness · question-level memory control
Memory settingParseBenchVision2CodeMATH-VisionBabyVisionTime · MASTime · memoryTokens · MASTokens · memory
No Memory (1 attempt)52.453.742.220.11.000.001.000.00
MC (ours 2B; ≤5 attempts)55.157.851.423.32.100.132.443.23
MC (Qwen3.8-27B; ≤5 attempts)64.565.861.635.12.470.162.653.25
DeepSeek-Harness · visual-memory ablation
Memory settingParseBenchVision2CodeMATH-VisionBabyVisionTime · MASTime · memoryTokens · MASTokens · memory
No Memory (1 attempt)52.453.742.220.11.000.001.000.00
Text only51.855.147.421.21.330.131.252.37
Frozen visual · always51.257.550.220.52.520.213.133.75
Frozen visual · adaptive52.857.350.821.21.950.132.413.58
Evolving visual · always52.157.051.821.22.530.213.153.72
Evolving visual · adaptive55.157.851.423.32.100.132.443.23

06 / Citation

BibTeX

@misc{zeng2026epicon,
  title         = {EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory},
  author        = {Ziyun Zeng and Hang Hua and Shaden Alshammari and Rogerio Feris and William T. Freeman and Jiebo Luo},
  year          = {2026},
  archivePrefix = {arXiv},
  eprint        = {2609.37923},
  url           = {https://arxiv.org/abs/2609.37923}
}