Status: an observation based on public research and engineering experience; not yet validated on a unified benchmark.

Early agent memory was often reduced to “put the conversation in a vector database.” Long-horizon tasks expose three harder problems:

  1. Write policy: which state deserves persistence?
  2. Conflict management: what happens when old and new memories disagree?
  3. Task-aware retrieval: semantic similarity does not guarantee decision value.

The next lab will compare raw history, summarized memory, and structured state across task success, cost, and error accumulation.