Status: an observation based on public research and engineering experience; not yet validated on a unified benchmark.
Early agent memory was often reduced to “put the conversation in a vector database.” Long-horizon tasks expose three harder problems:
- Write policy: which state deserves persistence?
- Conflict management: what happens when old and new memories disagree?
- Task-aware retrieval: semantic similarity does not guarantee decision value.
The next lab will compare raw history, summarized memory, and structured state across task success, cost, and error accumulation.