A useful distinction between chat and an agent is whether the system maintains explicit intermediate state and can respond to failed decisions.

What planning adds

  • a representation of the current objective;
  • a sequence of executable steps;
  • observations from tools or environments;
  • a policy for revising the plan after failure.

This also changes evaluation. Final-answer accuracy is insufficient: we need task success, step efficiency, recovery quality, tool errors, and unnecessary actions.

Evidence state: conceptual synthesis. The long-horizon agent topic will test these dimensions in a minimum task environment.