[{"content":"This topic follows the system problems that emerge when agents move from short conversations to sustained work.\nValidation map Memory: write, update, conflict, and forgetting State: reliable task recovery Tools: permissions, idempotency, and error handling Evaluation: process quality beyond final success Human intervention: when the agent must stop and ask The topic will combine frontier notes, minimum experiments, and engineering decision checklists.\n","permalink":"https://yanlunai.com/en/topics/long-horizon-agents/","summary":"Tracking memory, state, tool use, recovery, and human intervention in long-running tasks.","title":"Long-horizon agents: from answers to sustained work"},{"content":" Status: an observation based on public research and engineering experience; not yet validated on a unified benchmark.\nEarly agent memory was often reduced to “put the conversation in a vector database.” Long-horizon tasks expose three harder problems:\nWrite policy: which state deserves persistence? Conflict management: what happens when old and new memories disagree? Task-aware retrieval: semantic similarity does not guarantee decision value. The next lab will compare raw history, summarized memory, and structured state across task success, cost, and error accumulation.\n","permalink":"https://yanlunai.com/en/frontier/agent-memory-2026/","summary":"The hard part is not storing history. It is deciding what to write, when to retrieve it, and when to forget.","title":"Agent memory is becoming a decision problem—not a storage problem"},{"content":"This topic asks three questions: which tasks are suitable for delegation, how completion quality can be verified, and how domain expertise changes agent performance.\nPlanned labs compare manual execution, AI assistance, and agent delegation using elapsed time, edit count, test results, and review cost.\n","permalink":"https://yanlunai.com/en/topics/ai-coding/","summary":"How delegation, verification evidence, and expertise determine the real return from coding agents.","title":"AI coding and human collaboration"},{"content":" Status: an industry signal that must be validated against each team\u0026rsquo;s workflow.\nOnce coding agents can modify multiple files, run tests, and iterate, the evaluation unit changes:\nCan the agent respect a clear task boundary? Are failures discovered quickly? Is human review cheaper than the execution time saved? Does the change retain verifiable evidence? The useful metric is no longer “how much code was generated,” but “how many tasks were reliably completed under verifiable conditions.”\n","permalink":"https://yanlunai.com/en/industry/ai-coding-adoption/","summary":"Teams now need to evaluate task boundaries, verification cost, and accountability—not only generation speed.","title":"AI coding adoption is shifting from completion rate to task delegation"},{"content":"This topic builds a tiered task set, selects the lowest-cost model that satisfies the quality threshold for each tier, and records the real cost of routing mistakes.\n","permalink":"https://yanlunai.com/en/topics/inference-routing/","summary":"Moving beyond one benchmark to task quality, latency, throughput, and cost per successful task.","title":"Inference efficiency and model routing"},{"content":" Status: experiment design demo; measurements will be added after real model endpoints are connected.\nModel routing cannot optimize for price per call alone. A minimum experiment records:\nMetric Why it matters Task success Whether the output is actually usable Time to first token Interactive responsiveness Total latency End-to-end throughput Cost per success Avoiding models that are cheap but fail more often The experiment separates extraction, structured generation, and multi-step reasoning, then selects the lowest-cost model that satisfies the quality threshold for each tier.\n","permalink":"https://yanlunai.com/en/labs/model-routing-lab/","summary":"A reproducible routing skeleton that records task success, time to first token, total latency, and cost per successful task.","title":"A minimum model-routing lab: quality, latency, and cost together"},{"content":"A trustworthy agent is not a safety sentence in the prompt. It is a system of permissions, tools, audit evidence, and human confirmation.\nThe topic begins with read-only work, then tests the additional controls required for files, external systems, and irreversible actions.\n","permalink":"https://yanlunai.com/en/topics/trustworthy-agents/","summary":"Designing for prompt injection, permission boundaries, tool side effects, and high-risk actions.","title":"Trustworthy agents and system security"},{"content":"RAG failures are rarely solved by a larger model alone. The pipeline must be stable end to end.\nA minimum viable RAG pipeline Raw sources → cleaning and structural chunking Embeddings → indexing Retrieval → candidate evidence Reranking → relevance convergence Generation → citation constraints Evaluation → traceable quality signals If only one improvement can be funded, start with reranking plus evaluation. Better ranking reduces irrelevant evidence; evaluation tells you whether the change helped the actual task.\nEvidence state: engineering synthesis. A reproducible comparison will be added in a later lab.\n","permalink":"https://yanlunai.com/en/posts/tools/2025-12-16-rag-toolchain-overview/","summary":"Treat RAG as a pipeline: data, chunking, indexing, retrieval, reranking, generation, and evaluation. Any stage can become the bottleneck.","title":"The RAG toolchain: a minimum loop from data to evaluation"},{"content":"A useful distinction between chat and an agent is whether the system maintains explicit intermediate state and can respond to failed decisions.\nWhat planning adds a representation of the current objective; a sequence of executable steps; observations from tools or environments; a policy for revising the plan after failure. This also changes evaluation. Final-answer accuracy is insufficient: we need task success, step efficiency, recovery quality, tool errors, and unnecessary actions.\nEvidence state: conceptual synthesis. The long-horizon agent topic will test these dimensions in a minimum task environment.\n","permalink":"https://yanlunai.com/en/posts/papers/2025-12-15-agent-planning-view/","summary":"An agent is more than a multi-turn chat interface: it maintains state, decomposes goals, uses tools, and recovers from failure.","title":"From chat to agents: planning is the dividing line"},{"content":"Trustworthy RAG can be reduced to three engineering questions:\nCan every important claim be traced to evidence? Is the answer consistent with that evidence? Can the system abstain when the evidence is insufficient? Each question requires a different control: source identifiers and versioning, grounded-generation checks, and an explicit uncertainty or refusal policy.\nThe practical goal is not to eliminate every hallucination. It is to make failures observable, measurable, and safer to handle.\nEvidence state: engineering framework. Metric definitions and a small evaluation set will follow.\n","permalink":"https://yanlunai.com/en/posts/papers/2025-12-15-clean-rag-engineering/","summary":"Trustworthy RAG is not a prompt. It is a system for traceable evidence, grounded answers, and abstention under uncertainty.","title":"An engineering view of trustworthy RAG"}]