LectionesPart XII
Agents
Where the model stops answering and starts doing.
Read Part X first. An agent is the reasoning work from that part with the model allowed to act between steps, and the papers make far more sense in that order.
The first paper is the loop that every framework you have used implements, usually without attribution. The middle two are about tools: one argues from first principles that a language model should not be doing arithmetic, the other shows a model teaching itself when to call an API with no human labelling the calls.
The last two are where the series ends, and they end it deliberately. Generative Agents is the memory architecture long-running systems still use. AutoGen is engineering rather than science, which is the honest place to stop: past here you are reading documentation, and documentation goes stale faster than a reading list can track.
The reading
- ReAct
ReAct: Synergizing Reasoning and Acting in Language Models
Yao et al. · ICLR 2023 · 2022
- Claim
- Interleaving reasoning traces with actions in an environment beats either reasoning or acting alone, and cuts hallucination on knowledge tasks.
- Why
- Thought, action, observation, repeat: the loop nearly every agent framework implements. The prompts in the appendix are the actual contribution.
- Read
- Section 2 and the appendix prompts.
- MRKL
MRKL Systems: A Modular, Neuro-Symbolic Architecture that Combines Large Language Models, External Knowledge Sources and Discrete Reasoning
Karpas et al. · AI21 Labs · 2022
- Claim
- A router in front of a set of discrete modules — a calculator, a database, a solver — fixes what a language model is structurally bad at.
- Why
- A position paper rather than a result, and it described tool use before tool use existed. Short, and the argument about arithmetic is still correct.
- Read
- Section 2.
- Toolformer
Toolformer: Language Models Can Teach Themselves to Use Tools
Schick et al. · NeurIPS · 2023
- Claim
- A model can learn when to call an API in a self-supervised way, by keeping only the calls that reduce its own loss on the text that follows.
- Why
- The filtering criterion is the elegant part: no human labels a single tool call, the loss does it.
- Read
- Section 2.
- Generative Agents
Generative Agents: Interactive Simulacra of Human Behavior
Park et al. · UIST · 2023
- Claim
- Agents with memory, reflection and planning produce believable individual behaviour and emergent social behaviour over days of simulated time.
- Why
- The memory architecture — scored by recency, importance and relevance — is the reusable piece, and it is roughly what long-running agents still do.
- Read
- Section 4, the memory stream.
- AutoGen
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Wu et al. · COLM 2024 · 2023
- Claim
- Framing applications as conversations between configurable agents — model, tool or human — covers a wide range of tasks with a single abstraction.
- Why
- Read it for the abstraction, not the interface. The interface has changed several times since; the framing has not.