Panorama of Agent

1. What is Agent & Harness?

First, the Large Language Model(LLM) is very mature now. Its essence is a prediction model:

  • You input a string of token, and it outputs a string of token.

The LLM nowadays can answer our questions but it can not finish a job independently. For example, you can ask LLM how to create a new folder on PC and it can give you an answer. However, the LLM cannot create the folder directly because the only thing it can do is output tokens.

Now we have LLM to describe the workflow of a job. But we need to finish the workflow ourselves. Agent is the software that can resolve problems independently.

Agent is usually composed of two parts: LLM and Harness.

Agent ≈ LLM-based Policy + Harness

The LLM provides the main reasoning/policy capability, while the Harness provides orchestration, state/context management, tools, environment interaction, execution, and runtime controls.

e.g. LLM thought it needs to read a file, then the Harness part will read the file and return the result to LLM.

But an Agent does not exist in a vacuum. An Agent is defined jointly with the environment it acts in. The Harness interacts with the environment through an Interface: it renders the environment’s state into observations the LLM can read, and translates the LLM’s decisions into actions the environment accepts.

Interface:  O = Φ(S),   A ⊆ 𝒯

So the well-known decomposition describes the Agent’s capability, while the environment class determines the shape of that capability:

  • Different environments expose different observation formats → different Context & Memory designs
  • Different environments accept different action spaces → different Tools & ACI designs
  • Different environments offer different verification signals → different training methods available

This is why SWE Agents and Web Agents look so different even when they share the same LLM: they are two extreme points of the same Interface design axis, not two unrelated species of Agent.

In Agent, LLM is more likely a role of “Brain”, and the Harness is the “Limb”.

2. Agent’s Pursuit

Agent is designed to solve problem independently. Everything we do now is to improve Agent’s ability of solving problems.

We hope Agent can solve particular problems in daily life and complex problems in specific area like software engineering.

3. Optimizing an Agent

Agent is composed of LLM and Harness.

So it is obvious that we should enhance LLM’s performance in Agent area. And optimize the Harness part. Then we need benchmarks to judge the Agent’s performance.

This article is organised as six parts, following one chain of reasoning:

Part Question it answers
1. What is an Agent? What is it made of?
2. The Pursuit of Agent Why do we build it?
3. Optimizing an Agent How do we make it stronger?
4. Agent Categories Which kinds exist, and how is each one actually built?
5. Benchmark How do we measure it?
6. Multi-Agent What changes when there are several?

Parts 3 and 4 are complementary and deliberately separated:

  • Part 3 is the abstract layer – it asks which components an Agent has, and how to improve each one. The seven Harness layers in 3.2.2 are meant to apply to every Agent.
  • Part 4 is the concrete layer – it asks how those same components are actually instantiated in each environment: what the observation interface looks like, what the action space allows, what the classic implementation does, and where the real engineering pain is.

The same Orchestration layer from Part 3 looks completely different in a software repository than in a browser. Part 4 is about that difference.

This is the original version of Agent Panorama:

This is the second version which considering the relation between Env and Agent:

3.1 LLM Tuning

In Agent area, the most popular post-training methods are:

  • Imitation Learning (SFT)
  • Online RL
  • Offline RL
  • Preference Learning (DPO)

3.1.1 Imitation Learning (SFT)

ℒSFT = −∑tlog Pθ(tokent ∣ token < t)

3.1.1.1 Trajectory SFT

Core Process:

1
2
3
4
5
6
7
Expert Agent Trajectory
↓
Teacher Forcing
↓
Cross Entropy
↓
LLM

Trajectory (ReAct) :

1
2
3
4
5
6
7
task
→ thought
→ action
→ observation
→ action
→ ...
→ final answer

Relative papers:

3.1.1.2 Critical-Step SFT

A trajectory includes many steps, but not all steps are necessary. One approach is to apply SFT only to the critical steps within a trajectory.

  • ATLaS — (Findings ACL 2025), only train critical steps

3.1.1.3 Synthetic Trajectory

One problem of SFT is that high-quality training data is few, so there is a research direction that aims to synthesize trajectories.

  • SWE-smith (NeurIPS 2025), Synthetic Trajectory

3.1.1.4 Distillation

Distillation is to extract reasoning trajectories from different teachers and distill into a student model.

1
2
3
4
5
6
7
8
9
10
11
12
13
Teacher A
│
├── Tool-based reasoning
│
Teacher B
│
└── Text reasoning
│
▼
Distill
│
▼
Student Agent
  • Agentic-R1 (EMNLP 2025)

3.1.2 Online RL

3.1.2.1 Reward Design

The reward is a criterion to judge how good a trajectory is.

There are many ways to calculate reward:

1
2
3
4
5
6
7
            Reward
│
┌───────────┼────────────┐
│ │ │
Human Verifier Process Reward
│ │ │
RLHF RLVR PRM/PR

RLVR (Reinforcement Learning Verifiable Reward) is an important type, because many engineering problems are verifiable.

  • SWE-RL (NeurIPS 2025)
  • WebAgent-R1 (EMNLP 2025)

Process Reward Model(PRM) is another research roadmap; There are many steps in a trajectory. The traditional reward is trajectory-level, which means thatthe LLM cannot tell which step in a trajectory matters.

1
2
3
4
5
6
7
8
9
Step 1
Step 2
Step 3
...
Step 50
↓
Success
↓
Reward = 1
  • LLM cannot know if step 3 is really important.
1
2
3
4
5
Outcome Reward
↓
Process Reward
↓
Step-level Credit Assignment
  • Rewarding Progress (ICLR 2025)
  • PURE (NeurIPS 2025)
  • ReasonFlux-PRM (NeurIPS 2025)

3.1.2.2 RL Optimization Algorithm

After calculating the reward of a trajectory, we need to update the weights of LLM.

There are three ways to update weights, but the PMD is not commonly used.

1
2
3
4
5
6
7
8
9
10
11
12
         Reward
│
▼
Policy Optimization
│
┌─────────┼─────────┐
│ │ │
PPO GRPO PMD
│ │ │
└─────────┼─────────┘
▼
Update LLM
3.1.2.2.1 GRPO

GRPO has become one of the most widely used policy-optimization methods for reasoning-oriented LLM post-training. A key difference from standard PPO is that GRPO estimates relative advantages from a group of sampled responses instead of maintaining a separate value/critic model.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
Question x
│
├── rollout 1 → R1
├── rollout 2 → R2
├── rollout 3 → R3
├── ...
└── rollout G → RG
│
▼
Group Relative Advantage
│
▼
GRPO
│
▼
LLM
  • SWE-RL
3.1.2.2.2 PPO

The core thought of PPO (Proximal Policy Optimization) is:

  • Need to make LLM more inclined to generate high-reward responses, while ensuring that a single update does not deviate too far from the original model.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
Question x
│
▼
Policy πold
│
▼
rollout → trajectory
│
├── Reward R
│
└── States / Actions
│
▼
Value / Critic V(s)
│
▼
Advantage A
│
▼
PPO Objective
│
┌─────┴─────┐
│ │
Probability Clip
Ratio
│ │
└─────┬─────┘
▼
Loss
│
▼
Update Policy πθ

3.1.3 Offline RL

  • OREO (ACL 2025): Offline Reinforcement Learning for LLM Multi-step Reasoning
  • Reward-Weighted Fine-Tuning (NeurIPS 2025): Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization

3.1.4 Preference Learning (DPO)

DPO (Direct Preference Optimization) is widely used in LLM post-training, but in Agent-specific post-training it is generally less dominant than SFT and online RL.

SFT teaches the LLM what the right answer is while Preference Learning just tells LLM: A is better than B.

1
2
3
4
Task

Trajectory A -> good
Trajectory B -> bad

Train:

1
2
πθ(A|x) ↑
πθ(B|x) ↓

Basic DPO process:

1
2
3
4
5
6
7
Prompt
├── Chosen
└── Rejected
↓
DPO
↓
LLM

Traditional DPO compares entire trajectories. Nowadays the tendency is to research which certain step makes trajectory better.

1
2
3
4
5
6
7
Trajectory-level
↓
Step-level
↓
Action-group level
↓
Critical-step preference
  • HPL (ICLR 2026): Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
  • Online DPO (ICLR 2025)
  • SDPO (ACL 2025)

3.2 Harness Structure

3.2.1 Environment Classes —— The First-Order Classification

Before listing the Harness layers, we must first answer: what kind of world does the Agent live in? This is the first-order classification of the field, because the environment class fixes three things at once.

Axis Spectrum
Observability(观测形态) plain text tokens → structured tree (DOM / AX) → pixel screenshots → video / point cloud / proprioception
Action space(动作自由度) open primitives, new tools writable on the fly → closed enumerated actions → continuous motor control
Verifiability(可验证性) test pass/fail → state assertion → LLM judge → human rating

The taxonomy induced by these three axes:

Environment Class Observability Action Space Verifiability
Classic / Text-World
ALFWorld, WebShop
text description few discrete text actions environment state
Software Engineering
SWE-bench
files, stdout, diff curated ACI + shell (open) test suite (pass/fail)
Tool / API / Enterprise
BFCL, τ-bench, TheAgentCompany
JSON tool returns tool calls (closed) state assertion + user simulator
Web / Browser
WebArena, Mind2Web
DOM / AX tree / screenshot click / type / scroll / goto (closed) state assertion
Computer-Use / OS
OSWorld
screenshot (+ AX tree) click(x,y) / type / key (closed) state assertion
Mobile
AndroidWorld
screenshot + view tree tap / swipe / back (closed) state assertion
Deep Research
BrowseComp
web text search / fetch / synthesize LLM judge (weak)
Embodied
ALFRED, Habitat
video / point cloud / proprioception continuous motor primitives physical / simulator

Key observation — the action space and the environment class are two views of one thing. An environment is its action space generator: Aenv = ActionSpace(Environment)

So “designing an Agent” and “designing an action space” are the same activity seen from two sides. SWE Agents collapse the action space into one escape hatch (a shell) plus a few curated commands; Web Agents are handed a closed, enumerated action set. This is not a difference in Agent architecture — it is a difference in what the environment chose to expose.

Second observation — the available training method follows verification, not environment complexity. SWE has a free correctness signal (run the tests), which is why SWE-RL and SWE-Gym advanced fastest and why RLVR works there. Deep Research is a far simpler environment, yet there is no cheap ground truth, so the field is left with LLM-as-judge and weak process rewards (see 3.1.2.1 and 3.1.2.2).

For this research map, the Harness is organized into seven engineering-oriented layers. This is a practical taxonomy rather than a universally accepted academic standard, and the layers can overlap in real systems.

Part Abstract Description
Orchestration
Loop / Planning / ReAct / Reflection / Verification
Control How should the agent decide what to do?
Context & Memory
Context Management / Retrieval / Compression / Memory
State What should the agent remember?
Tools & Actions
Tool Calling / Discovery / Selection / Code / Browser / Computer
Action What can the agent do?
ACI & Environment
State / Action Space / Observation / Error / Feedback
Interface What state, observations, actions, errors, and feedback does the environment expose to the agent?
Skills & Composition
Skills / Workflow / Subagents / Agent-as-a-Tool
Capability How can reusable capabilities and workflows be represented, retrieved, composed, and reused?
Protocol & Interoperability
MCP / A2A / Agent API / Session Protocol
Interoperability How do agents/tools/services communicate?
Runtime & Safety
Sandbox / Execution / Permission / Guardrail / Reliability / Observability / Cost
Runtime & Governance Where and how are actions executed?
How do we keep agents safe, reliable, observable, recoverable, and economical?
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
Agent
│
├── LLM
│
└── Harness
│
├── 1. Orchestration
│ Loop / Planning / ReAct / Reflection / Verification
│
├── 2. Context & Memory
│ Context Management / Retrieval / Compression / Memory
│
├── 3. Tools & Actions
│ Tool Calling / Discovery / Selection / Code / Browser / Computer
│
├── 4. ACI & Environment
│ State / Action Space / Observation / Error / Feedback
│
├── 5. Skills & Composition
│ Skills / Workflow / Subagents / Agent-as-a-Tool
│
├── 6. Protocol & Interoperability
│ MCP / A2A / Agent API / Session Protocol
│
└── 7. Runtime & Safety
Sandbox / Execution / Permission / Guardrail /
Reliability / Observability / Cost

3.2.2 Orchestration / Control

Q: How should the Agent decide what to do?

1
2
3
4
5
6
Orchestration
├── Loop
├── Planning
├── ReAct
├── Reflection
└── Verification

Earlier Agent(Classic ReAct):

1
LLM → Tool → LLM → Tool

Nowadays Agent:

1
Plan → Act → Verify → Reflect → Re-plan → Act

From simple ReAct Loop to a Loop with planning / verification / recovery.

  • PlanGEN (EMNLP 2025): This work combines planning and iterative verification.

  • ReflAct (EMNLP 2025): The traditional ReAct structure will accumulate little errors which will make agent drift from correct path. (ALFWorld + 27.7% than ReAct)

  • 1
    2
    3
    4
    5
    6
    7
    Current State
    ↕
    Goal State
    ↓
    Reflection
    ↓
    Action
  • MetaAgent-P (ACL 2025)

3.2.3 Context & Memory / State

Q: What should the Agent know at the current moment, and what past information should be retained ?

1
2
3
4
5
6
7
Orchestration
↓
“What should I do?”

Context & Memory
↓
“What should I know?”

The Context & Memory is not just ‘Chat Log’. It store the Agent State.

1
2
3
4
5
6
7
8
9
Current Context
+
Working Memory
+
Long-term Memory
+
Retrieved Experience
+
Compressed History

The tradition memory system:

1
2
3
Memory
=
Vector DB + Embedding + Retrieval

Nowadays:

1
2
3
4
5
6
Agent decides:
store?
update?
delete?
retrieve?
summarize?

Memory Management will be considered into Agent Policy.

  • Memory OS of AI Agent, (EMNLP 2025)

    This work design the Memory System like operating system hierarchy.

    1
    2
    3
    4
    5
    Short-term
    ↓
    Mid-term
    ↓
    Long-term
  • Coarse-to-Fine Grounded Memory, (EMNLP 2025)

    This work use memory to replan. Memory is used to store useful things.

    1
    2
    3
    4
    5
    6
    7
    Past Experience
    ↓
    Grounded Memory
    ↓
    Current Situation
    ↓
    Planning
  • Agentic Memory, ACL 2026

  • Memory-R1,ACL 2026

One problem:

How Memory Management Impacts LLM Agents, ACL 2026:

The agent will clearly exhibit “experience-following”: after retrieving historical experience from similar tasks, it is easy to reproduce similar behaviors. So if there exists wrong memory, the Agent is inclined to repeat the wrong operations.

3.2.4 Tools & Actions / Action

What can Agent do ?

The tradition actions system is simply call tools like:

1
2
3
4
5
6
7
8
Search
Code
Database
Browser
Filesystem
Shell
Computer
API

The new tendency of actions area is:

1
2
3
4
5
Tool Discovery
Tool Selection
Tool Composition
Tool Creation
Computer Use
  • LLM Agents Making Agent Tools,ACL 2025

    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    Paper + GitHub Code
    ↓
    Agent
    ↓
    Install dependencies
    ↓
    Generate tool
    ↓
    Execute
    ↓
    Debug / Self-correct

    In this work, the Agent creates its own tools

  • Adaptive Tool Use,ACL 2025

    More tools do not necessarily mean a better Agent. If one tool frequently returns errors, the Agent should realize that.

  • ToolScope,ACL 2026

    In large tool ecosystem, there exists a problem:

    1
    2
    3
    4
    5
    6
    7
    Too Many Tools 
    +
    Repeated Tools
    +
    Tedious Schema
    +
    Trouble Choosing

    Solution:

    1
    2
    3
    4
    5
    Tool Merging
    +
    Tool Retrieval
    +
    Context-aware Filtering

3.2.5 ACI & Environment / Interface

Agent-Computer Interface.

This part is to solve the problem: What form should the external world take for the agent, and how should the agent act upon it?

A common API-agent interface takes a form like:

1
search(query)

A computer-use Agent instead exposes a human-like interaction loop:

1
2
3
4
5
6
7
8
9
Screenshot
↓
LLM
↓
click(x,y)
↓
Screenshot
↓
...

The key shift is from tool-specific APIs to general-purpose interaction interfaces. Instead of designing a dedicated tool for every task, the Agent can operate through interfaces already available to humans, such as GUI, CLI, and Web.

  • OS Agents Survey,ACL 2025:

    The paper defines OS Agent that interact with:

    1
    2
    3
    4
    5
    GUI
    CLI
    Web
    Mobile
    Desktop
  • OSWorld-Human, ICML 2025:

    To finish a certain task, the Agent takes more steps/operations than human.

    1
    2
    3
    4
    5
    Agent Benchmark
    ↓
    Can you finish?
    +
    Can you finish efficiently?

3.2.6 Skills & Composition / Capability

How can we extract complex, repeatable capabilities from a specific agent and turn them into reusable, composable capabilities?

1
2
3
Tool = an executable capability or interface for an action

Skill = reusable procedural knowledge or capability that can guide or execute a class of tasks.
1
2
3
Tools → Atomic Capability

Skills → procedural capability

Tendency: Skills are evolving from prompts into software artifacts.

The traditional Skill is a text:

1
2
3
4
5
To finish ... task, you should:
1. ...
2. ...
3. ...
4. ...

It’s a stable text describing a fixed workflow.

  • Agent Skills,Anthropic 2025

    Skill is no more a huge text file. It’s new structure:

    1
    2
    3
    4
    Skill/
    ├── instructions
    ├── scripts
    └── resources

    e.g:

    1
    2
    3
    4
    5
    6
    7
    8
    9
    pdf-to-excel/
    ├── SKILL.md
    ├── scripts/
    │ ├── extract_table.py
    │ ├── clean_data.py
    │ └── validate.py
    └── resources/
    ├── excel_template.xlsx
    └── examples/

    There’s a small Skill metadata like tool definitions will be read by Agent when system starts:

    1
    2
    3
    4
    5
    6
    7
    Available Skill:

    Name:
    pdf-to-excel

    Description:
    Extract structured tables from PDFs and produce validated Excel files.

    When handling a certain task:

    1
    2
    3
    4
    5
    Task
    ↓
    Skill Retrieval
    ↓
    Load Relative Skill, e.g. pdf-to-excel/SKILL.md

    In Skill.md, there may exists content like:

    1
    2
    3
    4
    5
    6
    7
    ...

    Run:

    python scripts/extract_table.py report.pdf

    ...

    In effect, the scripts act as tools for skills.

    GitHub - anthropics/skills: Public repository for Agent Skills · GitHub

  • SkillWeaver, arXiv:2504.07079

    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    12
    13
    14
    15
    16
    17
    Web Task
    ↓
    Agent Exploration
    ↓
    Successful Trajectory
    ↓
    Skill Discovery
    ↓
    Skill Synthesis
    ↓
    Skill Practice / Debugging
    ↓
    Reusable Skill API
    ↓
    Skill Library
    ↓
    Future Tasks
  • SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning (arXiv:2602.08234, Feb 2026)

    Distills raw experience into a hierarchical SkillBank, then lets the skill library co-evolve with the policy during RL, instead of storing raw trajectories.

    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    12
    13
    14
    15
    Experience
    ↓
    Skill Distillation
    ↓
    Hierarchical SkillBank
    ↓
    Retrieval
    ↓
    RL Training
    ↓
    Policy Improvement
    ↓
    New Experience
    ↓
    SkillBank Evolution
  • SkillCraft: Can LLM Agents Learn to Use Tools Skillfully? (arXiv:2603.00718, Feb 2026)

    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    Task A
    ↓
    Learn Skill
    ↓
    Task B
    ↓
    Reuse Skill
    ↓
    Task C
    ↓
    Reuse / Compose Skill
  • Generative Skill Composition for LLM Agents (arXiv:2606.32025, Jun 2026)

    Formalizes structured skill composition — which skills, how many, and in what order — as task-conditioned skill sequence prediction, decoded jointly rather than by retrieval plus reranking.

  • Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward (arXiv:2602.12430, Feb 2026), Agent Skills ’26 Workshop @ ACM CAAS 2026

    A survey of the skill abstraction layer, covering SKILL.md, progressive context loading, the complementary roles of Skills and MCP, and security (it reports that 26.1% of community-contributed skills contain vulnerabilities).

3.2.7 Protocol & Interoperability / Interoperability

Solve the problems of communication between Agents, Tools, Data Source and Services.

  • MCP (Model Context Protocol):

    1
    2
    3
    Agent
    ↕
    Tools / Data / Services
  • A2A (Agent to Agent):

    Different Agents should have the capability to work together.

    1
    Agent A ↔ Agent B
  • MCP-AgentBench, AAAI 2026

3.2.8 Runtime & Safety

Make the Agent actually run in a controlled, secure, observable, and recoverable environment.

Traditional safety check:

1
2
3
4
5
Prompt
↓
Output
↓
Safety Check

Agent workflow:

1
2
3
4
5
6
7
8
9
Prompt
↓
Agent
↓
Tool Call
↓
External Data
↓
Side Effect

So we should check safety of not only prompt and output but also Action, Tool Call, State Transition and Permission.

  • Task Shield,ACL 2025

    Is this action really relevant to the user’s task?

  • IPIGuard,EMNLP 2025

    Create a Tool Dependency Graph, and limit the tool call by the dependency graph.

3.2.9 Conclusion

Field 2025 Tendency in 2026
Orchestration ReAct、Planning、Reflection Adaptive / verified / long-horizon control
Context & Memory RAG、hierarchical memory、context compression Agentic memory、learned memory management
Tools & Actions Tool calling、tool selection Tool discovery、tool management、tool creation、computer use
ACI & Environment Browser / GUI agents Efficient action spaces、long-horizon computer interaction、Secure ACI
Skills & Composition Skills、workflows、multi-agent Reusable capability、skill lifecycle、agent composition
Protocol & Interoperability MCP、A2A large ecosystem、discovery、identity、security、task protocols
Runtime & Safety Sandbox、guardrails、tracing Durable runtime、containment、permission、agent security、cost/reliability

4. Agent Categories

4.1 How to Read This Chapter

Part 3 was about the abstract Agent: which components it has, and how to improve each one. This part is about the concrete Agent: how those same components are instantiated in each environment, and what the classic implementation of each category actually looks like.

Every category below is described with the same seven fields, so they can be compared side by side:

Field What it captures
Observation Interface What the Agent can perceive
Action Space What the Agent is allowed to do
Classic Agent The canonical, widely-copied implementation
Harness Implementation How the layers from 3.2 are actually realised here
Practical Usage When this category is the right tool
Verification & Training What signal is available, so what training works
Current Limits Where it breaks today

Scope note. Categories are chosen by influence within the field and by the availability of a public reference implementation, not by current SOTA. Embodied and Mobile Agents are out of scope for this article; the Generalist (“cross-environment”) setting is discussed at the end as an evaluation axis rather than as a category of its own.

4.2 Environment Classes, Revisited

3.2.1 already established the first-order classification. For convenience:

Category Environment Observation Action Space Verification Representative Benchmarks
Software Engineering repository + OS sandbox files, stdout, diff curated ACI + shell (open) test suite (pass/fail) SWE-bench, SWE-bench Verified, SWE-Bench Pro
Web / Browser website DOM / AX tree / screenshot click / type / scroll / goto (closed) state assertion WebArena, Mind2Web, WebVoyager
Computer-Use / OS desktop OS, cross-application screenshot (+ AX tree) click(x,y) / type / key (closed) state assertion OSWorld, OSWorld-Human, WindowsAgentArena
Deep Research open web + documents web text search / fetch / synthesize LLM judge (weak) BrowseComp, BrowseComp-Plus, MLR-Bench
Tool / API / Enterprise borrowed – no dedicated world JSON tool returns tool calls (closed) state assertion + user simulator BFCL, τ-bench / τ²-bench, MCP-Bench, TheAgentCompany

The single most useful thing to notice in this table is the open vs. closed action space column. It is the sharpest dividing line in the whole field, and it explains most of the downstream differences: whether the Agent can create new tools mid-task, how much of the work lives in the Harness, and how safely the Agent can be deployed.

4.3 Software Engineering Agent

Observation Interface. A repository plus a shell. The Agent reads files, runs commands, and inspects stdout, stderr, exit codes and diffs. Crucially, output can be filtered before it reaches the model: grep, head, tail, and truncation are all legal, so this is a token-thrifty environment.

Action Space. Open. The Agent can write a new script and execute it. This is the defining property of the category: the Agent can create new tools for itself in the middle of a task.

1
2
3
4
5
6
7
8
9
write repro.py
↓
python repro.py ← a tool that did not exist 30 seconds ago
↓
read the traceback
↓
edit the fix
↓
run the project's test suite

Classic Agent: SWE-agent. Its contribution was not a better model but a better interface: the Agent-Computer Interface (ACI). Instead of handing the model a raw shell, SWE-agent defines a small set of structured commands (open, search_file, search_dir, edit, goto) whose output is bounded and whose syntax is validated. The lesson generalises: when the model is weak at an interface, redesign the interface rather than fine-tune the model.

Harness Implementation.

Layer How it is realised here
Orchestration ReAct loop, plus a verification step (run tests) and a reflection step on failure
Context & Memory File slices on demand; aggressive trajectory compression; the “last N observations” rule
Tools & Actions A small curated command set, plus an escape hatch to the shell
ACI & Environment This is the heart of the category – command design, output truncation, edit validation
Skills & Composition Reusable recipes (“how to run this project’s tests”), plus subagents for exploration
Protocol MCP servers for issue trackers, CI, and code search
Runtime & Safety Container sandbox, filesystem scoping, cost caps

Practical Usage. Bug fixing in an existing repository, test generation, dependency upgrades, code review, and repository-scale refactors. It is the category with the best cost-benefit ratio today, because verification is free and the environment is reproducible.

Verification & Training. Tests give a free 0/1 reward, which is why this category leads in RL: SWE-RL, SWE-Gym and SWE-Lego all exploit it. Trajectory SFT remains the base, and RLVR is the differentiator. This is the only category where RL is straightforwardly worth the engineering cost.

Current Limits. Localisation is still the bottleneck, not editing. Performance is also contamination-sensitive: as shown in 5.2.7, models score far better on SWE-bench Verified than on comparable fresh benchmarks, which suggests part of the reported progress is memorisation rather than skill.

4.4 Web / Browser Agent

Observation Interface. A live web page, exposed as some combination of DOM, accessibility tree, and screenshot. Unlike SWE, the observation is expensive: a full page state may cost thousands of tokens and cannot be filtered the way stdout can.

Action Space. Closed. A fixed set of primitives – click(element), type(element, text), scroll, goto(url), go_back. The Agent cannot add to this set.

Classic Agent: WebArena + BrowserGym. WebArena provides self-hosted, reproducible websites (e-commerce, forum, GitLab, CMS) so that results are comparable; BrowserGym standardises the observation/action interface so that different agents can be evaluated on the same environment. WebVoyager is the canonical end-to-end demonstration on the open web.

Harness Implementation.

Layer How it is realised here
Orchestration ReAct-style loop, but with an explicit grounding step before each action
Context & Memory Observation compression is the main cost driver: DOM pruning, AX-tree only, or screenshot + Set-of-Mark
Tools & Actions A fixed action vocabulary; parallel tabs for search-and-compare workflows
ACI & Environment Element grounding: mapping “the login button” to a concrete element id or (x,y)
Skills & Composition Site-specific recipes; skill synthesis (e.g. SkillWeaver) turns successful trajectories into reusable APIs
Protocol Browser drivers (CDP/Playwright), plus MCP servers for specific sites
Runtime & Safety Domain allow-lists, credential isolation, confirmation before irreversible actions

Practical Usage. Information gathering across many sites, form filling, price and inventory monitoring, and any task where no API exists. When an API does exist, prefer the Tool/API category – a browser Agent driving a UI is strictly more fragile than a typed function call.

Verification & Training. Final state assertions (did the order appear? did the row change?), plus LLM judges for open-ended tasks. Training is dominated by SFT on trajectories, with online RL on the simulated WebArena suite as a growing direction. Note the asymmetry with SWE: WebArena gives a simulated deterministic website, and that simulator is exactly what makes training possible at all.

Current Limits. Grounding remains the dominant error source – small or dynamic elements cause drift and hallucinated clicks. Long horizons amplify it, and unlike SWE there is usually no clean rollback.

4.5 Computer-Use / OS Agent

Observation Interface. Pixels. The Agent sees a screenshot of the whole desktop and must locate UI elements visually, across arbitrary applications. This is the highest-dimensional, least structured observation of the digital categories.

Action Space. Closed and low-level. click(x, y), type(text), key(combo), scroll, screenshot. Note that actions are addressed by coordinates, not by semantic element identifiers – which is precisely what makes grounding hard.

Classic Agent: OSWorld + Claude Computer Use. OSWorld (NeurIPS 2024) is the reference environment: a real Ubuntu desktop with browser, office applications, filesystem and terminal, with execution-based evaluation. Claude’s Computer Use was the first widely available productised implementation of the same interface, and established the screenshot-to-action loop as an industry pattern.

Harness Implementation.

Layer How it is realised here
Orchestration Tight perceive-act loop; limited planning, because each step is expensive
Context & Memory Screenshot history with keyframe retention; most past frames must be dropped
Tools & Actions Coordinate-level primitives, sometimes augmented with OCR or an AX tree
ACI & Environment Grounding: turning “the Save button” into (x, y); resolution and scaling normalisation
Skills & Composition Application-specific macros; replaying demonstrated workflows
Protocol VM/container orchestration, screen-capture and input-injection services
Runtime & Safety Full VM isolation, snapshot/rollback, strict egress control

Practical Usage. Automating legacy software with no API, cross-application workflows (copy from a PDF into a spreadsheet and email it), and desktop QA. It is the last resort category: if an API or a CLI exists, use it – OS Agents are strictly less reliable and far more expensive per step.

Verification & Training. State assertions on the filesystem or application state. Training combines SFT on human demonstrations with RL on simulators; grounding models are often trained separately as a vision task.

Current Limits. Efficiency is the headline problem: even when top agents succeed, they take roughly 1.4-2.7x as many steps as humans (OSWorld-Human, ICML 2025). Add latency, brittleness to layout changes, and the fact that every step is an irreversible side effect, and the safety burden is the highest among the digital categories.

Boundary note. Web and Computer-Use overlap: a browser is an application running on an OS. The distinction is which interface the Agent targets – the page’s structured representation, or the desktop’s pixels.

4.6 Deep Research Agent

Observation Interface. Web text, retrievable by query. It is the only category that is purely read-only and side-effect free.

Action Space. search(query), fetch(url), plus internal actions such as extract, summarise, compare and cite. There is no state to mutate.

Classic Agent. This category is unusual in that its canonical implementations are products rather than papers – OpenAI Deep Research and Gemini Deep Research defined what users expect the interaction to look like. On the research side, Search-R1 is the reference for treating retrieval as a trainable action inside an RL loop, and BrowseComp is the benchmark that made the category measurable.

Harness Implementation.

Layer How it is realised here
Orchestration Decompose into sub-questions, then iterate: search, read, cross-check, identify gaps, re-search
Context & Memory Document cache on disk; never put whole documents in context – return extracts or summaries
Tools & Actions Search, fetch, extract; heavy use of parallel search over independent sub-questions
ACI & Environment Search-result ranking, deduplication, source credibility scoring
Skills & Composition Sub-agents used for context isolation, not for collaboration
Protocol Search APIs, MCP servers for scholarly databases
Runtime & Safety Low risk (read-only); the real risks are fabricated citations and source bias

The subtlest design point is context isolation. Production systems split retrieval from writing: each sub-question gets its own context that digests pages and returns only conclusions plus citations, and the writer never sees raw pages. Without this split, a few rounds of full documents exhaust the window. This is the clearest real use of the Subagents pattern from 3.2.6 – and note that it is used for context management, not because the task needs multiple minds.

Practical Usage. Market and literature surveys, competitive analysis, “what is the current state of X” questions, and any task whose output is a cited report rather than a single answer.

Verification & Training. This is the weakest spot in the entire field. There is no test suite and no deterministic end state, so evaluation falls back on LLM-as-judge plus citation-faithfulness checks, both of which are unreliable. The consequence is structural: because there is no cheap ground truth, RLVR does not apply, and progress has been much slower than in SWE – even though this environment is far simpler. This is the cleanest illustration of the rule stated in 3.2.1: the available training method follows verifiability, not environment complexity.

Representative benchmarks: BrowseComp (questions deliberately hard to answer by direct search), BrowseComp-Plus (separates retrieval quality from reasoning quality), and MLR-Bench (open-ended research tasks).

Important distinction. There are two very different research agents here. Search agents (BrowseComp) find existing answers on the open web. Scientific agents (MLR-Bench) run experiments and therefore do have executable verification, which makes them closer to the SWE category than to BrowseComp. Do not treat them as one thing.

Current Limits. Gap detection – knowing what it still does not know – is the main failure mode; agents stop early and deliver a confident-looking report with holes. Fabricated or mismatched citations are the second. Both are made worse by the fact that neither is cheaply detectable.

4.7 Tool / API / Enterprise Agent

Observation Interface. Structured tool returns (usually JSON, sometimes XML). Observation is cheap and bounded – but see “large ecosystems” below.

Action Space. Closed and externally defined. The Agent can call the tools it is given. It cannot invent one.

This category has no environment of its own. Every other category has a world: SWE has a repository, Web has a page, OS has a desktop. Tool agents have none – they borrow whatever the tools reach. This is why the category is really a cross-cutting capability rather than an environment class, and it is also why it appears inside every other category.

That said, when tools are the main interface, three problems appear that are specific to this category.

Problem 1: the action space is defined by someone else. Because the Agent cannot create tools, all the engineering moves from making actions to finding, understanding and sequencing them:

1
2
3
Tool Discovery / Retrieval   ← too many tools to choose from
Schema Understanding ← parameters are structured and unforgiving
Dependency Orchestration ← wrong order wastes the whole run

Problem 2: the tool inventory itself is a context budget problem.

1
2
3
10 tools    → put them all in the system prompt
100 tools → selection accuracy drops, context is eaten
250+ tools → retrieval is mandatory (MCP-Bench uses 250 tools over 28 MCP servers)

The mechanism is identical to RAG: embed the tool descriptions, retrieve the top-k, and inject only those definitions. ToolScope (which combines tool merging, retrieval and context-aware filtering) is the research counterpart of this engineering pattern. Tool registries need the same care as document indexes – deduplicate, merge near-duplicates, and rewrite descriptions for retrievability.

Problem 3: real users do not name the tools. MCP-Bench addresses this directly by generating fuzzy task variants that omit tool names and execution steps, and by attaching ten distractor servers (100+ irrelevant tools) to every task. In production, inferring intent from an underspecified request is the common case, not the edge case.

Classic Agent. There is no single canonical implementation; the category is defined by its benchmarks instead. The lineage runs:

1
2
3
4
5
ReAct                  → the prototype: interleave reasoning with tool calls
Toolformer / Gorilla → learn when to call, and which tool to call
BFCL → becomes the de facto standard for function calling
τ-bench / τ²-bench → add a simulated user and policy compliance
MCP + MCP-Bench → from "a set of APIs" to "a tool ecosystem"

Harness Implementation.

Layer How it is realised here
Orchestration Plan-then-execute or multi-turn plan-act-observe; dependency-aware scheduling
Context & Memory Tool retrieval; observation compression (tool outputs can be long); state externalised
Tools & Actions Schema design, validation, retry, idempotency keys
ACI & Environment Error semantics: distinguishing a bad argument from a transient failure
Skills & Composition Tool composition into higher-level capabilities; agent-as-a-tool
Protocol MCP is the centre of gravity – discovery, auth, and governance live here
Runtime & Safety Permission scoping, audit logging, human confirmation for destructive calls

Practical Usage. Enterprise workflow automation: CRM updates, ticketing, scheduling, payments, and any SaaS operation with an API. This is the category with the most immediate commercial value, because the environment already exists and only needs wrapping.

Verification & Training. End-state assertions on a database, plus pass^k (did it succeed every time across k runs) – the right metric for anything touching money or records. τ-bench adds policy compliance: the tool permits a call that business rules forbid (refunding an already-refunded order), so the Agent must read the policy, not just the schema. This is where the Permission and Guardrail layers of 3.2.8 get their real justification.

Current Limits. Failure is dominated by the plumbing rather than by reasoning: wrong tool, wrong arguments, wrong order, missing dependency, duplicated non-idempotent call, or a hallucinated tool name. The research frontier has largely moved from model capability to interface standards – early work asked “can the model call a tool?”, and the current question is “the MCP ecosystem is huge, so how do we discover, authenticate and govern tools?”

Common failure modes and their engineering answers:

Failure Countermeasure
Wrong tool selected tool retrieval, merging duplicates, better descriptions
Schema violation strict validation plus one automatic retry
Wrong dependency order explicit DAG orchestration
Duplicate / non-idempotent call idempotency keys, write de-duplication
Hallucinated tool name allow-list validation
Out-of-scope operation permission sandbox, confirmation, audit log

4.8 What Changes Across Categories

Reading the categories side by side exposes four regularities. They are the reason this part exists.

1. Tool creation is the sharpest dividing line. Only the SWE Agent can build a new tool mid-task. Everywhere else the action space is closed, so the Agent can only select and compose. This is why skill-library and skill-synthesis research is driven mainly by Web and general agents – if you cannot create tools at runtime, you must manufacture reusable ones offline.

2. Observation cost determines action granularity. Text observations can be filtered (grep, head, truncation), so the SWE Agent can take semantically coarse actions. Screenshots cannot be filtered, so OS and Web agents are forced into fine-grained actions and therefore into long trajectories. The awkwardness of computer-use agents is not a modelling failure; it follows directly from the cost of their observations.

3. Verifiability determines which training method is available.

Category Signal Therefore
SWE test suite RLVR works; fastest RL progress
Web / OS state assertion simulated environments make training feasible
Tool / Enterprise end-state + policy pass^k and user simulation
Deep Research LLM judge weak rewards; slow progress

This is the same rule stated in 3.2.1, now with the evidence attached: training method follows verifiability, not environment complexity.

4. Irreversibility determines safety design. A repository has version control, so a bad edit is recoverable. A submitted form, an overwritten file, a sent email and a physical action are not. The Runtime & Safety layer of 3.2.8 is therefore not an optional add-on – it is derived from the environment’s reversibility, and it grows in importance along exactly the axis from SWE to Tool/Enterprise to OS to Embodied.

Generalisation is an evaluation axis, not a category. GAIA and AgentBench do not describe a kind of world; they test whether one Agent can operate across several worlds. That makes them a measure of transfer, and they belong with the cross-environment structures of 5.2.2 rather than in the list above.

5. Benchmark

5.1 What Does an Agent Benchmark Evaluate?

Agent benchmarks evaluate more than final-answer correctness. Depending on the benchmark, they may measure task success, long-horizon execution, tool use, environment interaction, efficiency, reliability, safety, and generalization.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
Agent Evaluation
│
├── Effectiveness
│ └── Did the task succeed?
│
├── Efficiency
│ ├── Steps / Actions
│ ├── Tokens
│ ├── Time
│ └── Cost
│
├── Reliability
│ └── Can the agent succeed repeatedly?
│
├── Safety
│ └── Did it violate constraints or permissions?
│
└── Generalization
└── Can it handle unseen tasks and environments?

The categories below follow directly from the environment classes defined in 3.2.1. Each environment class induces its own benchmark family; multi-agent and safety are cross-environment concerns rather than environment classes of their own.

5.2.1 Environment Classes and Their Benchmarks

Environment Class Representative Benchmarks
Classic / Text-World ALFWorld, WebShop
Software Engineering SWE-bench, SWE-bench Verified, SWE-Bench Pro
Tool / API / Enterprise BFCL, τ-bench / τ²-bench, TheAgentCompany, AgencyBench
Web / Browser WebArena, Mind2Web, TurkingBench, X-WebAgentBench
Computer-Use / OS OSWorld, OSWorld 2.0, OSWorld-Human, WindowsAgentArena
Mobile AndroidWorld, MobileAgentBench
Deep Research BrowseComp, BrowseComp-Plus, MLR-Bench
Generalist / Cross-Environment GAIA, CRAB
Embodied ALFRED, Habitat, BEHAVIOR

5.2.2 Cross-Environment Structures and Evaluation Axes

Category Kind Representative Benchmarks
Multi-Agent structural axis — crosses every environment class MultiAgentBench, AgentVerse-related benchmarks
Agent Safety / Security evaluation axis — crosses every environment class Agent safety / tool-use security benchmarks
Benchmark Methodology meta-evaluation BrowseComp-Plus and related benchmark methodology work

Note on Long-Horizon. Long-horizon execution is not an environment class — it is a property that every environment class can exhibit (SWE, Web, OS, and Deep Research all have long-horizon variants). It is therefore folded into each class above rather than listed as its own category.

5.2.3 Generalist Agent Benchmark

Generalist benchmarks evaluate whether an agent can combine reasoning, tool use, information gathering, and multi-step execution across heterogeneous tasks. GAIA is a representative benchmark in this category.

The agent workflow nowadays can look like:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
Goal
↓
Planning
↓
Tool calls
↓
Environment interaction
↓
State changes
↓
Observation
↓
Recovery
↓
More tool calls
↓
Final artifact

Traditional benchmarks often judge if the final answer is correct.

This research direction is about:

1
Can the Agent still complete the task after taking dozens, hundreds, or even thousands of actions?

5.2.4 Long-Horizon / Real-World Agent Benchmark

Long-horizon is a property, not an environment class (see the note in 5.2.2). This section is kept separate only because these benchmarks are usually discussed together, and because they are the clearest stress test of sustained execution. In terms of environment class they belong to Tool / API / Enterprise (TheAgentCompany, AgencyBench), and their distinguishing feature is the horizon, not the world they run in.

These benchmarks emphasize sustained execution over many interacting steps, often across realistic workplace or application environments.

  • TheAgentCompany, NeurIPS 2025

    This benchmark builds a company system:

    1
    2
    3
    4
    5
    6
    7
    8
    Company Environment
    │
    ├── Browser
    ├── Code repository
    ├── Terminal
    ├── Internal websites
    ├── Communication
    └── Coworkers
  • AgencyBench, ACL 2026

5.2.5 Computer-Use / GUI Agent Benchmark

1
Can an Agent really operate a computer instead of merely describing the operations?
  • OSWorld, NeurIPS 2024.

    The most important Computer-Use Benchmark.

    1
    2
    3
    4
    5
    6
    7
    8
    Agent
    │
    ├── Desktop OS Environment
    ├── Browser
    ├── Office Applications
    ├── File System
    ├── Terminal
    └── Desktop Applications
  • OSWorld-Human, ICML 2025 Workshop on Computer-Use Agents

    Calculate efficiency of Agent.

    Even when they successfully complete the task, top agents still require 1.4–2.7× as many steps as humans.

  • OSWorld 2.0

  • CRAB, ACL 2025

    Cross-Environment Agent Benchmark for Multimodal Language Model

    Thought: Benchmark should not be bound by certain environment.

5.2.6 Tool Use / Function Calling / MCP Benchmark

Tools Evaluation Process:

1
2
3
4
5
6
7
8
9
Function Calling
↓
Tool Use
↓
Stateful Tool Use
↓
Agentic Tool Use
↓
MCP / Interoperability
  • BFCL, ICML 2025

    BFCL v4 Leaderboard

    It judges:

    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    12
    13
    Single-turn
    ↓
    Multiple functions
    ↓
    Parallel calls
    ↓
    Multi-turn
    ↓
    Memory
    ↓
    Dynamic decision making
    ↓
    Long-horizon agentic behavior
  • ToolSandbox, NAACL 2025 Findings

  • τ-bench, ICLR 2025

  • MCP-Bench, ICLR 2026

    1
    2
    3
    4
    5
    6
    7
    8
    9
    10
    11
    Tool Discovery
    ↓
    Tool Selection
    ↓
    Parameter Control
    ↓
    Multi-hop Planning
    ↓
    Cross-tool Coordination
    ↓
    Task Completion

5.2.7 Software Engineering Agent Benchmark

  • SWE-bench, ICLR 2024 Oral

  • SWE-bench Multimodal, ICLR 2025

  • SWE-Bench Pro, ICML 2026.

    For Long-Horizon software engineering.

There’s a problem: Contamination.

The paper finds that some models’ performance on SWE-bench Verified may be affected by training data contamination:

Does SWE-Bench-Verified Test Agent Ability or Model Memory?

5.2.8 Web / Browser Agent Benchmark

The core of Web Agent is not “Search” but: The Agent observes webpages, plans, clicks, types, navigates, calls APIs, and ultimately completes tasks in real-world web environments.

Classic workflow:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
User Goal
↓
Browser
↓
Observe webpage
↓
Plan
↓
Click / Type / Search
↓
New webpage
↓
Reason
↓
...
↓
Task Success
  • WebArena, ICLR 2024.
  • TurkingBench, NAACL 2025 Long.
  • X-WebAgentBench, ACL 2025 Findings

5.2.9 Deep Research / Research Agent Benchmark

Normal Web Research:

1
2
3
4
5
6
7
8
9
Web Agent

Goal
↓
Browse
↓
Action
↓
Done

Deep Research:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
Research Question
↓
Decompose
↓
Search
↓
Read
↓
Cross-check
↓
Identify information gaps
↓
Search again
↓
Synthesize
↓
Cite sources
↓
Report
  • BrowseComp

  • BrowseComp-Plus, ACL 2026

  • MLR-Bench, NeurIPS 2025:

    1
    2
    3
    4
    5
    Web Research
    ↓
    Deep Research
    ↓
    Scientific Research Agent

    Can AI do research itself?

5.2.10 Multi-Agent Benchmark

Multi-agent benchmarks evaluate coordination, communication, task decomposition, role assignment, negotiation, and collaborative problem solving among multiple agents. This is a distinct evaluation axis from single-agent capability.

Multi-Agent is a structural axis, not an environment class (see 5.2.2). Its benchmark family is therefore orthogonal to the environment classes above.

To be added…

5.2.11 Agent Safety / Security Benchmark

Agent safety benchmarks evaluate whether agents remain within task, permission, and security constraints while interacting with tools, external data, and stateful environments. Important dimensions include prompt injection, unsafe tool use, excessive permissions, policy violations, and recovery from adversarial conditions.

5.3 Benchmark Methodology

Benchmark design itself is an important research direction. A useful benchmark should measure the intended capability without being dominated by contamination, brittle verifiers, uncontrolled environment changes, or hidden implementation details.

Important questions include:

1
2
3
4
5
6
7
8
9
10
11
Benchmark Quality
│
├── Contamination / Data Leakage
├── Reproducibility
├── Dynamic vs. Fixed Environment
├── Verifier Reliability
├── Difficulty Calibration
├── Human Baselines
├── Efficiency / Cost
├── Robustness
└── Agent-vs-Model Disentanglement

BrowseComp-Plus is a representative example of this direction, emphasizing controlled and reproducible evaluation for deep-research agents.

5.4 Benchmark Deep Dive: AgentBench

5.4.1 AgentBench

AgentBench (ICLR 2024) is a multi-environment benchmark suite, rather than a single environment-specific benchmark. It includes:

1
2
3
4
5
6
7
8
9
10
AgentBench
│
├── OS
├── Database
├── Knowledge Graph
├── Digital Card Game
├── Lateral Thinking Puzzle
├── ALFWorld (ICLR 2021)
├── WebShop (NeurIPS 2022)
└── Mind2Web (NeurIPS 2023)

Classic Agent Environments

ALFWorld, WebShop, and Mind2Web have an important historical role in LLM-agent research and should not be treated as merely implementation details of AgentBench. They are independently used environments/benchmarks that also appear within the AgentBench suite.

  • ALFWorld — interactive household tasks in a text-based embodied environment.
  • WebShop — simulated e-commerce tasks involving search, browsing, product selection, and purchase.
  • Mind2Web — web-agent tasks grounded in real-world websites and interaction trajectories.

This distinction avoids mixing a benchmark suite (AgentBench) with individual environments (ALFWorld/WebShop/Mind2Web) at the same conceptual level.

6. Multi-Agent

This part is still being written.

Multi-Agent is a structural axis, not an environment class: every category in Part 4 can be built as a single Agent or as several. Planned sections: 6.1 Why Multiple Agents, 6.2 Communication Topologies, 6.3 Frameworks, 6.4 Failure Modes and Cost.

本文链接:https://wangyier.top/panorama-of-agent/

版权声明:本博客所有文章除特别声明外,均采用 CC BY-NC-SA 4.0 许可协议。转载请注明来自 The Great Library!