Paper Summary – HuggingGPT
1. Introduction
This paper is published in year 2023. At that time, the conception of the Agent of nowadays is immature. This paper could be seen as the origin (~2600 citations, NIPS 2023) of the modern Multimodal Agent though it utilizes a text-only LLM instead of a VLM(Vision Language Model) to empower the agent.
This paper aims to solve the problems:
LLM is limited to the input and output forms of text generation, lacking the ability to process the multimodal tasks such as vision and speech.
In real-world scenarios, complex tasks are usually composed of sub-tasks, and thus require the scheduling and cooperation of multiple models, which are also beyond the capability of language models.
[!NOTE]
Notice: the models mentioned in the CV area, can be seen as individual tools completing specific tasks.
For some challenging tasks, the LLM performs weaker than experts.
To solve these problems, the LLM should be able to coordinate with external models. So the problem is :
- How to choose a middleware between LLM and AI Models (CV Tools).
An available method is to inject the description of those models into prompt of the LLM. But the cost of collecting suitable description is very huge. But thankfully, there are many existing descriptions on Hugging Face.


2. HuggingGPT
HuggingGPT is a collaborative system for solving AI tasks, composed of:
- a large language model (LLM)
- numerous expert models from ML communities.

2.1 Stage 1: Task Planning
Here’s a task planning prompt design consisting of 2 parts:
- specification-based instruction
- demonstration-based parsing
To better represent the expected tasks of user requests and use them in the subsequent stages, the authors expect the LLM to parse tasks by adhering to specific specifications (e.g., JSON format).
The authors design a standardized template for tasks and instruct the LLM to conduct task parsing through slot filing:



To better understand the intention and criteria for task planning, HuggingGPT incorporates multiple demonstrations in the prompt. Each demonstration consists of:
- A user request
- Its corresponding output, which represents the expected sequence of parsed tasks.
To support more complex scenarios (e.g., multi-turn dialogues), the authors include chat logs in the prompt by appending the following instruction:
1 | “To assist with task planning, the chat history is available as {{ Chat Logs }}, where you can trace the user-mentioned resources and incorporate them into the task planning.” |
2.2 Model Selection
After task planning, the HuggingGPT should match tasks with models:
- Selecting the most appropriate model for each task in the parsed task list
This stage uses model descriptions as the language interface to connect each model.
- Gather the descriptions of expert models from the ML community (e.g., Hugging Face).
- Employ a dynamic in-context task-model assignment mechanism to choose models for the tasks.
In-context Task-model Assignment:
Available models are presented as options within a given context.
However, due to the limits of maximum context length, it is not feasible to encompass the information of all relevant models within one prompt.
It’s not feasible to encompass the information of all relevant models within one prompt.
So:
- First, filter out models based on their task type to select the ones that match the current task.
- Rank them based on the number of downloads.
- Select the top-K models as the candidates.

2.3 Task Execution
After assigning models for subtasks, the next step is to execute the task.
However, there’s a resource dependency problem: Since the outputs of the prerequisite tasks are dynamically produced, HuggingGPT also needs to dynamically specify the dependent resources for the task before launching it.
During the task planning stage, if some tasks are dependent on the outputs of previously executed tasks (e.g., task_id), HuggingGPT sets this symbol (i.e., <resource>-task_id) to the corresponding resource subfield in the arguments.
2.4 Response Generation
In stage 4, HuggingGPT integrates all the information from the previous three stages:
- Task Planning
- Model Selection
- Task Execution
into a concise summary in this stage, including the list of planned tasks, the selected models for the tasks, and the inference results of the models.
Most important among them are the inference results, which are the key points for HuggingGPT to make the final decisions.

3. Conclusion
Workflow:
