首页 / 文章 / 从零开始构建ReAct:无框架的思考、行动、暂停与观察

从零开始构建ReAct:无框架的思考、行动、暂停与观察

手动构建一个极小的 Groq ReAct 循环——先进行人工观察,再通过正则表达式触发工具调用——以此了解现代智能体框架为何会呈现这样的特性。

1373 词

简介

可以询问通俗语言模型Nvidia的实时股价,以及10万美元能购买多少股股票。在没有工具辅助的情况下,它会从训练数据中找出过时的报价,然后对错误的输入进行错误的分算。它没有内置“我不知道——请查询”的功能,因为查询并非一次简单的向前传播过程。

随着响应质量的提升,这些模型被证明适用于一次性推理任务。但仅靠规划仍无法实现自主性。定义智能体的关键问题包括:模型知识之外的是什么、何种行动能够获取这些知识、如何产生结果、如何调整计划以及何时停止。决策→行动→检查→重新规划构成一个循环,而非直线过程。正是这个循环催生了现代人工智能智能体。ReAct(推理+行动)模式为此奠定了基础。最初的演示并非LangChain应用——它实际上是由一个大型语言模型、结构化的提示语,以及执行动作并反馈观察结果的人类(或后来的脚本)共同构成的。

ReAct登场:推理与行动的结合

以下是一个逐步演示的实际案例:

from groq import Groq
import re
from dotenv import load_dotenv

_ = load_dotenv()
client = Groq()
message = client.chat.completions.create(
    model="openai/gpt-oss-120b",  # free, fast open-weight model hosted on Groq
    max_tokens=1000,
    messages=[
        {"role": "user", "content": "Hello, GPT!"}
    ],
)
print(message.choices[0].message.content) # Check the client
class Agent:
    def __init__(self, system=""):
        self.system = system
        self.messages = []
        if self.system:
            self.messages.append({"role": "system", "content": system})

    def __call__(self, message):
        self.messages.append({"role": "user", "content": message})
        result = self.execute()
        self.messages.append({"role": "assistant", "content": result})
        return result

    def execute(self):
        response = client.chat.completions.create(
            model="openai/gpt-oss-120b",
            max_tokens=1000,
            messages=self.messages,
        ).choices[0].message.content
        return response
prompt = """
You run in a loop of Thought, Action, PAUSE, Observation.
At the end of the loop you output an Answer
Use Thought to describe your thoughts about the question you have been asked.
Use Action to run one of the actions available to you - then return PAUSE.
Observation will be the result of running those actions.

Your available actions are:

calculate:
e.g. calculate: 4 * 7 / 3
Runs a calculation and returns the number - uses Python so be sure to use floating point syntax if necessary

fish_weight:
e.g. fish_weight: Shark
returns weight of a fish when given the breed

Example session:

Question: How much does a shark weigh?
Thought: I should look the fish weight using fish_weight
Action: fish_weight: Shark
PAUSE

You will be called again with this:

Observation: A Great white shark weights 41000 lbs

You then output:

Answer: A Great white shark weights 41000 lbs
""".strip()
def calculate(expression):
    return eval(expression)

def fish_weight(name):
    if "Shark" in name:
        return("Great white shark weighs 41000 lbs")
    elif "Puffer" in name:
        return("A puffer fish weighs 20 lbs")
    else:
        return("A fish can weight upto 47000 lbs")

known_actions = {
    "calculate": calculate,
    "fish_weight": fish_weight
}
abot = Agent(prompt)
result = abot("How much does a Shark weigh?")
print(result)
result = fish_weight("Shark")
result
next_prompt = "Observation: {}".format(result)
abot(next_prompt)
abot.messages # View the list of messages containing the conversation
# Present a compund query
abot = Agent(prompt)
# New query
question = """I have 2 fishes, a Shark and a puffer fish. \
What is their combined weight"""
abot(question)
next_prompt = "Observation: {}".format(fish_weight("Shark"))
print(next_prompt)
abot(next_prompt)
next_prompt = "Observation: {}".format(fish_weight("Puffer fish"))
print(next_prompt)
abot(next_prompt)
next_prompt = "Observation: {}".format(eval("41000 + 20"))
print(next_prompt)
abot(next_prompt)
## Add loop (automate the reasoning + act process)
action_re = re.compile('^Action: (\w+): (.*)
) # python regular expression to selection action
def query(question, max_turns=5):
    i = 0
    bot = Agent(prompt)
    next_prompt = question
    while i < max_turns:
        i += 1
        result = bot(next_prompt)
        print(result)
        actions = [
            action_re.match(a)
            for a in result.split('\n')
            if action_re.match(a)
        ]
        if actions:
            # There is an action to run
            action, action_input = actions[0].groups()
            if action not in known_actions:
                raise Exception("Unknown action: {}: {}".format(action, action_input))
            print(" -- running {} {}".format(action, action_input))
            observation = known_actions[action](action_input)
            print("Observation:", observation)
            next_prompt = "Observation: {}".format(observation)
        else:
            return
question = """I have 2 fishes, a Shark and a puffer fish. \
What is their combined weight"""
query(question)
  1. 导入Groq客户端、用于模式匹配的re函数以及load_dotenv函数;加载环境变量,以避免将API密钥硬编码在其中。
  2. 构建一个能够从环境变量中读取GROQ_API_KEY的Groq客户端。
  3. 向已配置的模型发送简单的“Hello”消息,以确认连接正常。
  4. 创建一个Agent,用于保存系统提示语及不断增长的消息历史记录:每次调用时都会添加用户的回复内容,发送完整的对话历史,再添加模型的回复,最后返回处理后的文本。
  • 编写强制遵循“思考→行动→暂停→观察”流程的系统提示词。
  • 在known_actions中注册诸如用于算术运算的calculate()以及用于小型查找表的fish_weight()等工具。
  • 使用ReAct系统提示词实例化智能体。
  • 提问“鲨鱼的重量是多少?”,预期得到的应是Action: fish_weight: Shark这样的行动指令,而非单纯的数字。
  • 先以手动模式运行fish_weight("Shark")并记录结果。
  • 将Observation: …反馈给智能体,使其能够给出答案。
  • 查看abot.messages以了解模型所看到的完整对话记录。
  • 通过复合问题(鲨鱼与河豚的重量)启动一个新的智能体。 13–15. 手动提供每一条观察结果,必要时还包括总和。
  • 一旦有观测结果,就让智能体合成综合答案。
  • 添加一个正则表达式,自动识别Action: name: input格式。
  • 将整个循环封装在query()函数中:创建新智能体,发送问题,打印输出结果,解析操作指令,与known_actions列表进行验证,执行操作并观测结果,重复此过程直至得到最终答案或达到轮次限制。
  • 对复合问题调用query()函数,无需手动操作即可查看“思考→行动→观测→答案”的完整流程。
  • 结论

    智能体的实现始于循环结构,而非框架。链式思考方法教会模型大声推理;ReAct则按照人类的工作方式将思考、行动和观测串联起来。常见的智能体框架便是将这一理念工业化应用的结果。一旦亲手构建循环结构,就会永久改变对框架的使用感受。相关原始论文地址为https://arxiv.org/abs/2210.03629。

    行业中的经验是,框架会将运行器和工具节点背后的循环隐藏起来,但核心规则依然不变:模型必须输出可解析的动作,运行时只能执行允许使用的工具,而观察结果必须以一级消息的形式重新进入对话记录。只要缺少其中任何一点,得到的就只是会产生副作用的聊天机器人,而非真正的智能体。

    行业中的经验是,框架会将运行器和工具节点背后的循环隐藏起来,但核心规则依然不变:模型必须输出可解析的动作,运行时只能执行允许使用的工具,而观察结果必须以一级消息的形式重新进入对话记录。只要缺少其中任何一点,得到的就只是会产生副作用的聊天机器人,而非真正的智能体。

    行业经验表明,框架虽然将运行器和工具节点背后的循环隐藏起来,但核心规则依然不变:模型必须输出可解析的动作,运行时只能执行允许使用的工具,而观察结果必须以一级消息的形式重新进入对话记录。只要缺少其中任何一点,得到的就只是具有副作用的聊天机器人,而非真正的智能体。

    行业经验表明,框架虽然将运行器和工具节点背后的循环隐藏起来,但核心规则依然不变:模型必须输出可解析的动作,运行时只能执行允许使用的工具,而观察结果必须以一级消息的形式重新进入对话记录。只要缺少其中任何一点,得到的就只是具有副作用的聊天机器人,而非真正的智能体。