Home / Articles / ReAct from scratch: thought, action, pause, observation without a framework

This article is published in English.

ReAct from scratch: thought, action, pause, observation without a framework

Build a tiny Groq ReAct loop by hand—manual observations first, then regex-driven tool calls—to see why modern agent frameworks feel the way they do.

1373 words

Introduction

Ask a plain language model for Nvidia’s live price and how many shares $100,000 buys. Without tools it invents a stale quote from training data, then divides correctly on wrong inputs. It has no built-in move that means “I do not know—look it up,” because lookup is not a single forward pass.

As response quality rose, models proved useful for one-shot reasoning. Planning alone still does not grant autonomy. The questions that define agents are: what is outside model knowledge, which action acquires it, how to produce a result, how to revise the plan, and when to stop. Decision → act → check → replan is a loop, not a straight line. That loop is the origin of modern AI agents. ReAct (Reason + Act) set the pattern. The original demos were not LangChain apps—they were an LLM, a structured prompt, and a human (or later a script) executing actions and feeding observations back.

Enter ReAct: Reasoning + Act

A worked example, step by step:

from groq import Groq
import re
from dotenv import load_dotenv

_ = load_dotenv()
client = Groq()
message = client.chat.completions.create(
    model="openai/gpt-oss-120b",  # free, fast open-weight model hosted on Groq
    max_tokens=1000,
    messages=[
        {"role": "user", "content": "Hello, GPT!"}
    ],
)
print(message.choices[0].message.content) # Check the client
class Agent:
    def __init__(self, system=""):
        self.system = system
        self.messages = []
        if self.system:
            self.messages.append({"role": "system", "content": system})

    def __call__(self, message):
        self.messages.append({"role": "user", "content": message})
        result = self.execute()
        self.messages.append({"role": "assistant", "content": result})
        return result

    def execute(self):
        response = client.chat.completions.create(
            model="openai/gpt-oss-120b",
            max_tokens=1000,
            messages=self.messages,
        ).choices[0].message.content
        return response
prompt = """
You run in a loop of Thought, Action, PAUSE, Observation.
At the end of the loop you output an Answer
Use Thought to describe your thoughts about the question you have been asked.
Use Action to run one of the actions available to you - then return PAUSE.
Observation will be the result of running those actions.

Your available actions are:

calculate:
e.g. calculate: 4 * 7 / 3
Runs a calculation and returns the number - uses Python so be sure to use floating point syntax if necessary

fish_weight:
e.g. fish_weight: Shark
returns weight of a fish when given the breed

Example session:

Question: How much does a shark weigh?
Thought: I should look the fish weight using fish_weight
Action: fish_weight: Shark
PAUSE

You will be called again with this:

Observation: A Great white shark weights 41000 lbs

You then output:

Answer: A Great white shark weights 41000 lbs
""".strip()
def calculate(expression):
    return eval(expression)

def fish_weight(name):
    if "Shark" in name:
        return("Great white shark weighs 41000 lbs")
    elif "Puffer" in name:
        return("A puffer fish weighs 20 lbs")
    else:
        return("A fish can weight upto 47000 lbs")

known_actions = {
    "calculate": calculate,
    "fish_weight": fish_weight
}
abot = Agent(prompt)
result = abot("How much does a Shark weigh?")
print(result)
result = fish_weight("Shark")
result
next_prompt = "Observation: {}".format(result)
abot(next_prompt)
abot.messages # View the list of messages containing the conversation
# Present a compund query
abot = Agent(prompt)
# New query
question = """I have 2 fishes, a Shark and a puffer fish. \
What is their combined weight"""
abot(question)
next_prompt = "Observation: {}".format(fish_weight("Shark"))
print(next_prompt)
abot(next_prompt)
next_prompt = "Observation: {}".format(fish_weight("Puffer fish"))
print(next_prompt)
abot(next_prompt)
next_prompt = "Observation: {}".format(eval("41000 + 20"))
print(next_prompt)
abot(next_prompt)
## Add loop (automate the reasoning + act process)
action_re = re.compile('^Action: (\w+): (.*)
) # python regular expression to selection action
def query(question, max_turns=5):
    i = 0
    bot = Agent(prompt)
    next_prompt = question
    while i < max_turns:
        i += 1
        result = bot(next_prompt)
        print(result)
        actions = [
            action_re.match(a)
            for a in result.split('\n')
            if action_re.match(a)
        ]
        if actions:
            # There is an action to run
            action, action_input = actions[0].groups()
            if action not in known_actions:
                raise Exception("Unknown action: {}: {}".format(action, action_input))
            print(" -- running {} {}".format(action, action_input))
            observation = known_actions[action](action_input)
            print("Observation:", observation)
            next_prompt = "Observation: {}".format(observation)
        else:
            return
question = """I have 2 fishes, a Shark and a puffer fish. \
What is their combined weight"""
query(question)
  1. Import the Groq client, re for pattern matching, and load_dotenv; load env so the API key is not hardcoded.
  2. Construct a Groq client that reads GROQ_API_KEY from the environment.
  3. Send a trivial “Hello” to a configured model to confirm connectivity.
  4. Build an Agent that keeps a system prompt and growing message history: each call appends the user turn, sends the full history, appends the model reply, and returns the text.
  5. Write the system prompt that enforces Thought → Action → PAUSE → Observation.
  6. Register tools such as calculate() for arithmetic and fish_weight() for a tiny lookup table in known_actions.
  7. Instantiate the agent with the ReAct system prompt.
  8. Ask “How much does a shark weigh?” Expect a proposed Action: fish_weight: Shark rather than a bare number.
  9. Manually run fish_weight("Shark") and keep the result (manual mode first).
  10. Feed Observation: … back so the agent can answer.
  11. Inspect abot.messages to see the full transcript the model sees.
  12. Start a fresh agent with a compound question (shark + puffer fish weights). 13–15. Manually supply each observation, including the summed total if needed.
  13. Let the agent synthesize the combined answer once observations exist.
  14. Add a regex that detects Action: name: input automatically.
  15. Wrap the loop in query(): new agent, send question, print output, parse action, validate against known_actions, execute, observe, repeat until final answer or turn limit.
  16. Run query() on the compound question and watch Thought → Action → Observation → Answer without manual steps.

Conclusion

Agents begin with a loop, not a framework. Chain-of-thought taught models to reason aloud; ReAct structured think–act–observe the way people work. Popular agent frameworks industrialize that idea. Building the loop by hand once permanently changes how frameworks feel. The original paper is at https://arxiv.org/abs/2210.03629.

The industrial lesson is that frameworks hide the loop behind runners and tool nodes, but the contract remains identical: the model must emit a parseable action, the runtime must execute only allowlisted tools, and observations must re-enter the transcript as first-class messages. Skip any of those and you have a chatbot with side effects, not an agent.

The industrial lesson is that frameworks hide the loop behind runners and tool nodes, but the contract remains identical: the model must emit a parseable action, the runtime must execute only allowlisted tools, and observations must re-enter the transcript as first-class messages. Skip any of those and you have a chatbot with side effects, not an agent.

The industrial lesson is that frameworks hide the loop behind runners and tool nodes, but the contract remains identical: the model must emit a parseable action, the runtime must execute only allowlisted tools, and observations must re-enter the transcript as first-class messages. Skip any of those and you have a chatbot with side effects, not an agent.

The industrial lesson is that frameworks hide the loop behind runners and tool nodes, but the contract remains identical: the model must emit a parseable action, the runtime must execute only allowlisted tools, and observations must re-enter the transcript as first-class messages. Skip any of those and you have a chatbot with side effects, not an agent.