Home / Articles / Practical notes: Agentic AI with LangChain — Part 3: Tool Calling in LangChain

This article is published in English.

Practical notes: Agentic AI with LangChain — Part 3: Tool Calling in LangChain

Operable walkthrough of Practical notes: Agentic AI with LangChain — Part 3: Tool Calling in LangChain: contracts, checks, and drop-in code slots for teams shipping this pattern.

1916 words

Use this as an operator-facing rebuild of the ideas in “Agentic AI with LangChain — Part 3: Tool Calling in LangChain”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Tool Calling Concepts

For the Tool Calling Concepts stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

1. System Configurations

For the 1 System Configurations stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

OPENAI_API_KEY="<Your OpenAI API Key>"
TAVILY_API_KEY=<Your TAVILY API Key>
class BaseConfig(BaseSettings):
    OPENAI_API_KEY: Optional[str]
    PINECONE_API_KEY: Optional[str]
    TAVILY_API_KEY: Optional[str]

model_config = SettingsConfigDict(env_file=".env", extra="ignore")

2. Project Structure

For the 2 Project Structure stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

project/
|
├── tools
│   ├── get_sum.py    # a simple LangChain tool example
|   └── weather.py    # get_weather LangChain tool using open source APIs
|
|── tool_call.py      # implement tool-calling loop using LangChain tools
├── config.py         # pydantic BaseConfig
├── .env              # environment variable definition
├── .gitignore
└── requirements.txt  # package requirements

3. LangChain Tools

For the 3 LangChain Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

3.1. What Is a LangChain Tool?

For the 3 1 What Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

from langchain.tools import tool


@tool
def get_sum(a: int, b: int) -> int:
    """
    get the summation of two integers
    :param a: int, input integer
    :param b: int, input integer
    :return: int, the sum of a and b
    """
    return a + b

3.2. Implement a Weather Tool

For the 3 2 Implement a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

3.3. Concept of Tool Calling

For the 3 3 Concept of stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

3.4. Tool Calling in LangChain

For the 3 4 Tool Calling stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the 3 4 Tool Calling stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

3.4.1. Setting Up the LLM and Tools

When working through the 3 4 1 Setting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

class WeatherAssistant:
    def __init__(self):

        # initialize llm
        self.llm = ChatOpenAI(api_key=api_key, model="gpt-4o-mini", temperature=0)

        # initialize a tool dictionary
        self.tools = {"get_weather": get_weather,
                      "tavily_search": TavilySearch(max_results=3, tavily_api_key=TAVILY_API_KEY)}
        # bind LangChain tools to llm
        self.llm_with_tools = self.llm.bind_tools(list(self.tools.values()))

        # initialize messages to store message list
        self.messages = []

        # System prompt
        self.system_prompt = f"""You are a helpful assistant for question-answering tasks.
        When users ask about weather, use the get_weather tool to get weather. For other questions,
        use web_search. If you don't know the answer, just say that you don't know.
        Be conversational and helpful in your responses."""

        self.messages.append(SystemMessage(content=self.system_prompt))

3.4.2. Implementing the Tool-Calling Loop

When working through the 3 4 2 Implementing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

async def chat(self, message: str):
    # Wrap User message in a HumanMessage and add it to message list
    self.messages.append(HumanMessage(content=message))

    # Get AI response (it may or may not contain tool calls)
    response = await self.llm_with_tools.ainvoke(self.messages)
    self.messages.append(response)

    # If there is any tool calls in the AI response
    if response.tool_calls:

        # process tool calls
        for tool_call in response.tool_calls:

            # retrieve function name from tool_call, then
            # retrieve the tool from tool dictionary, and invoke it,
            # append resulting tool message to message list
            tool = self.tools[tool_call["name"]]
            tool_result = await tool.ainvoke(tool_call)
            self.messages.append(tool_result)

        # Get final response after tool execution
        final_response = await self.llm_with_tools.ainvoke(self.messages)
        self.messages.append(final_response)

3.4.3 Testing the Loop

When working through the 3 4 3 Testing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the 3 4 3 Testing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

async def main():
    print("hello tool calling!")
    assistant = WeatherAssistant()
    message = "What is the temperature in Tokyo?"
    await assistant.chat(message)
    for msg in assistant.messages:
        msg.pretty_print()
if __name__ == "__main__":
    asyncio.run(main())

4. Limitations of the Basic Tool-Calling Loop

The 4 Limitations of the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

5. Closing Summary

The 5 Closing Summary stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for d7ca1ebeb899: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.