This article is published in English.
Teams Compared 6 Python AI Agent Frameworks So You Don’t Have To: LangGraph vs
Operable walkthrough of Teams Compared 6 Python AI Agent Frameworks So You Don’t Have To: LangGraph vs CrewAI vs PydanticAI vs OpenAI SDK vs Smolagents vs Google AD: contracts,.
The following notes reconstruct a practical path around “one Compared 6 Python AI Agent Frameworks So You Don’t Have To: LangGraph vs CrewAI vs PydanticAI vs OpenAI SDK vs Smolagents vs Google ADK”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing.
you built the same research orchestrator six times. Only two of these stacks remained viable the weekend.
When working through you built the same research orchestrator six times. Only two of these stacks remained viable the weekend., write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
The Setup: What you Actually Built
When working through The Setup: What you Actually Built, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Framework 1: LangGraph — The Control Freak’s Paradise
When working through Framework 1: LangGraph — The Control Freak’s Paradise, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Framework 2: CrewAI — The Fast Prototype Machine
When working through Framework 2: CrewAI — The Fast Prototype Machine, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
researcher = Agent(
role="Financial Research Analyst",
goal="Find and verify recent financial data",
backstory="You're a senior analyst at a hedge fund...",
)
Framework 3: PydanticAI — The Quiet Overachiever
When working through Framework 3: PydanticAI — The Quiet Overachiever, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through Framework 3: PydanticAI — The Quiet Overachiever, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
agent = Agent(
"openai:gpt-4o",
result_type=CompanyAnalysis, # Pydantic model
system_prompt="You are a financial research assistant.",
)
Here’s Where Things Got Interesting
Here’s Where Things Got Interesting works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
Framework 4: OpenAI Agents SDK — The Sleeper Hit
Framework 4: OpenAI Agents SDK — The Sleeper Hit works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
agent = Agent(
name="Researcher",
instructions="You are a financial research assistant.",
tools=[search_tool, db_tool],
handoffs=[summary_agent],
)
Framework 5: Smolagents — The Open-Source Purist’s Dream
Framework 5: Smolagents — The Open-Source Purist’s Dream works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. Framework 5: Smolagents — The Open-Source Purist’s Dream works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
agent = CodeAgent(
tools=[search_tool, db_tool],
model=InferenceClientModel(),
)
result = agent.run("Analyze recent financial news for Acme Corp")
Framework 6: Google ADK — The Enterprise Sleeper
For Framework 6: Google ADK — The Enterprise Sleeper, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
from google.adk.agents import Agent
root_agent = Agent(
model="gemini-2.5-flash",
name="financial_analyst",
instruction="You are a financial research assistant.",
tools=[search_tool, db_tool],
)
The Verdict: It Depends (But Not the Way You Think)
For The Verdict: It Depends (But Not the Way You Think), define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
The Full Comparison Table
For The Full Comparison Table, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For The Full Comparison Table, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
| Metric | LangGraph | CrewAI | PydanticAI | OpenAI SDK | Smolagents | Google ADK |
| ----------------- | ----------- | --------------- | ------------- | ---------------- | --------------- | --------------- |
| Lines of code | ~210 | ~340 | ~130 | ~150 | ~95 | ~180 |
| Time to prototype | 3 hrs | 45 min | 1.5 hrs | 1 hr | 30 min | 2 hrs |
| Avg tokens/run | 2,847 | 4,216 | 2,912 | 2,791 | 3,340 | 3,102 |
| Multi-agent | Yes (graph) | Yes (teams) | Manual | Yes (handoffs) | Yes (hierarchy) | Yes (AgentTeam) |
| Type safety | TypedDict | Pydantic config | Full generics | Generic context | Minimal | Standard |
| MCP support | Yes | Limited | Native + A2A | Native | Yes | Yes |
| Model-agnostic | Yes | Yes | Yes (20+) | Yes (100+) | Yes (LiteLLM) | Gemini-first |
| Best debugger | LangSmith | Logs | IDE/types | Built-in tracing | Code output | ADK Web UI |
| GitHub stars | ~48K | ~44K | ~15K | ~16K | ~26K | ~23K |
| 2 AM debug | 9/10 | 5/10 | 8/10 | 7/10 | 8/10 | 6/10 |
The One Thing you Wish you Knew Before Starting
When working through The One Thing you Wish you Knew Before Starting, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
Operational checklist
For Operational checklist, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Favor boring reliability over clever one-off demos.
Batch note for d8a5e6e43262: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For hardening note 0, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 0/819: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through hardening note 1, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 1/819: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
hardening note 2 works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 2/819: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For hardening note 3, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 3/819: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through hardening note 4, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 4/819: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
hardening note 5 works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 5/819: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
Rewrite marker 1 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 2 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 3 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 4 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 5 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 6 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 7 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 8 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 9 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.
Rewrite marker 10 for d8a5e6e43262: rephrase surrounding claims with operator language, keep [[CODE_n]] slots intact, and avoid echoing source sentences.