This article is published in English.
Practical notes: Upgrade Your Deep Agent With a Local Open-Source Sandbox
Operable walkthrough of Practical notes: Upgrade Your Deep Agent With a Local Open-Source Sandbox: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Upgrade Your Deep Agent With a Local Open-Source Sandbox”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
What a sandbox is, and why agents need one
For the What a sandbox is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The catch
For the The catch stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Enter OpenSandbox
For the Enter OpenSandbox stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Enter OpenSandbox stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Deep Agents sandbox extension point
When working through the Deep Agents sandbox extension stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
class BaseSandbox(ABC):
def execute(self, command: str, *, timeout: int | None = None) -> ExecuteResponse: ...
@property
def id(self) -> str: ...
def upload_files(self, files: list[tuple[str, bytes]]) -> list[FileUploadResponse]: ...
def download_files(self, paths: list[str]) -> list[FileDownloadResponse]: ...
Building the integration
When working through the Building the integration stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
How OpenSandbox work?
When working through the How OpenSandbox work stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the How OpenSandbox work stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
# Generate a starter config
uvx opensandbox-server init-config ~/.sandbox.toml --example docker
# Start the server
uvx opensandbox-server
Simplified integration code
The Simplified integration code stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
.create()
The create stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
class MinimalOpenSandboxBackend(BaseSandbox):
def __init__(self, sandbox: Sandbox, runner: AsyncRunner):
self._sandbox = sandbox
self._runner = runner
from opensandbox import Sandbox
from opensandbox.config import ConnectionConfig
IMAGE = "opensandbox/code-interpreter:v1.1.0"
ENTRYPOINT = ["/opt/code-interpreter/code-interpreter.sh"]
@classmethod
def create(cls, api_key: str | None = None) -> "MinimalOpenSandboxBackend":
runner = AsyncRunner()
config = ConnectionConfig(domain="localhost:8080", api_key=api_key)
sandbox = runner.run(
Sandbox.create(IMAGE, entrypoint=ENTRYPOINT, connection_config=config, timeout=timedelta(minutes=30))
)
return cls(sandbox, runner)
.id()
The this stage stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The this stage stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
@property
def id(self) -> str:
return self._sandbox.id
.execute()
For the execute stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
from deepagents.backends.protocol import ExecuteResponse
def execute(self, command: str, *, timeout: int | None = None) -> ExecuteResponse:
execution = self._runner.run(self._sandbox.commands.run(command))
stdout = "\n".join(c.text for c in execution.logs.stdout)
stderr = "\n".join(c.text for c in execution.logs.stderr)
output = "\n".join(p for p in (stdout, stderr) if p)
return ExecuteResponse(output=output, exit_code=execution.exit_code or 0)
.upload_files()
For the uploadfiles stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
from opensandbox.models import WriteEntry
from deepagents.backends.protocol import FileUploadResponse
def upload_files(self, files: list[tuple[str, bytes]]) -> list[FileUploadResponse]:
entries = [WriteEntry(path=path, data=data, mode=644) for path, data in files]
try:
self._runner.run(self._sandbox.files.write_files(entries))
return [FileUploadResponse(path=p) for p, _ in files]
except Exception as exc:
return [FileUploadResponse(path=p, error=str(exc)) for p, _ in files]
.download_files()
For the downloadfiles stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
from deepagents.backends.protocol import FileDownloadResponse
def download_files(self, paths: list[str]) -> list[FileDownloadResponse]:
results = []
for path in paths:
try:
content = self._runner.run(self._sandbox.files.read_bytes(path))
results.append(FileDownloadResponse(path=path, content=content))
except Exception as exc:
results.append(FileDownloadResponse(path=path, error=str(exc)))
return results
For the downloadfiles stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
.kill()
When working through the kill stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
def kill(self) -> None:
self._runner.run(self._sandbox.kill())
self._runner.shutdown()
The sync/async bridge
When working through the The sync async bridge stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
A data analysis agent in the sandbox
When working through the A data analysis agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the A data analysis agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
import asyncio
import nest_asyncio
import threading
from datetime import timedelta
from pathlib import Path
from deepagents import create_deep_agent
from deepagents.backends.protocol import ExecuteResponse, FileDownloadResponse, FileUploadResponse
from deepagents.backends.sandbox import BaseSandbox
from langchain.chat_models import init_chat_model
from opensandbox import Sandbox
from opensandbox.config import ConnectionConfig
from opensandbox.models import WriteEntry
# nest_asyncio for running async functions in Jupyter.
nest_asyncio.apply()
IMAGE = "opensandbox/code-interpreter:v1.1.0"
ENTRYPOINT = ["/opt/code-interpreter/code-interpreter.sh"]
backend = MinimalOpenSandboxBackend.create(api_key="SANDBOX_API_KEY")
print("Sandbox ready:", backend.id)
llm = init_chat_model(
model="gemini-3.5-flash",
model_provider="google_genai",
api_key=os.environ["GOOGLE_API_KEY"],
max_tokens=14750,
max_retries=5,
)
agent = create_deep_agent(
model=llm,
system_prompt=(
"You are a Python coding assistant with sandbox access. "
"You specialize in performing data analysis and data visualization with python,"
"you generate clear reports with good looking charts using seaborn."
),
backend=backend,
)
csv_bytes = Path("customers-1000.csv").read_bytes()
results = backend.upload_files([("/workspace/customers-1000.csv", csv_bytes)])
for r in results:
if r.error:
print(f"Upload failed for {r.path}: {r.error}")
else:
print(f"Uploaded {r.path}")
result = agent.invoke({
"messages": "Perform a deep exploratory data analysis on the customers-1000.csv file "
"and summarize the findings in a markdown report with clear charts."
})
# Deep Exploratory Data Analysis: Customer Acquisition and Profiling
**Dataset:** `customers-1000.csv`
**Analysis Period:** Jan 2020 – May 2022
---
## 1. Executive Summary
This report presents a comprehensive exploratory data analysis (EDA) of a customer database containing 1,000 unique records. The analysis delves into geographical distributions, sign-up temporal trends, domain & technical alignments, and name demographics to uncover actionable insights for strategic growth.
### Key Takeaways
1. **Unprecedented Global Reach:** The customer base is extraordinarily decentralized, spanning **240 countries** across all **7 continents** (including Antarctica). No single country represents more than 1.2% of the customer base. Africa (24.7%) and Asia (22.6%) are the leading regions, followed by Europe (18.5%) and North America (16.1%).
2. **Stable Acquisition Trends:** Customer subscriptions are remarkably stable, averaging roughly **34-35 new customers per month** across 2020 and 2021. This consistency is maintained across all continents year-over-year, indicating a highly standardized, globally distributed customer acquisition channel.
3. **Mid-Week and Weekend Consistency:** Subscriptions are evenly spread across the days of the week, with a minor peak on Friday and Saturday, and a minor trough on Thursday.
4. **B2B / Synthetic Profile Characteristics:** The dataset shows zero domain overlap between customer email domains and company websites (0.0% exact match across 923 unique domains). Combined with the near 1-to-1 ratio of customers to companies, this suggests a highly B2B-centric profile (one representative per enterprise) or synthetically generated profiles with randomized fields.
5. **Standardized TLD Footprint:** The `.com` top-level domain (TLD) dominates both emails (61.2%) and corporate websites (58.8%). The remaining distribution is evenly split among `.org`, `.net`, `.biz`, and `.info`.
---
## 2. Dataset Structure & Data Integrity
The initial dataset contains **1,000 rows** and **12 columns**. An inspection of data integrity reveals excellent completeness:
- **Zero Missing Values:** Every column is 100% populated.
- **Zero Duplicates:** There are no duplicate rows, and the `Customer Id` column contains 1,000 unique identifiers.
- **Data Types:** All columns are stored as object/string types except for `Index` (integer).
### Data Preprocessing & Feature Engineering
To enable deep exploratory analysis, several features were engineered:
1. **Temporal Features:** `Subscription Date` was parsed as a datetime object, allowing the extraction of `Sub_Year`, `Sub_Month`, `Sub_Month_Name`, `Sub_Day_of_Week`, and `Sub_Year_Month` (period).
2. **Geographical Mapping:** Using the `pycountry` and `pycountry-convert` libraries, coupled with a manual fallback dictionary for territories, each of the 240 countries was successfully mapped to its respective **Continent**.
3. **Domain & Technical Profiles:** Email domains (`Email_Domain`), email TLDs (`Email_TLD`), and website TLDs (`Website_TLD`) were extracted to analyze the technical profiling of users.
---
## 3. Geographical Analysis
### Continent-Level Distribution
The geographic reach of this customer base is truly global. Rather than being concentrated in a single dominant market like North America or Europe, customers are spread across all continents:
| Continent | Customer Count | Percentage |
| :--- | :---: | :---: |
| **Africa** | 247 | 24.7% |
| **Asia** | 226 | 22.6% |
| **Europe** | 185 | 18.5% |
| **North America** | 161 | 16.1% |
| **Oceania** | 107 | 10.7% |
| **South America** | 54 | 5.4% |
| **Antarctica** | 20 | 2.0% |
#### Chart 1: Customer Distribution by Continent

### Country-Level Distribution (Top 15 Countries)
The country-level distribution exhibits a heavy tail, with the 1,000 customers distributed across **240 distinct nations**. This indicates that the average number of customers per country is only **4.17**.
The top countries by customer density are:
- **Liechtenstein:** 12 customers (1.2%)
- **Gabon:** 10 customers (1.0%)
- **China, Bangladesh, Reunion, Nigeria, Luxembourg:** 9 customers each (0.9%)
This extreme dispersion suggests a borderless, digital-first product that appeals universally across jurisdictions without localized geographic bias.
#### Chart 2: Top 15 Countries by Customer Count

---
## 4. Temporal Analysis (Subscription Trends)
... (Trimmed to keep blog (estimated read time short))
From notebook to a PyPi package
The From notebook to a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
!pip install deepagents-opensandbox-backend
Check it out
The Check it out stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 43662eb4f13d: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 0/814: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.