Практычныя прытамулі: Створыце свой першы агент AI для DevOps за дапамою LangChain + інструментаў Bash
Практычныя нарады: Как створыць свага першага агента AI для DevOps за дапамою LangChain + інструментаў Bash: контракты, перакрыцчы і блокі коду для команд, якія викорыстоўваюць гэты патэрн.
У гэтым карыце практычным напамінанні перакладзенаецца шлях ад сыр'ёў да рабочай системы для проекту «Стварэнне вашага першага AI-агента для DevOps за дапамою LangChain + інструментаў Bash». Акцэнт ставіцца на практычныя крокі, чысткія пераконтанні і код, які можна проста дадаць у репазітарый без неабясненых спадчыных параметраў. Авар у стадію агульнага відзору павінен з'явіць вхідныя даны, апрацоўчыка кроку і критэрыя завершэння працы перш чым змяніць код. Апрацоўчыкі павінны магчымае перадзванаць крок з вядомай точкі контролю без неабясненняя скрытых станоў. Неабходна адразу задокументаваць як шлях успеху, так і шлях вяснавання проблем. Перапрыбуткі, людзкія пераконтанні і обробка некоректных паведамленняў є часткай продукту, а не яго пазнейшай дапрацоўкі.
Што мы насправды ствараем?
Калі працюеце над стадзіяй «Што мы будуем ствараць», спачатку запісайце умовы кантракту: неабяжлівыя даны, сігнал успеху і тое, што выходзіць у разе частковага неудачы. Такі список контроля дапамагае заліцварыць пазнейшыя змены ў кодзе. Валіце маленькія, тэставаныя елементы замест большых скрыптов. Калі якісь крок не выйшае, неудача должна вказваць на адну конкрэтную адпаведальнасць, а не на заплутаны ланцюг задач. Запісвайце назву інструмента, хэш аргументаў, час адклікання і рэзультат кожнага вызову. Без такога следу дэбагаванне агента займае гадзіны.
Chatbot: "Your nginx config might have a syntax error."
Agent: [runs `nginx -t`] → "Confirmed. Line 42 in /etc/nginx/sites-enabled/app.conf
has an unclosed bracket. Here's the fix."
Чаму самэ ЛангЧейн для гэтага?
Калі працюеце над раздзелам «Чаму LangChain для гэтага этапу», спачатку запісайце контракт: неабяжлівыя даннэ, сігнал успеху і тое, што выходзіць у разе частковага нявыпання. Такі список контроля дапамагае заліцьварыць пазнейшыя змены ў кодзе. Спрэцьвачайце гэты этап як контракт межа даннемі і перакананымі выходамі. Дайце назву рэзультатам, задаце критэрыя успеху і не прымайце часткова завершанне без паведамлення. Зявляйце логі з назвай інструмента, хэшам параметраў, часам адклікання і рэзультатам кожнага вызову. Без такога следу дэбагаванне агента, які застряг у цыкле, займае гадзіны.
Прыямыя прычынкі
Калі працуеце над стадзіяй «Прыямнія умовы», спачатку запісайце «кантракт»: неабяжлівыя данні, сігнал успеху і тое, што выходзіць у разе частковага нявыпання. Такі список перакладоў заходзіць пазнейшыя змены коду ў правільным напрамку. Запісвайце час выканання і кост токена або запыту праза функцыйнальныя рэзултаты. Відразлівасць костаў з самага пачатку запобегае неспакойным рахункам, калі процес пераходзіць з дэмаверсіі ў спяльныя среды. Запішвайце назву інструмента, хэш аргументаў, час затрымкі і рэзултат кожнага вызову. Без такога лёгкага следу дэбагаванне агента губіць гады.
# Python 3.11+
python --version
# Install dependencies
pip install \\
langchain \\
langchain-anthropic \\
langchain-community \\
anthropic \\
python-dotenv \\
--break-system-packages
# Set your API key
export ANTHROPIC_API_KEY="your-key-here"
# Or create a .env file
echo "ANTHROPIC_API_KEY=your-key-here" > .env
Калі працуеце над стадзіяй «Прыямнія умовы», спачатку запісайце «кантракт»: неабяжлівыя данні, сігнал успеху і тое, што выходзіць у разе частковага нявыпання. Такі список перакладоў заходзіць пазнейшыя змены коду ў правільным напрамку. Аддзекументавайце як «шчаслівы» шлях, так і шлях вяснавання. Перапрыбуткі, людзкія контралі і обработка нерабочых паведамленняў є часткай продукту, а не чымсь, што дадаецца пазней.
devops-agent/
├── agent/
│ ├── __init__.py
│ ├── tools/
│ │ ├── __init__.py
│ │ ├── bash_tools.py # Core bash execution tools
│ │ ├── kubernetes_tools.py # kubectl wrapper tools
│ │ ├── log_tools.py # Log reading and analysis
│ │ └── system_tools.py # System info tools
│ ├── agent.py # Agent definition and loop
│ └── safety.py # Command safety checks
├── config.py
├── main.py
└── .env
Цыкл агента ReAct
Этап цыклу агента ReAct працюе найэфектывней, калі яго розглядаць як меравальную плошчу. Зафіксавайце адны ідеальны прыклад работы, адну справу з бягам і прыметку па поверненню да пачатковага стану пры розширэнні масштаба. Валіце маленькія, тэставальныя элементы замест большых скрыптов. Калі якісь крок не выйшае, прычына бягу должна вказываць на адну конкрэтную адпаведнасць, а не на заплутаны ланцюг задач. Адкройце інструменты з вузкімі схемамі та чысткімі пазначэннямі побачных наследкаў. Адпаведальныя за эксплуатацыю должны знать, якія вызовы мутуюць стан, прытаму яны можаць автаматычна санкціонаваць іх.
User: "Why is the API service slow?"
Agent Loop:
┌─────────────────────────────────────────────┐
│ THOUGHT: I need to check the service first │
│ ACTION: run_command("systemctl status api")│
│ OBSERVE: service is running, high CPU │
├─────────────────────────────────────────────┤
│ THOUGHT: High CPU - check what's running │
│ ACTION: run_command("top -bn1 | head -20")│
│ OBSERVE: api process at 95% CPU │
├─────────────────────────────────────────────┤
│ THOUGHT: Check recent logs for errors │
│ ACTION: read_logs("/var/log/api/app.log") │
│ OBSERVE: Massive DB query loop in logs │
├─────────────────────────────────────────────┤
│ THOUGHT: I have enough info now │
│ FINAL ANSWER: "The API is slow because..." │
└─────────────────────────────────────────────┘
Крок 1 — Безпека прыоритэт
Этап «Безпека пры ўсё» пасля першага крока даўа найлепыя рэзультаты, калі яго спрыяваць як меравальную плошчу. Зафіксавайце адна «золатая» транскрыпцыю, адин прыклад неудачы і запіс пра вярнэнне да пачатковага стану, перш чым расширваць сферу дзеяння. Спрыяйце гэтаму этапу як кантракту межаў між вхіднымі дадзеннямі і паўнастацыянымі выходамі. Дайце назвы артыфактам, задаць критэрыя успеху і не падтрымайце тыхню частковую завершэннасць. Адкройце інструменты з вузкімі схемамі та чысткімі пазначэннямі побачных эфектаў. Адпаведальныя за эксплуатацыю должны знать, якія вызовы мутуюць стан, перш чым автаматычна ўзяць іх на спрыт.
import re
from typing import Tuple
# Commands that are NEVER allowed regardless of context
BLOCKED_COMMANDS = [
r"rm\\s+-rf\\s+/", # rm -rf /
r"dd\\s+if=", # disk wipe
r"mkfs\\.", # format filesystem
r">\\s*/dev/sd", # write to disk
r"chmod\\s+-R\\s+777\\s+/", # open all permissions
r"passwd\\s+root", # change root password
r"userdel\\s+", # delete users
r"iptables\\s+-F", # flush all firewall rules
r"shutdown", # shutdown/reboot
r"reboot",
r"halt",
r"curl.*\\|\\s*bash", # pipe curl to bash
r"wget.*\\|\\s*bash", # pipe wget to bash
r"eval\\s+", # eval injection
r"base64\\s+--decode.*\\|", # decode and execute
]
# Commands that require extra caution (logged but allowed)
SENSITIVE_COMMANDS = [
"sudo", "su ", "ssh ", "scp ", "rsync",
"systemctl stop", "systemctl disable",
"apt remove", "yum remove", "pip uninstall",
"kubectl delete", "kubectl drain",
"terraform destroy", "ansible-playbook"
]
# Read-only commands - always safe
READONLY_COMMANDS = [
"cat ", "less ", "head ", "tail ", "grep ",
"find ", "ls ", "ps ", "top ", "df ", "du ",
"netstat", "ss ", "lsof ", "curl -s",
"systemctl status", "systemctl list",
"kubectl get", "kubectl describe", "kubectl logs",
"docker ps", "docker inspect", "docker logs",
"ansible --list-hosts",
"terraform plan", "terraform show",
"aws ec2 describe", "aws s3 ls",
"git log", "git status", "git diff",
"ping ", "traceroute", "nslookup", "dig "
]
def check_command_safety(command: str) -> Tuple[bool, str, str]:
"""
Returns: (is_allowed, risk_level, reason)
risk_level: "safe" | "sensitive" | "blocked"
"""
command_lower = command.lower().strip()
# Check blocked patterns
for pattern in BLOCKED_COMMANDS:
if re.search(pattern, command_lower):
return False, "blocked", f"Command matches blocked pattern: {pattern}"
# Check sensitive commands
for sensitive in SENSITIVE_COMMANDS:
if sensitive in command_lower:
return True, "sensitive", f"Sensitive command detected: {sensitive}"
# Check if it's read-only
for readonly in READONLY_COMMANDS:
if command_lower.startswith(readonly):
return True, "safe", "Read-only command"
# Default: allow but flag as unknown
return True, "unknown", "Command not in any list - proceeding with caution"
Крок 2 — Асновныя інструменты Bash
Этап Step 2 Core Bash працюе найкраща, калі яго розглядаць як параметрызаваную плошчу. Зберагучы адна ідеальная транскрыпцыя, адзін прыклад неудачы і запіс пра вярнэнне да пачатковага стану, спачатку трэба расширыць сферу дзеяння. Запісваючы час выканання і косць токенаў або запытак праза функцыйнальнымі рэзултатамі, можна заздалегідь бачыць косцы, чаго лепш утримацца ад неспакоўных рахунков, калі процес пераходзіць з дэмаверсіі ў спяльныя сераўы. Неабходна выклікаць інструменты з вузкімі схемамі і чыткімі пазначэннямі парадуктываў. Адпаведальныя за хосты должны знать, якія вызовы мутуюць стан, перш чым автаматычна схваліць іх. Этап Step 2 Core Bash працюе найкраща, калі яго розглядаць як параметрызаваную плошчу. Зберагучы адна ідеальная транскрыпцыя, адзін прыклад неудачы і запіс пра вярнэнне да пачатковага стану, спачатку трэба расширыць сферу дзеяння. Неабходна адначасова задокументаваць успішны і вярнэнчы маршруты. Перапрыбуткі, людзкія контрольныя пункты і обработка неканальных паведамленняў є часткай продукту, а не чымсь, што дадаецца пазней.
import subprocess
from langchain_core.tools import tool
from agent.safety import check_command_safety
import logging
logger = logging.getLogger(__name__)
@tool
def run_command(command: str) -> str:
"""
Run a bash command and return its output.
Use this to check system status, read configs, inspect processes.
Examples:
- run_command("systemctl status nginx")
- run_command("ps aux | grep python")
- run_command("df -h")
- run_command("netstat -tlnp")
"""
is_allowed, risk_level, reason = check_command_safety(command)
if not is_allowed:
return f"BLOCKED: Command not allowed. Reason: {reason}"
if risk_level == "sensitive":
logger.warning(f"SENSITIVE command executed: {command}")
try:
result = subprocess.run(
command,
shell=True,
capture_output=True,
text=True,
timeout=30
)
output = result.stdout or result.stderr
# Truncate very long output
if len(output) > 3000:
output = output[:3000] + "\\n... [output truncated]"
if result.returncode != 0:
return f"Command failed (exit {result.returncode}):\\n{output}"
return output or "(no output)"
except subprocess.TimeoutExpired:
return "Command timed out after 30 seconds"
except Exception as e:
return f"Error running command: {str(e)}"
@tool
def read_file(file_path: str) -> str:
"""
Read the contents of a file.
Use this to inspect configs, logs, scripts.
Examples:
- read_file("/etc/nginx/nginx.conf")
- read_file("/var/log/syslog")
- read_file("/etc/systemd/system/myapp.service")
"""
try:
with open(file_path, 'r', errors='replace') as f:
content = f.read()
# Truncate large files
if len(content) > 4000:
# Show beginning and end for log files
content = content[:2000] + "\\n\\n... [middle truncated] ...\\n\\n" + content[-2000:]
return content
except PermissionError:
return f"Permission denied: {file_path}"
except FileNotFoundError:
return f"File not found: {file_path}"
except Exception as e:
return f"Error reading file: {str(e)}"
@tool
def read_logs(service_name: str, lines: int = 50) -> str:
"""
Read recent logs for a systemd service using journalctl.
Use this to diagnose service errors, crashes, or warnings.
Examples:
- read_logs("nginx")
- read_logs("docker", lines=100)
- read_logs("postgresql")
"""
command = f"journalctl -u {service_name} -n {lines} --no-pager -o short-iso"
try:
result = subprocess.run(
command, shell=True,
capture_output=True, text=True, timeout=15
)
return result.stdout or result.stderr or "(no logs found)"
except Exception as e:
return f"Error reading logs: {str(e)}"
@tool
def grep_logs(log_path: str, pattern: str, lines_before: int = 2, lines_after: int = 2) -> str:
"""
Search for a pattern in a log file with context lines.
Use this to find specific errors, IPs, or events in logs.
Examples:
- grep_logs("/var/log/nginx/error.log", "upstream timed out")
- grep_logs("/var/log/auth.log", "Failed password", lines_after=0)
"""
command = f"grep -i -B {lines_before} -A {lines_after} '{pattern}' {log_path} | tail -100"
is_allowed, _, _ = check_command_safety(f"grep {log_path}")
if not is_allowed:
return "BLOCKED: Cannot read this log file"
try:
result = subprocess.run(
command, shell=True,
capture_output=True, text=True, timeout=20
)
return result.stdout or f"No matches found for '{pattern}' in {log_path}"
except Exception as e:
return f"Error searching logs: {str(e)}"
Этап 3 — Інструменты Kubernetes
У стадії «Інструменты Kubernetes, крок 3» аўтары должны вказаць параметры вхідных дадзеных, адпаведальнага за крок і крэтарыя завершэння працы перад змінайом коду. Аператары павинны магчымаць перзапуск кроку з вядомай точкі контролю, не прабуючы спадарожваць схованы стан. Лепш выбіраць маленькія, тэставаныя елементы замест большых скрыптов. Калі крок не выйшаў, прычына неудачы павінна вказываць на адну конкрэтную адпаведальнасць, а не на заплутаны процес. Автентыфікуйцеся на воратах і паўторна автарызуйцеся на роўні дадзеных. Толькі токэн-носіцель не є межай адпаведальнасцяў.
import subprocess
from langchain_core.tools import tool
def kubectl(command: str) -> str:
"""Run a kubectl command safely."""
# Only allow read operations from the agent
read_verbs = ["get", "describe", "logs", "top", "explain", "version", "cluster-info"]
cmd_parts = command.strip().split()
if cmd_parts and cmd_parts[0] not in read_verbs:
return f"BLOCKED: Only read operations allowed. Got: {cmd_parts[0]}"
result = subprocess.run(
f"kubectl {command}",
shell=True, capture_output=True, text=True, timeout=30
)
return result.stdout or result.stderr
@tool
def k8s_get_pods(namespace: str = "default") -> str:
"""
List all pods in a namespace with their status.
Use this to check if pods are running, crashing, or pending.
Examples:
- k8s_get_pods("production")
- k8s_get_pods("monitoring")
"""
return kubectl(f"get pods -n {namespace} -o wide")
@tool
def k8s_describe_pod(pod_name: str, namespace: str = "default") -> str:
"""
Get detailed info about a specific pod including events.
Use this when a pod is failing to understand why.
Examples:
- k8s_describe_pod("api-worker-6d4f9b", "production")
"""
return kubectl(f"describe pod {pod_name} -n {namespace}")
@tool
def k8s_get_logs(pod_name: str, namespace: str = "default", lines: int = 50) -> str:
"""
Get logs from a Kubernetes pod.
Use this to see application errors inside a pod.
Examples:
- k8s_get_logs("api-worker-6d4f9b", "production", lines=100)
"""
return kubectl(f"logs {pod_name} -n {namespace} --tail={lines}")
@tool
def k8s_get_events(namespace: str = "default") -> str:
"""
Get recent Kubernetes events in a namespace.
Sorted by time - shows warnings, errors, pod restarts.
Examples:
- k8s_get_events("production")
"""
return kubectl(f"get events -n {namespace} --sort-by='.lastTimestamp'")
@tool
def k8s_top_pods(namespace: str = "default") -> str:
"""
Get CPU and memory usage for all pods in a namespace.
Use this to find resource-hungry pods.
Examples:
- k8s_top_pods("production")
"""
return kubectl(f"top pods -n {namespace} --sort-by=memory")
Крок 4 — Інструменты інфармаціі пра систему
У стадії «Інфармацыя працоўнікага системы», якая ўжо ў чатвертам кроку, перш чым зменіць код, неабходна вказаць параметры вхідных дадзеных, адпаведальнага за крок і крэтарыі завершэння. Аперацыйныя працоўнікі должны магчымае запускаць крок з вядомай точкі контролю, не прабуючы спадарожваць схованы стан. Спрыяйце цій стадіі як даговору межы вхідных дадзеных і перакананых выходных рэзультатаў. Даць назвы артыфактам, визначыць крэтарыі успеху і не падтрымваць тыхя частковых завершэнняў. Автентыфікуйцеся на в’язку і паўтарна автарызуйцеся на роўні дадзеных. Толькі токэн-носіцель не є межай арендаванага ресурсу.
import subprocess
import psutil
from langchain_core.tools import tool
from datetime import datetime
@tool
def get_system_overview() -> str:
"""
Get a full system health snapshot:
CPU, memory, disk usage, load average, uptime.
Always call this first when diagnosing a system issue.
"""
try:
cpu = psutil.cpu_percent(interval=1)
memory = psutil.virtual_memory()
disk = psutil.disk_usage('/')
load = subprocess.run(
"uptime", capture_output=True, text=True
).stdout.strip()
return f"""System Overview ({datetime.now().strftime('%H:%M:%S')}):
CPU: {cpu}% used
Memory: {memory.percent}% used ({memory.used // 1024**3}GB / {memory.total // 1024**3}GB)
Disk: {disk.percent}% used ({disk.used // 1024**3}GB / {disk.total // 1024**3}GB)
Load: {load}"""
except Exception as e:
return f"Error getting system overview: {e}"
@tool
def check_service_status(service_name: str) -> str:
"""
Check if a systemd service is running and get its status.
Examples:
- check_service_status("nginx")
- check_service_status("postgresql")
- check_service_status("docker")
"""
result = subprocess.run(
f"systemctl status {service_name}",
shell=True, capture_output=True, text=True
)
return result.stdout or result.stderr
@tool
def check_port(port: int) -> str:
"""
Check what process is listening on a specific port.
Use this to verify services are bound to expected ports
or to find unexpected listeners.
Examples:
- check_port(80)
- check_port(5432)
- check_port(6379)
"""
result = subprocess.run(
f"ss -tlnp sport = :{port}",
shell=True, capture_output=True, text=True
)
if not result.stdout.strip():
return f"Nothing is listening on port {port}"
return result.stdout
@tool
def get_top_processes(sort_by: str = "cpu") -> str:
"""
Get the top 10 processes by CPU or memory usage.
Use this to find runaway processes causing system load.
Args:
sort_by: "cpu" or "memory"
"""
sort_flag = "-%cpu" if sort_by == "cpu" else "-%mem"
result = subprocess.run(
f"ps aux --sort={sort_flag} | head -11",
shell=True, capture_output=True, text=True
)
return result.stdout
@tool
def check_disk_usage(path: str = "/") -> str:
"""
Check disk usage for a path and show largest directories.
Use this when disk space is running low.
Examples:
- check_disk_usage("/")
- check_disk_usage("/var/log")
- check_disk_usage("/home")
"""
df = subprocess.run(
f"df -h {path}",
shell=True, capture_output=True, text=True
).stdout
du = subprocess.run(
f"du -sh {path}/* 2>/dev/null | sort -rh | head -10",
shell=True, capture_output=True, text=True
).stdout
return f"Disk usage for {path}:\\n{df}\\nLargest directories:\\n{du}"
5-й крок — Стварэнне агента
Аварыя на кроку 5: падчас змены коду неабходна стварэнне сцэны, апрацоўка параметраў вхідных дадзеных, вызначэнне адпаведальнага за крок і критэрыя завершэння. Аператары должны магчымае перзапускіць крок з вядомай точкі контролю, не падозрываючы прыхованы стан. Неабходна фіксацыя часу выкарыстоўвання і косту токенаў чы выкарыстоўвання запитаў разам з функцыональнымі рэзултатамі. Відразлівае паказанне костаў запобегае неспадзяваным расчыткам, калі процес пераходзіць з дэмавай версіі ў спяльныя сераўы. Аутентыфікацыя трэба адбывацца на воратах, а прабачэнне правоў — на роўні дадзеных. Толькі токен-носіцель не є межай адпаведнага тэнантства.
from langchain_anthropic import ChatAnthropic
from langchain.agents import AgentExecutor, create_react_agent
from langchain_core.prompts import PromptTemplate
from langchain_core.tools import BaseTool
from typing import List
import os
from agent.tools.bash_tools import run_command, read_file, read_logs, grep_logs
from agent.tools.kubernetes_tools import (
k8s_get_pods, k8s_describe_pod, k8s_get_logs,
k8s_get_events, k8s_top_pods
)
from agent.tools.system_tools import (
get_system_overview, check_service_status,
check_port, get_top_processes, check_disk_usage
)
SYSTEM_PROMPT = """You are an expert DevOps engineer and SRE with 10 years of experience.
You have access to tools that let you inspect systems, read logs, check services,
and query Kubernetes clusters.
Your approach:
- Start with broad system checks, then narrow down to specifics
- Always check logs when a service is misbehaving
- Look for patterns - one error often points to another
- Explain what you're doing and why as you go
- Give clear, actionable recommendations at the end
- Never run destructive commands - only read and inspect
When diagnosing issues:
1. Get system overview first if it's a general performance issue
2. Check the specific service/pod status
3. Read recent logs for errors
4. Cross-reference with system metrics
5. Provide root cause + recommended fix
You have access to these tools:
{tools}
Use this format:
Thought: [what you're thinking and why]
Action: [tool name]
Action Input: [tool input]
Observation: [what the tool returned]
... (repeat as needed)
Thought: I now have enough information to answer
Final Answer: [clear explanation + recommendations]
Begin!
Question: {input}
{agent_scratchpad}"""
def build_devops_agent(
include_kubernetes: bool = True,
verbose: bool = True
) -> AgentExecutor:
"""Build and return the DevOps AI agent."""
# Initialize Claude
llm = ChatAnthropic(
model="claude-sonnet-4-20250514",
temperature=0, # deterministic for ops tasks
max_tokens=4096,
anthropic_api_key=os.getenv("ANTHROPIC_API_KEY")
)
# Collect tools
tools: List[BaseTool] = [
run_command,
read_file,
read_logs,
grep_logs,
get_system_overview,
check_service_status,
check_port,
get_top_processes,
check_disk_usage,
]
if include_kubernetes:
tools.extend([
k8s_get_pods,
k8s_describe_pod,
k8s_get_logs,
k8s_get_events,
k8s_top_pods,
])
# Build prompt
prompt = PromptTemplate(
input_variables=["input", "tools", "tool_names", "agent_scratchpad"],
template=SYSTEM_PROMPT
)
# Create ReAct agent
agent = create_react_agent(llm, tools, prompt)
# Wrap in executor
return AgentExecutor(
agent=agent,
tools=tools,
verbose=verbose, # streams thinking to terminal
max_iterations=15, # prevent infinite loops
handle_parsing_errors=True, # recover from malformed outputs
return_intermediate_steps=True
)
Крок 6 — Точка входу
import os
from dotenv import load_dotenv
from agent.agent import build_devops_agent
load_dotenv()
def run_interactive():
"""Interactive mode - chat with your agent."""
print("\\n🤖 DevOps AI Agent")
print(" Powered by Claude + LangChain")
print(" Type 'exit' to quit\\n")
print(" Examples:")
print(" → Why is nginx returning 502 errors?")
print(" → Which pod is using the most memory in production?")
print(" → Check disk space and find what's eating it")
print(" → Is the database service healthy?\\n")
agent = build_devops_agent(
include_kubernetes=True,
verbose=True
)
while True:
try:
user_input = input("\\n💬 You: ").strip()
if not user_input:
continue
if user_input.lower() in ["exit", "quit", "q"]:
print("👋 Goodbye!")
break
print("\\n🤖 Agent thinking...\\n")
result = agent.invoke({"input": user_input})
print(f"\\n{'='*60}")
print("FINAL ANSWER:")
print('='*60)
print(result["output"])
except KeyboardInterrupt:
print("\\n\\n👋 Interrupted. Goodbye!")
break
except Exception as e:
print(f"\\n❌ Agent error: {e}")
def run_single_query(query: str):
"""Run a single query - useful for scripting."""
agent = build_devops_agent(verbose=False)
result = agent.invoke({"input": query})
print(result["output"])
if __name__ == "__main__":
import sys
if len(sys.argv) > 1:
# Single query mode
run_single_query(" ".join(sys.argv[1:]))
else:
# Interactive mode
run_interactive()
Пазірце, як гэта працуе
python main.py
💬 You: The API service is responding slowly. Find out why.
🤖 Agent thinking...
Thought: I'll start with a system overview, then check the API service.
Action: get_system_overview
Action Input: {}
Observation:
CPU: 87% used
Memory: 94% used (7.5GB / 8GB)
Disk: 45% used
Load: load average: 6.2, 5.8, 4.1
Thought: Very high CPU and memory. Check what's consuming resources.
Action: get_top_processes
Action Input: {"sort_by": "memory"}
Observation:
USER PID %CPU %MEM COMMAND
api 14821 45.2 67.3 python api_worker.py
api 14822 38.1 21.4 python api_worker.py
...
Thought: Two API worker processes are using 88% of memory combined.
Let me check the service logs for clues.
Action: read_logs
Action Input: {"service_name": "api-worker", "lines": 100}
Observation:
2026-05-22 02:31:14 WARNING Slow query: SELECT * FROM events took 8432ms
2026-05-22 02:31:19 WARNING Slow query: SELECT * FROM events took 9103ms
2026-05-22 02:31:22 ERROR Connection pool exhausted (pool_size=10)
2026-05-22 02:31:28 WARNING Slow query: SELECT * FROM events took 11201ms
Thought: DB connection pool is exhausted and queries are very slow.
Check if PostgreSQL is healthy.
Action: check_service_status
Action Input: {"service_name": "postgresql"}
Observation:
● postgresql.service - PostgreSQL RDBMS
Loaded: loaded (/lib/systemd/system/postgresql.service)
Active: active (running) since 2026-05-21 18:00:01 UTC
May 22 02:28:11 postgres[1842]: LOG: checkpoint taking longer than 30s
Thought: Checkpoints are taking too long - DB is under heavy write load.
I have enough information now.
Final Answer:
Root Cause: The API is slow because of a database bottleneck, not the app itself.
What's happening:
- PostgreSQL is performing slow checkpoints (>30s), causing all queries to queue
- The `events` table queries are taking 8-11 seconds each (should be <100ms)
- This has exhausted the connection pool (size: 10), causing all new requests to wait
Why it cascaded:
- Slow DB queries → workers hold connections longer
- Connection pool exhausts → new requests queue
- Queue builds up → high CPU from context switching → memory pressure
Immediate actions (do these now):
1. Check if the events table has a missing index:
→ psql -c "EXPLAIN ANALYZE SELECT * FROM events WHERE ..."
2. Temporarily increase connection pool:
→ Set DB_POOL_SIZE=25 and restart api-worker
3. Check for a long-running transaction blocking checkpoints:
→ psql -c "SELECT pid, age(clock_timestamp(), xact_start) FROM pg_stat_activity WHERE xact_start IS NOT NULL ORDER BY xact_start"
Root fix:
- Add an index on the events table (likely missing on timestamp or foreign key column)
- Consider query result caching for high-frequency event reads
Дадаць памяць — канверсаціі з калькама раундав
from langchain.memory import ConversationBufferWindowMemory
from langchain.agents import AgentExecutor, create_react_agent
# Add memory - keeps last 10 exchanges
memory = ConversationBufferWindowMemory(
k=10,
memory_key="chat_history",
return_messages=True
)
# Updated system prompt with memory
SYSTEM_PROMPT_WITH_MEMORY = """...(same as before but add)...
Previous conversation:
{chat_history}
Question: {input}
{agent_scratchpad}"""
# Pass memory to executor
agent_executor = AgentExecutor(
agent=agent,
tools=tools,
memory=memory,
verbose=True,
max_iterations=15,
)
💬 You: What's eating disk space on the server?
🤖 Agent: [checks disk] /var/log is using 47GB.
/var/log/nginx/access.log is 38GB alone.
💬 You: How long has it been growing?
🤖 Agent: [remembers context - checks log rotation config]
Log rotation isn't configured for nginx.
The file has been growing unrotated since March 14th.
💬 You: What's the fix?
🤖 Agent: [still has full context]
Configure logrotate for nginx. Here's the exact config...
Упакаваць у REST API
# api.py
from fastapi import FastAPI, BackgroundTasks
from pydantic import BaseModel
import uuid
from agent.agent import build_devops_agent
app = FastAPI(title="DevOps AI Agent API")
agent = build_devops_agent(verbose=False)
results_store = {} # In production: use Redis
class QueryRequest(BaseModel):
query: str
environment: str = "production"
class QueryResponse(BaseModel):
job_id: str
status: str
result: str = None
@app.post("/query", response_model=QueryResponse)
async def run_query(request: QueryRequest, background_tasks: BackgroundTasks):
"""Submit a query to the DevOps agent."""
job_id = str(uuid.uuid4())
results_store[job_id] = {"status": "running", "result": None}
def run_agent_task():
result = agent.invoke({"input": request.query})
results_store[job_id] = {
"status": "complete",
"result": result["output"]
}
background_tasks.add_task(run_agent_task)
return QueryResponse(job_id=job_id, status="running")
@app.get("/query/{job_id}", response_model=QueryResponse)
async def get_result(job_id: str):
"""Poll for query result."""
job = results_store.get(job_id)
if not job:
return QueryResponse(job_id=job_id, status="not_found")
return QueryResponse(
job_id=job_id,
status=job["status"],
result=job["result"]
)
# Run: uvicorn api:app --host 0.0.0.0 --port 8000
# From a CI/CD pipeline
curl -X POST <http://agent:8000/query> \\
-H "Content-Type: application/json" \\
-d '{"query": "Is the staging deployment healthy after this release?"}'
# From a Makefile
make check-deploy:
curl -s <http://agent:8000/query> \\
-d '{"query": "Run post-deploy checks on staging"}' | jq .result
Рэальныя прыклады выкарыстоўвання, якія можна працаваць зараз
"Nginx is returning 504 errors. Find the root cause."
"The payment service pod keeps crashing. What's happening?"
"Memory usage spiked 30 minutes ago. What caused it?"
"Run a health check on all services and give me a summary."
"Which pods in production are close to their memory limits?"
"Check if any SSL certificates are expiring in the next 30 days."
"At current growth rate, when will we run out of disk space in /var/log?"
"Which services are consistently above 80% CPU?"
"Find any processes listening on unexpected ports."
"Check auth logs for failed login attempts in the last hour."
"Are there any world-writable files in /etc?"
Звычныя памылкі і способы ўсунення
# Set a hard iteration limit
AgentExecutor(max_iterations=15, ...) # never go above 20
# Always truncate in your tools
if len(output) > 3000:
output = output[:3000] + "\\n... [truncated]"
# Don't rely on the LLM to self-police
# Enforce safety in the tool code itself (like we did in safety.py)
# The LLM cannot bypass code-level blocks
# Use streaming to show progress
async for chunk in agent.astream({"input": query}):
if "actions" in chunk:
for action in chunk["actions"]:
print(f" → Running: {action.tool}({action.tool_input})")