首页 / 文章 / 实用提示:使用 LangChain 与 Bash 工具构建你的第一个 DevOps AI 智能体

实用提示:使用 LangChain 与 Bash 工具构建你的第一个 DevOps AI 智能体

《实用笔记》操作指南:使用 LangChain 与 Bash 工具构建你的第一个 DevOps AI 智能体——为采用该架构的团队提供合同模板、检查清单以及可直接插入的代码片段。

4204 词

本指南将逐步构建从原材料到可运行系统的完整流程,内容为:使用 LangChain 和 Bash 工具打造你的第一个 DevOps AI 智能体。重点在于可操作的步骤、明确的检查点,以及可直接放入代码仓库的代码,无需猜测其用途。 在修改代码之前,作者需先确定概览阶段、输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品不可或缺的部分,而非后续需要补充的功能。

我们究竟在构建什么?

在“我们正在构建什么”这一阶段,首先写下相关约定:所需的输入参数、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 相比冗长的脚本,更应选择小型且可测试的单元。当某个步骤失败时,故障应能指向单一的责任模块,而非复杂的流程链。 需为每次调用记录工具名称、参数哈希值、延迟时间以及执行结果。没有这些记录,调试过程将会浪费大量时间。

Chatbot:  "Your nginx config might have a syntax error."
Agent:    [runs `nginx -t`] → "Confirmed. Line 42 in /etc/nginx/sites-enabled/app.conf
           has an unclosed bracket. Here's the fix."

为何选择 LangChain?

在处理“为何选择LangChain”这一阶段时,首先需列出相关契约:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与验证后输出之间的契约。为相关成果命名,明确成功判定标准,并杜绝无声的半完成状态。 需记录每次调用的工具名称、参数哈希值、延迟时间以及最终结果。没有这些记录,调试智能体循环问题将会耗费大量时间。

前置条件

在完成前置条件阶段时,首先写下接口规范:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 为每次调用记录工具名称、参数哈希值、延迟时间以及最终结果。没有这些记录,调试代理将陷入无止境的循环,耗费大量时间。

# Python 3.11+
python --version

# Install dependencies
pip install \\
  langchain \\
  langchain-anthropic \\
  langchain-community \\
  anthropic \\
  python-dotenv \\
  --break-system-packages

# Set your API key
export ANTHROPIC_API_KEY="your-key-here"

# Or create a .env file
echo "ANTHROPIC_API_KEY=your-key-here" > .env

在完成前置条件阶段时,首先写下接口规范:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。

devops-agent/
├── agent/
│   ├── __init__.py
│   ├── tools/
│   │   ├── __init__.py
│   │   ├── bash_tools.py       # Core bash execution tools
│   │   ├── kubernetes_tools.py # kubectl wrapper tools
│   │   ├── log_tools.py        # Log reading and analysis
│   │   └── system_tools.py     # System info tools
│   ├── agent.py                # Agent definition and loop
│   └── safety.py               # Command safety checks
├── config.py
├── main.py
└── .env

ReAct 代理循环

将ReAct Agent Loop阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一份理想的执行日志、一个失败案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤失败时,故障应指向单一责任模块,而非复杂的流程链。 使用结构明确的工具,并标注清晰的副作用信息。主机需要在自动批准之前知道哪些调用会修改状态。

User: "Why is the API service slow?"

Agent Loop:
┌─────────────────────────────────────────────┐
│  THOUGHT: I need to check the service first │
│  ACTION:  run_command("systemctl status api")│
│  OBSERVE: service is running, high CPU      │
├─────────────────────────────────────────────┤
│  THOUGHT: High CPU - check what's running   │
│  ACTION:  run_command("top -bn1 | head -20")│
│  OBSERVE: api process at 95% CPU            │
├─────────────────────────────────────────────┤
│  THOUGHT: Check recent logs for errors      │
│  ACTION:  read_logs("/var/log/api/app.log") │
│  OBSERVE: Massive DB query loop in logs     │
├─────────────────────────────────────────────┤
│  THOUGHT: I have enough info now            │
│  FINAL ANSWER: "The API is slow because..." │
└─────────────────────────────────────────────┘

第一步 — 安全第一

在将“第一步:安全第一”视为可度量的流程层面时,其效果最佳。在扩大范围之前,需记录一份优秀的处理结果、一个故障案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许默许不完整的处理结果。 使用具有严格结构定义和明确副作用标注的工具。在自动批准之前,系统管理员必须清楚哪些操作会改变状态。

import re
from typing import Tuple

# Commands that are NEVER allowed regardless of context
BLOCKED_COMMANDS = [
    r"rm\\s+-rf\\s+/",           # rm -rf /
    r"dd\\s+if=",               # disk wipe
    r"mkfs\\.",                 # format filesystem
    r">\\s*/dev/sd",            # write to disk
    r"chmod\\s+-R\\s+777\\s+/",  # open all permissions
    r"passwd\\s+root",          # change root password
    r"userdel\\s+",             # delete users
    r"iptables\\s+-F",          # flush all firewall rules
    r"shutdown",               # shutdown/reboot
    r"reboot",
    r"halt",
    r"curl.*\\|\\s*bash",        # pipe curl to bash
    r"wget.*\\|\\s*bash",        # pipe wget to bash
    r"eval\\s+",                # eval injection
    r"base64\\s+--decode.*\\|",  # decode and execute
]

# Commands that require extra caution (logged but allowed)
SENSITIVE_COMMANDS = [
    "sudo", "su ", "ssh ", "scp ", "rsync",
    "systemctl stop", "systemctl disable",
    "apt remove", "yum remove", "pip uninstall",
    "kubectl delete", "kubectl drain",
    "terraform destroy", "ansible-playbook"
]

# Read-only commands - always safe
READONLY_COMMANDS = [
    "cat ", "less ", "head ", "tail ", "grep ",
    "find ", "ls ", "ps ", "top ", "df ", "du ",
    "netstat", "ss ", "lsof ", "curl -s",
    "systemctl status", "systemctl list",
    "kubectl get", "kubectl describe", "kubectl logs",
    "docker ps", "docker inspect", "docker logs",
    "ansible --list-hosts",
    "terraform plan", "terraform show",
    "aws ec2 describe", "aws s3 ls",
    "git log", "git status", "git diff",
    "ping ", "traceroute", "nslookup", "dig "
]

def check_command_safety(command: str) -> Tuple[bool, str, str]:
    """
    Returns: (is_allowed, risk_level, reason)
    risk_level: "safe" | "sensitive" | "blocked"
    """
    command_lower = command.lower().strip()

    # Check blocked patterns
    for pattern in BLOCKED_COMMANDS:
        if re.search(pattern, command_lower):
            return False, "blocked", f"Command matches blocked pattern: {pattern}"

    # Check sensitive commands
    for sensitive in SENSITIVE_COMMANDS:
        if sensitive in command_lower:
            return True, "sensitive", f"Sensitive command detected: {sensitive}"

    # Check if it's read-only
    for readonly in READONLY_COMMANDS:
        if command_lower.startswith(readonly):
            return True, "safe", "Read-only command"

    # Default: allow but flag as unknown
    return True, "unknown", "Command not in any list - proceeding with caution"

第二步 — 核心 Bash 工具

将第二步的Core Bash阶段视为可度量的测试环境,效果最佳。在扩大范围之前,需记录一份理想的执行日志、一个失败案例以及回滚说明。 在功能结果旁同时记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 应使用结构清晰且带有明确副作用标注的工具。主机需要在自动批准之前知道哪些调用会改变系统状态。 将第二步的Core Bash阶段视为可度量的测试环境,效果最佳。在扩大范围之前,需记录一份理想的执行日志、一个失败案例以及回滚说明。 需同时记录正常流程和恢复流程的文档。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。

import subprocess
from langchain_core.tools import tool
from agent.safety import check_command_safety
import logging

logger = logging.getLogger(__name__)

@tool
def run_command(command: str) -> str:
    """
    Run a bash command and return its output.
    Use this to check system status, read configs, inspect processes.

    Examples:
    - run_command("systemctl status nginx")
    - run_command("ps aux | grep python")
    - run_command("df -h")
    - run_command("netstat -tlnp")
    """
    is_allowed, risk_level, reason = check_command_safety(command)

    if not is_allowed:
        return f"BLOCKED: Command not allowed. Reason: {reason}"

    if risk_level == "sensitive":
        logger.warning(f"SENSITIVE command executed: {command}")

    try:
        result = subprocess.run(
            command,
            shell=True,
            capture_output=True,
            text=True,
            timeout=30
        )

        output = result.stdout or result.stderr

        # Truncate very long output
        if len(output) > 3000:
            output = output[:3000] + "\\n... [output truncated]"

        if result.returncode != 0:
            return f"Command failed (exit {result.returncode}):\\n{output}"

        return output or "(no output)"

    except subprocess.TimeoutExpired:
        return "Command timed out after 30 seconds"
    except Exception as e:
        return f"Error running command: {str(e)}"

@tool
def read_file(file_path: str) -> str:
    """
    Read the contents of a file.
    Use this to inspect configs, logs, scripts.

    Examples:
    - read_file("/etc/nginx/nginx.conf")
    - read_file("/var/log/syslog")
    - read_file("/etc/systemd/system/myapp.service")
    """

    try:
        with open(file_path, 'r', errors='replace') as f:

            content = f.read()

        # Truncate large files
        if len(content) > 4000:
            # Show beginning and end for log files
            content = content[:2000] + "\\n\\n... [middle truncated] ...\\n\\n" + content[-2000:]

        return content
    except PermissionError:
        return f"Permission denied: {file_path}"
    except FileNotFoundError:
        return f"File not found: {file_path}"
    except Exception as e:
        return f"Error reading file: {str(e)}"

@tool
def read_logs(service_name: str, lines: int = 50) -> str:
    """
    Read recent logs for a systemd service using journalctl.
    Use this to diagnose service errors, crashes, or warnings.

    Examples:
    - read_logs("nginx")
    - read_logs("docker", lines=100)
    - read_logs("postgresql")
    """

    command = f"journalctl -u {service_name} -n {lines} --no-pager -o short-iso"

    try:
        result = subprocess.run(
            command, shell=True,
            capture_output=True, text=True, timeout=15
        )

        return result.stdout or result.stderr or "(no logs found)"
    except Exception as e:
        return f"Error reading logs: {str(e)}"

@tool
def grep_logs(log_path: str, pattern: str, lines_before: int = 2, lines_after: int = 2) -> str:
    """
    Search for a pattern in a log file with context lines.
    Use this to find specific errors, IPs, or events in logs.

    Examples:
    - grep_logs("/var/log/nginx/error.log", "upstream timed out")
    - grep_logs("/var/log/auth.log", "Failed password", lines_after=0)
    """

    command = f"grep -i -B {lines_before} -A {lines_after} '{pattern}' {log_path} | tail -100"
    is_allowed, _, _ = check_command_safety(f"grep {log_path}")

    if not is_allowed:
        return "BLOCKED: Cannot read this log file"

    try:
        result = subprocess.run(
            command, shell=True,
            capture_output=True, text=True, timeout=20
        )

        return result.stdout or f"No matches found for '{pattern}' in {log_path}"
    except Exception as e:
        return f"Error searching logs: {str(e)}"

第三步 — Kubernetes工具

在“步骤3:Kubernetes工具”阶段,作者应在修改代码之前明确输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏的状态。相比庞大的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。应在网关处进行身份验证,并在数据层再次授权——仅凭承载令牌并不足以界定租户边界。

import subprocess
from langchain_core.tools import tool

def kubectl(command: str) -> str:
    """Run a kubectl command safely."""

    # Only allow read operations from the agent
    read_verbs = ["get", "describe", "logs", "top", "explain", "version", "cluster-info"]

    cmd_parts = command.strip().split()

    if cmd_parts and cmd_parts[0] not in read_verbs:
        return f"BLOCKED: Only read operations allowed. Got: {cmd_parts[0]}"

    result = subprocess.run(
        f"kubectl {command}",
        shell=True, capture_output=True, text=True, timeout=30
    )

    return result.stdout or result.stderr

@tool
def k8s_get_pods(namespace: str = "default") -> str:
    """
    List all pods in a namespace with their status.
    Use this to check if pods are running, crashing, or pending.

    Examples:
    - k8s_get_pods("production")
    - k8s_get_pods("monitoring")
    """
    return kubectl(f"get pods -n {namespace} -o wide")

@tool
def k8s_describe_pod(pod_name: str, namespace: str = "default") -> str:
    """
    Get detailed info about a specific pod including events.
    Use this when a pod is failing to understand why.

    Examples:
    - k8s_describe_pod("api-worker-6d4f9b", "production")
    """

    return kubectl(f"describe pod {pod_name} -n {namespace}")

@tool
def k8s_get_logs(pod_name: str, namespace: str = "default", lines: int = 50) -> str:
    """
    Get logs from a Kubernetes pod.
    Use this to see application errors inside a pod.

    Examples:
    - k8s_get_logs("api-worker-6d4f9b", "production", lines=100)
    """
    return kubectl(f"logs {pod_name} -n {namespace} --tail={lines}")

@tool
def k8s_get_events(namespace: str = "default") -> str:
    """
    Get recent Kubernetes events in a namespace.
    Sorted by time - shows warnings, errors, pod restarts.

    Examples:
    - k8s_get_events("production")
    """

    return kubectl(f"get events -n {namespace} --sort-by='.lastTimestamp'")

@tool
def k8s_top_pods(namespace: str = "default") -> str:
    """
    Get CPU and memory usage for all pods in a namespace.
    Use this to find resource-hungry pods.

    Examples:
    - k8s_top_pods("production")
    """

    return kubectl(f"top pods -n {namespace} --sort-by=memory")

步骤4 — 系统信息工具

在“步骤4:系统信息”阶段,作者需在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应将此阶段视为输入与已验证输出之间的契约:为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。在网关处进行身份验证,在数据层再次授权——仅凭承载令牌并不足以界定租户边界。

import subprocess
import psutil
from langchain_core.tools import tool
from datetime import datetime

@tool
def get_system_overview() -> str:
    """
    Get a full system health snapshot:
    CPU, memory, disk usage, load average, uptime.
    Always call this first when diagnosing a system issue.
    """
    try:
        cpu = psutil.cpu_percent(interval=1)
        memory = psutil.virtual_memory()
        disk = psutil.disk_usage('/')

        load = subprocess.run(
            "uptime", capture_output=True, text=True
        ).stdout.strip()

        return f"""System Overview ({datetime.now().strftime('%H:%M:%S')}):
CPU:    {cpu}% used
Memory: {memory.percent}% used ({memory.used // 1024**3}GB / {memory.total // 1024**3}GB)
Disk:   {disk.percent}% used ({disk.used // 1024**3}GB / {disk.total // 1024**3}GB)
Load:   {load}"""
    except Exception as e:
        return f"Error getting system overview: {e}"

@tool
def check_service_status(service_name: str) -> str:
    """
    Check if a systemd service is running and get its status.

    Examples:
    - check_service_status("nginx")
    - check_service_status("postgresql")
    - check_service_status("docker")
    """

    result = subprocess.run(
        f"systemctl status {service_name}",
        shell=True, capture_output=True, text=True

    )

    return result.stdout or result.stderr

@tool
def check_port(port: int) -> str:
    """
    Check what process is listening on a specific port.
    Use this to verify services are bound to expected ports
    or to find unexpected listeners.

    Examples:
    - check_port(80)
    - check_port(5432)
    - check_port(6379)
    """

    result = subprocess.run(
        f"ss -tlnp sport = :{port}",
        shell=True, capture_output=True, text=True
    )

    if not result.stdout.strip():
        return f"Nothing is listening on port {port}"
    return result.stdout

@tool
def get_top_processes(sort_by: str = "cpu") -> str:
    """
    Get the top 10 processes by CPU or memory usage.
    Use this to find runaway processes causing system load.
    Args:
        sort_by: "cpu" or "memory"
    """
    sort_flag = "-%cpu" if sort_by == "cpu" else "-%mem"

    result = subprocess.run(
        f"ps aux --sort={sort_flag} | head -11",
        shell=True, capture_output=True, text=True
    )
    return result.stdout

@tool
def check_disk_usage(path: str = "/") -> str:
    """
    Check disk usage for a path and show largest directories.
    Use this when disk space is running low.

    Examples:
    - check_disk_usage("/")
    - check_disk_usage("/var/log")
    - check_disk_usage("/home")
    """

    df = subprocess.run(
        f"df -h {path}",
        shell=True, capture_output=True, text=True
    ).stdout

    du = subprocess.run(
        f"du -sh {path}/* 2>/dev/null | sort -rh | head -10",
        shell=True, capture_output=True, text=True
    ).stdout

    return f"Disk usage for {path}:\\n{df}\\nLargest directories:\\n{du}"

步骤5 — 构建代理

在修改代码之前,作者需完成第5步:构建执行阶段、明确输入参数、指定该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。

from langchain_anthropic import ChatAnthropic
from langchain.agents import AgentExecutor, create_react_agent
from langchain_core.prompts import PromptTemplate
from langchain_core.tools import BaseTool
from typing import List
import os

from agent.tools.bash_tools import run_command, read_file, read_logs, grep_logs
from agent.tools.kubernetes_tools import (
    k8s_get_pods, k8s_describe_pod, k8s_get_logs,
    k8s_get_events, k8s_top_pods
)
from agent.tools.system_tools import (
    get_system_overview, check_service_status,
    check_port, get_top_processes, check_disk_usage
)
SYSTEM_PROMPT = """You are an expert DevOps engineer and SRE with 10 years of experience.
You have access to tools that let you inspect systems, read logs, check services,
and query Kubernetes clusters.
Your approach:
- Start with broad system checks, then narrow down to specifics
- Always check logs when a service is misbehaving
- Look for patterns - one error often points to another
- Explain what you're doing and why as you go
- Give clear, actionable recommendations at the end
- Never run destructive commands - only read and inspect
When diagnosing issues:
1. Get system overview first if it's a general performance issue
2. Check the specific service/pod status
3. Read recent logs for errors
4. Cross-reference with system metrics
5. Provide root cause + recommended fix
You have access to these tools:
{tools}
Use this format:
Thought: [what you're thinking and why]
Action: [tool name]
Action Input: [tool input]
Observation: [what the tool returned]
... (repeat as needed)
Thought: I now have enough information to answer
Final Answer: [clear explanation + recommendations]
Begin!
Question: {input}
{agent_scratchpad}"""

def build_devops_agent(
    include_kubernetes: bool = True,
    verbose: bool = True
) -> AgentExecutor:

    """Build and return the DevOps AI agent."""

    # Initialize Claude
    llm = ChatAnthropic(
        model="claude-sonnet-4-20250514",
        temperature=0,              # deterministic for ops tasks
        max_tokens=4096,
        anthropic_api_key=os.getenv("ANTHROPIC_API_KEY")
    )

    # Collect tools
    tools: List[BaseTool] = [
        run_command,
        read_file,
        read_logs,
        grep_logs,
        get_system_overview,
        check_service_status,
        check_port,
        get_top_processes,
        check_disk_usage,
    ]

    if include_kubernetes:
        tools.extend([
            k8s_get_pods,
            k8s_describe_pod,
            k8s_get_logs,
            k8s_get_events,
            k8s_top_pods,
        ])

    # Build prompt
    prompt = PromptTemplate(
        input_variables=["input", "tools", "tool_names", "agent_scratchpad"],
        template=SYSTEM_PROMPT
    )

    # Create ReAct agent
    agent = create_react_agent(llm, tools, prompt)

    # Wrap in executor
    return AgentExecutor(
        agent=agent,
        tools=tools,
        verbose=verbose,              # streams thinking to terminal
        max_iterations=15,            # prevent infinite loops
        handle_parsing_errors=True,   # recover from malformed outputs
        return_intermediate_steps=True
    )

第6步 — 入口点

import os
from dotenv import load_dotenv
from agent.agent import build_devops_agent

load_dotenv()

def run_interactive():

    """Interactive mode - chat with your agent."""
    print("\\n🤖 DevOps AI Agent")
    print("   Powered by Claude + LangChain")
    print("   Type 'exit' to quit\\n")
    print("   Examples:")
    print("   → Why is nginx returning 502 errors?")
    print("   → Which pod is using the most memory in production?")
    print("   → Check disk space and find what's eating it")
    print("   → Is the database service healthy?\\n")

    agent = build_devops_agent(
        include_kubernetes=True,
        verbose=True
    )

    while True:
        try:
            user_input = input("\\n💬 You: ").strip()
            if not user_input:
                continue

            if user_input.lower() in ["exit", "quit", "q"]:
                print("👋 Goodbye!")
                break
            print("\\n🤖 Agent thinking...\\n")

            result = agent.invoke({"input": user_input})
            print(f"\\n{'='*60}")
            print("FINAL ANSWER:")
            print('='*60)
            print(result["output"])

        except KeyboardInterrupt:
            print("\\n\\n👋 Interrupted. Goodbye!")
            break

        except Exception as e:
            print(f"\\n❌ Agent error: {e}")

def run_single_query(query: str):
    """Run a single query - useful for scripting."""

    agent = build_devops_agent(verbose=False)

    result = agent.invoke({"input": query})
    print(result["output"])

if __name__ == "__main__":
    import sys
    if len(sys.argv) > 1:

        # Single query mode
        run_single_query(" ".join(sys.argv[1:]))
    else:

        # Interactive mode
        run_interactive()

实际应用演示

python main.py
💬 You: The API service is responding slowly. Find out why.
🤖 Agent thinking...

Thought: I'll start with a system overview, then check the API service.
Action: get_system_overview
Action Input: {}
Observation:
  CPU:    87% used
  Memory: 94% used (7.5GB / 8GB)
  Disk:   45% used
  Load:   load average: 6.2, 5.8, 4.1


Thought: Very high CPU and memory. Check what's consuming resources.
Action: get_top_processes
Action Input: {"sort_by": "memory"}
Observation:
  USER    PID    %CPU  %MEM  COMMAND
  api     14821  45.2  67.3  python api_worker.py
  api     14822  38.1  21.4  python api_worker.py
  ...


Thought: Two API worker processes are using 88% of memory combined.
         Let me check the service logs for clues.
Action: read_logs
Action Input: {"service_name": "api-worker", "lines": 100}
Observation:
  2026-05-22 02:31:14 WARNING  Slow query: SELECT * FROM events took 8432ms
  2026-05-22 02:31:19 WARNING  Slow query: SELECT * FROM events took 9103ms
  2026-05-22 02:31:22 ERROR    Connection pool exhausted (pool_size=10)
  2026-05-22 02:31:28 WARNING  Slow query: SELECT * FROM events took 11201ms


Thought: DB connection pool is exhausted and queries are very slow.
         Check if PostgreSQL is healthy.
Action: check_service_status
Action Input: {"service_name": "postgresql"}
Observation:
  ● postgresql.service - PostgreSQL RDBMS
     Loaded: loaded (/lib/systemd/system/postgresql.service)
     Active: active (running) since 2026-05-21 18:00:01 UTC
  May 22 02:28:11 postgres[1842]: LOG: checkpoint taking longer than 30s


Thought: Checkpoints are taking too long - DB is under heavy write load.
         I have enough information now.


Final Answer:


Root Cause: The API is slow because of a database bottleneck, not the app itself.

What's happening:
- PostgreSQL is performing slow checkpoints (>30s), causing all queries to queue
- The `events` table queries are taking 8-11 seconds each (should be <100ms)
- This has exhausted the connection pool (size: 10), causing all new requests to wait

Why it cascaded:
- Slow DB queries → workers hold connections longer
- Connection pool exhausts → new requests queue
- Queue builds up → high CPU from context switching → memory pressure

Immediate actions (do these now):

1. Check if the events table has a missing index:
   → psql -c "EXPLAIN ANALYZE SELECT * FROM events WHERE ..."

2. Temporarily increase connection pool:
   → Set DB_POOL_SIZE=25 and restart api-worker

3. Check for a long-running transaction blocking checkpoints:
   → psql -c "SELECT pid, age(clock_timestamp(), xact_start) FROM pg_stat_activity WHERE xact_start IS NOT NULL ORDER BY xact_start"

Root fix:
- Add an index on the events table (likely missing on timestamp or foreign key column)
- Consider query result caching for high-frequency event reads

添加内存功能 — 多轮对话支持

from langchain.memory import ConversationBufferWindowMemory
from langchain.agents import AgentExecutor, create_react_agent

# Add memory - keeps last 10 exchanges
memory = ConversationBufferWindowMemory(
    k=10,
    memory_key="chat_history",
    return_messages=True
)

# Updated system prompt with memory
SYSTEM_PROMPT_WITH_MEMORY = """...(same as before but add)...
Previous conversation:
{chat_history}
Question: {input}
{agent_scratchpad}"""

# Pass memory to executor
agent_executor = AgentExecutor(
    agent=agent,
    tools=tools,
    memory=memory,
    verbose=True,
    max_iterations=15,
)
💬 You: What's eating disk space on the server?

🤖 Agent: [checks disk] /var/log is using 47GB.
          /var/log/nginx/access.log is 38GB alone.

💬 You: How long has it been growing?
🤖 Agent: [remembers context - checks log rotation config]
          Log rotation isn't configured for nginx.
          The file has been growing unrotated since March 14th.

💬 You: What's the fix?
🤖 Agent: [still has full context]
          Configure logrotate for nginx. Here's the exact config...

将其封装为REST API

# api.py
from fastapi import FastAPI, BackgroundTasks
from pydantic import BaseModel
import uuid
from agent.agent import build_devops_agent

app = FastAPI(title="DevOps AI Agent API")

agent = build_devops_agent(verbose=False)

results_store = {}  # In production: use Redis

class QueryRequest(BaseModel):
    query: str
    environment: str = "production"

class QueryResponse(BaseModel):
    job_id: str
    status: str
    result: str = None

@app.post("/query", response_model=QueryResponse)
async def run_query(request: QueryRequest, background_tasks: BackgroundTasks):
    """Submit a query to the DevOps agent."""
    job_id = str(uuid.uuid4())

    results_store[job_id] = {"status": "running", "result": None}

    def run_agent_task():
        result = agent.invoke({"input": request.query})
        results_store[job_id] = {
            "status": "complete",
            "result": result["output"]
        }

    background_tasks.add_task(run_agent_task)

    return QueryResponse(job_id=job_id, status="running")

@app.get("/query/{job_id}", response_model=QueryResponse)
async def get_result(job_id: str):
    """Poll for query result."""
    job = results_store.get(job_id)

    if not job:
        return QueryResponse(job_id=job_id, status="not_found")

    return QueryResponse(
        job_id=job_id,
        status=job["status"],
        result=job["result"]
    )

# Run: uvicorn api:app --host 0.0.0.0 --port 8000
# From a CI/CD pipeline
curl -X POST <http://agent:8000/query> \\
  -H "Content-Type: application/json" \\
  -d '{"query": "Is the staging deployment healthy after this release?"}'

# From a Makefile
make check-deploy:
  curl -s <http://agent:8000/query> \\
    -d '{"query": "Run post-deploy checks on staging"}' | jq .result

立即可尝试的实际用例

"Nginx is returning 504 errors. Find the root cause."
"The payment service pod keeps crashing. What's happening?"
"Memory usage spiked 30 minutes ago. What caused it?"
"Run a health check on all services and give me a summary."
"Which pods in production are close to their memory limits?"
"Check if any SSL certificates are expiring in the next 30 days."
"At current growth rate, when will we run out of disk space in /var/log?"
"Which services are consistently above 80% CPU?"
"Find any processes listening on unexpected ports."
"Check auth logs for failed login attempts in the last hour."
"Are there any world-writable files in /etc?"

常见错误与解决方案

# Set a hard iteration limit
AgentExecutor(max_iterations=15, ...)  # never go above 20
# Always truncate in your tools
if len(output) > 3000:
    output = output[:3000] + "\\n... [truncated]"
# Don't rely on the LLM to self-police
# Enforce safety in the tool code itself (like we did in safety.py)
# The LLM cannot bypass code-level blocks
# Use streaming to show progress
async for chunk in agent.astream({"input": query}):
    if "actions" in chunk:
        for action in chunk["actions"]:
            print(f"  → Running: {action.tool}({action.tool_input})")

下一步该做什么

核心要点

操作检查清单