Home / Articles / Who writes the code when agents join the workflow?

This article is published in English.

Who writes the code when agents join the workflow?

Map coding assistants from autocomplete to swarms and match tier to blast radius, tests, and review gates.

2200 words

AI coding tools no longer stop at autocomplete. They span reactive pair programmers, task agents that execute a ticket, autonomous project agents that own a branch, and experimental swarms where multiple agents critique and patch each other. The useful question is not which brand to install; it is which tier matches the risk of the change.

From autocomplete to autonomy

Early assistants completed the line under your cursor. Newer agents open files, run tests, and open pull requests. Autonomy rises with blast radius. So should review depth, sandboxing, and the clarity of the goal you hand them.

Tier 1 — AI pair programmers

Reactive autocomplete and chat-in-IDE tools. You steer every commit. Best for boilerplate, renames, and explaining unfamiliar code. Weak at multi-file refactors without guidance.

# Conceptual representation of a Tier 3 agent's execution loop
def execute_project_goal(goal: str):
    plan = agent.generate_plan(goal)
    while not plan.is_complete():
        action = plan.get_next_action()
        result = workspace.run(action)      # Runs commands, edits files
        if result.has_errors():
            plan.replan(result.logs)       # Self-correction loop
        else:
            plan.mark_step_done()

Tier 2 — task-specific agents

Goal-oriented executors: “add logging to this handler,” “write tests for this module.” They plan short tool loops and stop when the goal check passes. Still human-gated for merge.

import time
from typing import Dict, Any

def run_agent_loop(task_prompt: str, max_iterations: int = 15) -> bool:
    # Initialize the agent's state, workspace, and execution context
    state: Dict[str, Any] = {
        "task": task_prompt,
        "workspace_files": get_project_files(),
        "history": [],
        "completed": False
    }

    for step in range(max_iterations):
        # 1. Perception: Observe current system state and tool outputs
        observation = observe_environment(state)

        # 2. Planning: Reason through the current state to generate a thought and next action
        thought, action = LLM_reasoning_engine(state, observation)
        state["history"].append({"step": step, "thought": thought, "action": action})

        if action.name == "task_complete":
            print(f"Task successfully completed in {step} steps.")
            return True

        # 3. Action: Execute the tool and capture the side effects
        try:
            action_result = execute_action(action)
            state = update_state(state, action_result)
        except Exception as execution_error:
            # Feed the error back to the LLM to allow for self-correction
            state = update_state(state, {"error": str(execution_error)})

        time.sleep(1) # Implement rate limiting and token management

    print("Agent failed: Reached maximum iteration budget.")
    return False

Tier 3 — autonomous project agents

End-to-end engineers for a scoped project: scaffold, implement, test, iterate. They need clear acceptance tests and a sealed environment. Without tests they thrash.

import subprocess
import json
import sys

def execute_agentic_workflow(specification_path: str) -> bool:
    """
    Orchestrates an agent by feeding it a structured specification,
    applying the generated code changes, and running unit tests to verify correctness.
    """
    # Step 1: Load the structured technical specification
    with open(specification_path, "r") as f:
        spec = json.load(f)

    print(f"🤖 Agent starting task: {spec['task_id']} - {spec['description']}")

    # Step 2: Agent generates code based on spec (simulated here)
    generated_code_diff = simulate_agent_generation(spec)

    # Step 3: Apply the generated patches to the codebase
    if not apply_patch(generated_code_diff):
        print("❌ Critical: Agent-generated patch failed to apply cleanly.")
        return False

    # Step 4: Run automated validation suites
    print("🧪 Running verification test suite...")
    test_result = subprocess.run(["pytest", "tests/test_agent_features.py"], capture_output=True, text=True)

    if test_result.returncode == 0:
        print("✅ Success: Agent changes verified successfully.")
        return True
    else:
        print("❌ Failure: Automated tests failed. Raw stderr output:")
        print(test_result.stderr)
        return False

def simulate_agent_generation(spec: dict) -> str:
    # Simulates returning a git patch block matching the spec constraints
    return "diff --git a/app.py b/app.py..."

def apply_patch(diff: str) -> bool:
    # Logic to apply git patch
    return True

if __name__ == "__main__":
    execute_agentic_workflow("specs/new_feature_spec.json")

Tier 4 — agentic swarms

Collaborative multi-agent networks: researcher, implementer, reviewer. Powerful and expensive. Coordination bugs become the new failure mode.

+----------------+      +-----------------+      +-----------------+      +-----------------+
|     Define     | ---> |     Prompt      | ---> |      Test       | ---> |     Verify      |
|  (Requirements |      | (Context, Specs |      |   (Automated    |      | (Human Approves |
|  & Interfaces) |      |   & Constraints)|      |  Suites & Runs) |      |   Final Diffs)  |
+----------------+      +-----------------+      +-----------------+      +-----------------+

Anatomy of a coding agent

Perception (repo tools), planning (task decomposition), action (edits/commands), and memory (scratchpads, PR context). Missing any one collapses the tier.

Generated by AI Agent: Provisions a secure, auto-scaling AWS ECS Fargate service
resource "aws_ecs_task_definition" "app" {
  family                   = "production-api"
  requires_compatibilities = ["FARGATE"]
  network_mode             = "awsvpc"
  cpu                      = "256"
  memory                   = "512"

  container_definitions = jsonencode([{
    name      = "api-service"
    image     = "backend-service:latest"
    essential = true
    portMappings = [{
      containerPort = 8080
      hostPort      = 8080
    }]
  }])
}

Choosing a tier

Match tier to repository criticality, test strength, and how reversible the change is. Prefer Tier 1–2 on payment paths until evaluation harnesses exist. Use Tier 3 where CI is strict and the sandbox cannot touch production secrets.

# safe_executor.py
import ast

# An allowlist of pre-approved libraries prevents hallucinated dependency injection.
ALLOWED_PACKAGES = {"requests", "pandas", "numpy", "json", "pydantic"}

def verify_agent_imports(agent_code: str) -> bool:
    """Parses agent code into an AST to audit imports before execution."""
    try:
        tree = ast.parse(agent_code)
        for node in ast.walk(tree):
            if isinstance(node, ast.Import):
                for alias in node.names:
                    base_package = alias.name.split('.')[0]
                    if base_package not in ALLOWED_PACKAGES:
                        raise SecurityError(f"Blocked unapproved import: {alias.name}")
            elif isinstance(node, ast.ImportFrom) and node.module:
                base_package = node.module.split('.')[0]
                if base_package not in ALLOWED_PACKAGES:
                    raise SecurityError(f"Blocked unapproved import from: {node.module}")
        return True
    except (SyntaxError, SecurityError) as e:
        print(f"Safety Gate Tripped: {e}")
        return False

class SecurityError(Exception):
    pass

# Example of agent output containing a hallucinated or malicious library.
untrusted_code = "import requests\nimport fast_json_validator_fake_pkg"
is_safe = verify_agent_imports(untrusted_code)
print(f"Is code safe to execute? {is_safe}") # Prints: Is code safe to execute? False

Workflow habits that keep humans in charge

Write acceptance checks before invoking an agent. Require diffs in reviewable chunks. Log every tool call. Ban unconstrained shell on sensitive hosts. Treat agent velocity as a liability metric when defect rates rise.

The future of “who writes code” is shared authorship with explicit gates—not unsupervised commits into main. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name. Keep fixtures green, measure each stage separately, and refuse to ship on vibes alone when the next release can reintroduce the same silent failure mode under a new name.