This article is published in English.
Practical notes: Build → Test → Fix → Repeat: Automating Laravel Development
Operable walkthrough of Practical notes: Build → Test → Fix → Repeat: Automating Laravel Development: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Build → Test → Fix → Repeat: Automating Laravel Development with AI Agents. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
A Few Terms, Before We Go Further
When working through the A Few Terms Before stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Why “Can AI Write Code” Is No Longer the Interesting Question
When working through the Why Can AI Write stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Developer → AI → Developer → Test → Developer → AI → Developer → Test → ...
Developer → Orchestrator → Build → Test → (failed? → Fix → Test again) → Review → Done
Meet the Two Main Tools
When working through the Meet the Two Main stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Laravel Boost — a translator between the AI and our application
When working through the Laravel Boost a translator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
composer require laravel/boost --dev
php artisan boost:install
OpenCode — where the agent actually does the work
When working through the OpenCode where the agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Hands-On: Wiring Boost into OpenCode
When working through the Hands-On Wiring Boost into stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
1. Connect OpenCode to Boost
When working through the 1 Connect OpenCode to stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"laravel-boost": {
"type": "local",
"command": ["php", "artisan", "boost:mcp"],
"enabled": true
}
}
}
opencode mcp list
2. Create agents with specific responsibilities
When working through the 2 Create agents with stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 2 Create agents with stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
---
description: Build Agent for new features
mode: primary
tools:
write: true
edit: true
---
You are the Build Agent.
Use Laravel Boost tools to understand the application structure
before writing any code. Never assume the database structure -
always check the schema first. Write tests for every new behavior.
At the end of your work, return ONLY JSON in this format:
{"status": "success", "files_changed": [...], "summary": "..."}
opencode agent list
3. Try running it manually first
The 3 Try running it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
opencode run --agent builder "Add a task assignment feature to TaskFlow according to the requirements in TASK.md"
The Real Case: A Task-Assignment Feature in “TaskFlow”
The The Real Case A stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
1. A task has an assignee column (belongs to a User).
2. A user can only assign tasks within a project they're a member of.
3. A user must not assign or access tasks from other projects.
4. All of the above must be covered by automated tests.
Why We Need a “Logbook” (Database) for This Process
The Why We Need a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Why We Need a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Schema::create('ai_tasks', function (Blueprint $table) {
$table->id();
$table->string('type');
$table->string('status')->default('pending');
$table->text('prompt');
$table->json('result')->nullable();
$table->unsignedTinyInteger('iteration')->default(0);
$table->timestamps();
});
class AiTask extends Model
{
protected $fillable = [
'type', 'status', 'prompt', 'result', 'iteration',
];
protected function casts(): array
{
return ['result' => 'array'];
}
}
pending → building → testing → (fixing → testing)* → reviewing → completed / failed
Bridging OpenCode into Our PHP Code
For the Bridging OpenCode into Our stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
interface AgentRunner
{
public function run(string $agent, string $prompt): array;
}
class OpenCodeAgentRunner implements AgentRunner
{
public function run(string $agent, string $prompt): array
{
$result = Process::timeout(600)->run([
'opencode', 'run',
'--agent', $agent,
'--format', 'json',
$prompt,
]);
if (! $result->successful()) {
return [
'status' => 'error',
'error' => $result->errorOutput(),
];
}
return json_decode($result->output(), true) ?? [
'status' => 'error',
'error' => 'Agent output was not valid JSON',
];
}
}
$this->app->bind(AgentRunner::class, OpenCodeAgentRunner::class);
A Job for Every Stage
For the A Job for Every stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Build Job
For the Build Job stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
class BuildFeature implements ShouldQueue
{
public function __construct(protected AiTask $task) {}
public function handle(AgentRunner $agent): void
{
$result = $agent->run(
agent: 'builder',
prompt: $this->task->prompt,
);
$this->task->update([
'status' => 'testing',
'result' => $result,
]);
RunTests::dispatch($this->task);
}
}
For the Build Job stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Test Job
When working through the Test Job stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
class RunTests implements ShouldQueue
{
public function __construct(protected AiTask $task) {}
public function handle(): void
{
$result = Process::timeout(300)->run('php artisan test --compact');
if ($result->successful()) {
$this->task->update(['status' => 'reviewing']);
ReviewFeature::dispatch($this->task);
return;
}
$this->task->update(['status' => 'failed']);
AnalyzeFailure::dispatch($this->task, $result->output());
}
}
Test
├── passed → move to Review
└── failed → move to Analyze
Don’t Hand Raw Errors to the Agent
When working through the Don t Hand Raw stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
{
"status": "failed",
"failed_tests": [
{
"name": "user_cannot_assign_task_from_other_project",
"error": "Expected response status code [403] but received 200."
}
]
}
You are the Debug Agent.
Analyze the failed tests below. DO NOT modify any files.
Inspect: relevant models, policies, migrations, and tests.
Return JSON with:
- root_cause
- affected_files
- recommended_fix
- risk_of_regression
{
"status": "analyzed",
"root_cause": "TaskPolicy does not verify project membership before allowing assignment",
"affected_files": ["app/Policies/TaskPolicy.php"],
"recommended_fix": "Add a project membership check inside the assign() method",
"next_action": "fix"
}
Fix, Then Test Again — Not Fix, Then Done
When working through the Fix Then Test Again stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
class FixFeature implements ShouldQueue
{
public function __construct(protected AiTask $task, protected array $analysis) {}
public function handle(AgentRunner $agent): void
{
$result = $agent->run(
agent: 'fixer',
prompt: json_encode($this->analysis),
);
$this->task->increment('iteration');
$this->task->update([
'status' => 'testing',
'result' => $result,
]);
RunTests::dispatch($this->task);
}
}
When working through the Fix Then Test Again stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The Orchestrator: The “Traffic Controller” Deciding What’s Next
The The Orchestrator The Traffic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Bus::chain([
new BuildFeature($task),
new RunTests($task),
new ReviewFeature($task),
])->dispatch();
class AgentOrchestrator
{
public function next(AiTask $task): void
{
match ($task->status) {
'pending' => BuildFeature::dispatch($task),
'testing' => RunTests::dispatch($task),
'failed' => AnalyzeFailure::dispatch($task),
'fixing' => RunTests::dispatch($task),
'reviewing' => ReviewFeature::dispatch($task),
'completed', 'stopped' => null,
default => throw new LogicException("Unknown status: {$task->status}"),
};
}
}
A Cap on Retries — So It Doesn’t Loop Forever
The A Cap on Retries stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Fix → Test fails → Fix → Test fails → Fix → Test fails → ...
class RunTests implements ShouldQueue
{
protected const MAX_ITERATIONS = 5;
public function handle(): void
{
if ($this->task->iteration >= self::MAX_ITERATIONS) {
$this->task->update(['status' => 'needs_human_review']);
return;
}
// ... run tests as usual
}
}
Green Tests ≠ Feature Done
The Green Tests Feature Done stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Green Tests Feature Done stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
public function assign(User $user, Task $task): bool
{
return $user->isAdmin();
}
Build → Tests pass → Deploy immediately
Build → Test → (if failed: Debug → Fix → Test again) → Human review → Deploy
If Multiple Agents Work in Parallel, Watch Out for File Conflicts
For the If Multiple Agents Work stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Agent A → writes Task.php
Agent B → reads Task.php (nearly at the same time)
Agent A → writes Task.php again
Research Agent → read-only
Review Agent → read-only
Security Agent → read-only
Build Agent → write access
Fix Agent → write access
Don’t Make One Agent Do Everything
For the Don t Make One stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Planner → designs the approach
Builder → writes the new feature's code
Tester → runs the test suite
Debugger → diagnoses failures (read-only)
Fixer → executes the fix
Reviewer → final quality check (read-only)
A Minimal Checklist Before Actually Using This
For the A Minimal Checklist Before stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the A Minimal Checklist Before stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
So What’s Laravel Boost’s Actual Role in All This?
When working through the So What s Laravel stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Developer (us)
│
▼
Orchestrator (traffic controller, running on Laravel Queue)
│
├── Build Agent
├── Test Agent
├── Debug Agent
└── Fix Agent
│
▼
OpenCode (where the agent does its work)
│
▼
Laravel Boost (context translator, via MCP)
│
▼
Our Laravel application
Conclusion: The Question Has Changed
When working through the Conclusion The Question Has stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 815696fa9b90: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.