This article is published in English.
Practical notes: Your Terraform Agent Is Probably Wrong Half the Time
Operable walkthrough of Practical notes: Your Terraform Agent Is Probably Wrong Half the Time: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Your Terraform Agent Is Probably Wrong Half the Time”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
agent/ the deepagents Terraform agent, its tools, prompts, and skills
eval/ the verifier, 38 tasks with Rego policies, and the benchmark
optimizer/ the DSPy program, metric, and GEPA compile
The Agent We’re Starting With
The The Agent We re stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
from deepagents import create_deep_agent
agent = create_deep_agent(
model="openrouter:openai/gpt-5.6-luna",
tools=[provider_schema, write_terraform, validate_config],
system_prompt=SYSTEM_PROMPT,
skills=["./agent/skills"], # SKILL.md — the thing we'll optimize
)
What failure actually looks like
The What failure actually looks stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
encryption {
kms_key_name = var.encryption_key_name
}
lifecycle_rules {
...
$ tofu validate
Error: Missing required argument
on main.tf line 18, in resource "google_storage_bucket" "terraform_state":
The argument "default_kms_key_name" is required, but no definition was found.
Error: Unsupported argument
on main.tf line 19, in resource "google_storage_bucket" "terraform_state":
An argument named "kms_key_name" is not expected here.
Error: Unsupported block type
on main.tf line 22, in resource "google_storage_bucket" "terraform_state":
Blocks of type "lifecycle_rules" are not expected here. Did you mean
"lifecycle_rule"?
$ uv run python -m agent.run --task eval/tasks/backend-var-interpolation.json
task backend-var-interpolation (opentofu)
model openai/gpt-5.6-luna engine=deepagents
files main.tf, variables.tf
tools {'write_terraform': 2, 'validate_config': 1, 'validate_pass': 1}
elapsed 56602ms
A Scorer Before a Framework: Terraform Grades Its Own Homework
The A Scorer Before a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The A Scorer Before a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
package main
import rego.v1
deny contains "bucket must use a customer-managed KMS key" if {
some name
bucket := input.resource.google_storage_bucket[name][_]
not bucket.encryption
}
$ cd eval/tasks && ./check_policies.sh
...
✓ vpc-subnet-firewall good=0 bad=2
✅ 38/38 policies verified in both directions
Adopting DSPy
For the Adopting DSPy stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
uv add dspy
uv sync
Phase 1: Programming
For the Phase 1 Programming stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
import dspy
class AuthorConfiguration(dspy.Signature):
"""Write a complete, valid Infrastructure-as-Code configuration."""
request: str = dspy.InputField(
desc="What the user wants built, in natural language."
)
target: str = dspy.InputField(
desc="Which dialect to target: 'terraform' or 'opentofu'. These have "
"diverged — code targeting the wrong one will fail validation."
)
config: str = dspy.OutputField(
desc="The complete configuration as fenced HCL blocks. Start each block "
"with a comment naming its file, e.g. '# main.tf'."
)
dspy.inspect_history(n=1)
System message:
Your input fields are:
1. `request` (str): What the user wants built, in natural language.
2. `target` (str): Which dialect to target: 'terraform' or 'opentofu'. ...
Your output fields are:
1. `config` (str): The complete configuration as fenced HCL blocks. ...
[[ ## request ## ]]
{request}
[[ ## target ## ]]
{target}
[[ ## config ## ]]
{config}
In adhering to this structure, your objective is:
Write a complete, valid Infrastructure-as-Code configuration ...
generate = dspy.Predict(AuthorConfiguration) # one shot
generate = dspy.ChainOfThought(AuthorConfiguration) # reason first
generate = dspy.ReAct(AuthorConfiguration, tools=[...]) # run a tool loop
from agent.tools import make_tools # the deployed agent's tools
class TerraformAuthoringAgent(dspy.Module):
def __init__(self, seed_instruction: str):
super().__init__()
# Seed the SIGNATURE before constructing ReAct, so DSPy appends its
# tool protocol to your instruction instead of replacing it.
seeded = AuthorConfiguration.with_instructions(seed_instruction)
lc_tools = make_tools(work_dir, binary="terraform", stats={})
self.react = dspy.ReAct(
seeded,
tools=[dspy.Tool(t.func, name=t.name, desc=t.description)
for t in lc_tools.values()],
max_iters=10,
)
def forward(self, request: str, target: str = "terraform"):
return self.react(request=request, target=target)
Phase 2: Evaluation
For the Phase 2 Evaluation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Phase 2 Evaluation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
dataset = [
dspy.Example(
task_id=t["id"],
request=t["prompt"],
target=t["target"],
rego=t["rego"], # path to this task's policy
).with_inputs("request", "target")
for t in tasks
]
def verify_metric(example, prediction, trace=None, pred_name=None, pred_trace=None):
"""Score with the SAME verifier that produces our benchmark numbers."""
result = verify(
config_dir=materialise(prediction.config),
policy_dir=example.rego,
target=example.target,
)
# Job 1 — bootstrapping (trace is set): a strict bool. Only outputs that
# FULLY pass may become worked examples.
if trace is not None:
return result.passed
# Job 2 — reflective optimization (pred_name is set): score AND feedback.
# GEPA reads the text to understand *why* a candidate failed.
if pred_name is not None:
return dspy.Prediction(score=result.score, feedback=format_failures(result))
# Job 3 — plain evaluation: a float.
return result.score
evaluate = dspy.Evaluate(devset=valset, metric=verify_metric,
num_threads=8, display_table=True)
evaluate(program)
Average Metric: 0.59 / 3 (19.6%): 100%|██████████| 3/3 [01:16<00:00, 25.60s/it]
INFO dspy.evaluate.evaluate: Average Metric: 0.5882 / 3 (19.6%)
WARNING dspy.evaluate.evaluate: Skipping table display since `pandas` is not installed.
Phase 3: Optimization
When working through the Phase 3 Optimization stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
gepa = dspy.GEPA(
metric=verify_metric,
max_metric_calls=150,
reflection_lm=dspy.LM("openrouter/openai/gpt-5.6-luna", max_tokens=16000),
num_threads=8,
track_stats=True,
)
compiled = gepa.compile(student=program, trainset=trainset, valset=valset)
compiled.save("compiled_state.json", save_program=False)
The workflow keeps paying after optimization
When working through the The workflow keeps paying stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
robust = dspy.Refine(module=program, N=3, reward_fn=verify_metric, threshold=1.0)
program.set_lm(dspy.LM("openrouter/qwen/qwen3-coder-30b-a3b-instruct"))
evaluate(program)
The Honest Result
When working through the The Honest Result stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Honest Result stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Conclusion
The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for c9e85f3f670f: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 2/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 3/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 4/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 5 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 5/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 6 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 6/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 7 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 7/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 8 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 8/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 9 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 9/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 10 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 10/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 11 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 11/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 12 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 12/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 13 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 13/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 14 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 14/880: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.