Home / Articles / Practical notes: Your Agent Can Fix Its Own Prompt. Here’s How.

This article is published in English.

Practical notes: Your Agent Can Fix Its Own Prompt. Here’s How.

Operable walkthrough of Practical notes: Your Agent Can Fix Its Own Prompt. Here’s How.: contracts, checks, and drop-in code slots for teams shipping this pattern.

4018 words

This walkthrough rebuilds the path from raw materials to a working system for: Your Agent Can Fix Its Own Prompt. Here’s How.. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Teach your agent to learn from its own mistakes and build a better version of itself

When working through the Teach your agent to stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

The agent

When working through the The agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

You are a helpful company information assistant.
You have the following knowledge about company policies:
- PTO: 20 days per year, accrued monthly. Up to 5 unused days roll over.
- Sick leave: 10 days per year, does not roll over.
- Remote work: Up to 3 days per week with manager approval.
- Benefits: The company offers competitive benefits.
Answer questions using only the information above. If a question is about
a topic not listed above, tell the user you do not have that information
and suggest they contact HR.

The building blocks

When working through the The building blocks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the The building blocks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The improvement cycle

The The improvement cycle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

git clone https://github.com/GoogleCloudPlatform/BigQuery-Agent-Analytics-SDK.git
cd examples/agent_improvement_cycle

export PROJECT_ID=<your-project-id>

./setup.sh
./run_cycle.sh               # single cycle, 10 questions, ~3-4 min
./run_cycle.sh --auto --cycles 3 --traffic-count 100

Step by step

The Step by step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Pre-flight: run the golden eval set

The Pre-flight run the golden stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Pre-flight run the golden stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

{
  "eval_cases": [
    {"id": "pto_balance",     "question": "How many PTO days do I get per year?",
     "category": "pto",        "expected_tool": "lookup_company_policy"},
    {"id": "sick_leave_days", "question": "How many sick days do I have?",
     "category": "sick_leave", "expected_tool": "lookup_company_policy"},
    {"id": "remote_work_days","question": "How many days can I work from home?",
     "category": "remote_work","expected_tool": "lookup_company_policy"}
  ]
}
▶ PRE-FLIGHT: Verifying golden eval set passes with current prompt

PASS: pto_balance
     Answer: You receive 20 PTO days per year, accrued monthly.
             Up to 5 unused days can roll over.
     Tools called: lookup_company_policy
PASS: sick_leave_days
     Answer: You have 10 sick days per year. They do not roll over.
     Tools called: lookup_company_policy
PASS: remote_work_days
     Answer: You can work from home up to 3 days per week
             with manager approval.
     Tools called: lookup_company_policy
All cases pass.

Step 1–2: Generate synthetic traffic and run through agent

For the Step 1 2 Generate stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

▶ STEP 1/5: GENERATE SYNTHETIC TRAFFIC
Goal:    Produce diverse user questions that differ from the golden eval set
  Method:  Gemini generates 100 questions

  ...

  Sample questions:
    - How many PTO days do I earn per month?
    - What's the maximum number of PTO days I can roll over into the next year?
    - I want to take a full week off for vacation. What is the advance notice required?

  ...

▶ STEP 2/5: RUN TRAFFIC THROUGH AGENT
Goal:    Send questions to the agent, log every session to BigQuery
...
  [62/100] traffic_expenses_window_to_submit: What is the timeframe for submitting expense reports?
           -> Business expenses must be submitted within 30 days.
  [64/100] traffic_benefits_health_company_contribution: What percentage of my health insurance premium does the company cover?
           -> The company covers 80% of health insurance premiums for employees.
  [65/100] traffic_benefits_dental_coverage: What does the dental insurance plan cover?
           -> I do not have specific information about what the dental insurance plan covers. The company offers competitive benefits, but details about specific plans like dental insurance are not something I have access to. Please contact HR for more information on dental insurance coverage.
  [67/100] traffic_benefits_401k_match: What is the company's 401k match policy?
           -> I do not have information about the company's 401k match policy. Please contact HR for details on this topic.
  [68/100] traffic_benefits_401k_vesting: How long do I need to be employed to be fully vested in the 401k match?
           -> I do not have information about 401k vesting. Please contact HR for details regarding the 401k match and vesting schedule.
  [69/100] traffic_benefits_parental_leave_primary: What is the parental leave duration for a primary caregiver?
           -> I do not have information about parental leave. Please contact HR for details on this policy.
  [66/100] traffic_benefits_vision_frames: How often can I get new frames under the vision plan?
           -> I do not have information about the vision plan or how often you can get new frames. Please contact HR for details about the vision plan.
  [60/100] traffic_expenses_receipt_required_over_x: Expenses over what amount require a receipt?
           -> I do not have information about the company's expense policy. Please contact HR for details.
  [89/100] traffic_holidays_next_company_holiday: When is the next company holiday?
           -> I do not have information about company holidays. Please contact HR for details.
...
Answer questions using only the information above. If a question is about
a topic not listed above, tell the user you do not have that information
and suggest they contact HR.

Step 3: Evaluate quality

For the Step 3 Evaluate quality stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

QUALITY SUMMARY
  Total sessions evaluated : 100
  Meaningful               : 64
  Partial                  : 1
  Unhelpful                : 35
  Unhelpful rate           : 35.0%

[response_usefulness]
    HELPFUL         :   64  ( 64.0%) ################################
    NOT HELPFUL     :   35  ( 35.0%) #################
    PARTIAL         :    1  (  1.0%)
  [task_grounding]
    GROUNDED        :   90  ( 90.0%) #############################################
    NOT GROUNDED    :    9  (  9.0%) ####
    NO TOOL NEEDED  :    1  (  1.0%)
UNHELPFUL Sessions (showing 10 of 35)

Session:     52282245-d2a4-4998-9ad9-4373612019e8  → company_info_agent
  Question:    For what amount do I need to submit a receipt for expenses?
  Response:    "I do not have information about the specific amount for which you need to submit a receipt for expenses. Please contact HR for details on the expense policy."
  Usefulness:    ❌ NOT HELPFUL
  Reason:        The agent explicitly states it does not have the requested information and directs the user elsewhere, thus not providing a substantive answer to the user's question.
  Grounding:     ❌ NOT GROUNDED
  Reason:        The agent's response indicates a lack of specific information and directs the user to HR, suggesting it did not retrieve specific data from a tool to answer the question. The response is a generic LLM statement rather than tool-derived data.

Session:     3821b327-43d8-47c6-b922-dd8d7988c366  → company_info_agent
  Question:    After how many days of illness is a doctor's note mandatory?
  Response:    "I do not have information about when a doctor's note is mandatory for illness. Please contact HR for this information."
  Usefulness:    ❌ NOT HELPFUL
  Reason:        The agent explicitly states it does not have the requested information and directs the user elsewhere, failing to answer the user's question substantively.
  Grounding:     ❌ NOT GROUNDED
  Reason:        The agent's response indicates a lack of information, meaning it was not grounded in actual data retrieved from its tools. The question required tool usage, but no relevant data was found or presented.
...
Metric              V1              Budget        Status
-----------         -----------     -----------   ------
Avg latency         1482.3 ms       10000 ms      PASS
Total tokens        811.4 tokens    50000 tokens  PASS
Turn count          1 turns         10 turns      PASS
Tool error rate     0.0 rate        0.1 rate      PASS

Step 4: Improve the prompt

For the Step 4 Improve the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Step 4 Improve the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Extracted 12 failed cases, added 12 new to golden set (15 total).
Q: As a primary caregiver, how many weeks of parental leave am I eligible for?
  Agent:   I do not have information about parental leave. Contact HR.
  Teacher: As a primary caregiver, you are eligible for 16 weeks of paid parental leave.
Q: What percentage of my contributions does the company match for 401k?
  Agent:   I do not have information about 401k matching. Contact HR.
  Teacher: The company matches 4% of your contributions to the 401k. Fully vested after 1 year.
Q: When is the next company holiday?
  Agent:   I do not have information about company holidays. Contact HR.
  Teacher: The next company holiday is May 25, 2026.
Q: What are the core hours I need to be available if I'm working remotely?
  Agent:   I do not have information about core hours. Contact HR.
  Teacher: Core collaboration hours for remote employees are 10am-3pm in your local timezone.
Calling Vertex AI Prompt Optimizer with 12 ground truth examples...
(The optimizer is a server-side job -- typically 2-4 minutes.)
... still optimizing (15s elapsed)
... still optimizing (30s elapsed)
... still optimizing (45s elapsed)
Optimizer returned a candidate prompt.
PASS: pto_balance
       Question: How many PTO days do I get per year?
       Answer: You get 20 PTO days per year, which are accrued monthly...
       Tools called: lookup_company_policy | Expected: lookup_company_policy
  PASS: extracted_as_a_primary_caregiver_how_many_weeks_of
       Question: As a primary caregiver, how many weeks of parental leave am I eligible for?
       Answer: As a primary caregiver, you are eligible for 16 weeks of paid parental leave.
       Tools called: lookup_company_policy | Expected: lookup_company_policy
  PASS: extracted_what_percentage_of_my_contributions_does
       Question: What percentage of my contributions does the company match for 401k?
       Answer: The company matches 4% of your 401(k) contributions, and you are fully vested after 1 year of employment.
       Tools called: lookup_company_policy | Expected: lookup_company_policy
  PASS: extracted_what_does_the_dental_insurance_plan_cove
       Question: What does the dental insurance plan cover?
       Answer: The dental insurance plan covers full preventive care and 80% for major procedures.
       Tools called: lookup_company_policy | Expected: lookup_company_policy
  PASS: extracted_when_is_the_next_company_holiday
       Question: When is the next company holiday?
       Answer: The next company holiday is Memorial Day on May 25, 2026.
       Tools called: lookup_company_policy, get_current_date | Expected: lookup_company_policy
  ...
  15/15 passed

Prompt:      V1 -> V2
  Golden set:  3 -> 15 cases

Step 5: Measure improvement

When working through the Step 5 Measure improvement stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

[44/100] traffic_expenses_004: I booked a flight for $600 for a business trip. Do I need pre-approval?
           -> Travel expenses over $500 require pre-approval from your manager. Since your flight is $600, you will need pre-approval.
  [80/100] traffic_benefits_020: If I leave the company after 6 months, will I get to keep the 401k match?
           -> No, the company 401(k) match is fully vested after 1 year of employment. If you leave after 6 months, you will not get to keep the 401k match.
  [84/100] traffic_holidays_004: Is the day after Thanksgiving a paid holiday?
           -> Yes, the day after Thanksgiving is a paid holiday.
  [19/100] traffic_sick_004: I was sick for two days; do I need to provide any documentation?
           -> Employees receive 10 sick days per year. A doctor's note is required for absences longer than 3 consecutive days. Since you were sick for two days, you do not need to provide any documentation.
QUALITY SUMMARY
  Total sessions evaluated : 100
  Meaningful               : 99
  Partial                  : 1
  Unhelpful                : 0
  Unhelpful rate           : 0.0%

[response_usefulness]
    HELPFUL         :   99  ( 99.0%) #################################################
    PARTIAL         :    1  (  1.0%)
  [task_grounding]
    GROUNDED        :   99  (100.0%) ##################################################
Session:     4e0ea11d-d4dc-4a59-b337-30415a595c90  → company_info_agent
  Question:    If I'm sick for more than 3 days, what kind of doctor's note is required?
  Response:    "If you are sick for more than 3 consecutive days, a doctor's note is required."
  Usefulness:    PARTIAL
  Reason:        The response confirms a doctor's note is required but does not specify
                 the 'kind' of note, which was part of the user's question.
  Grounding:     GROUNDED
  Reason:        The agent's response directly reflects the information retrieved
                 from the 'lookup_company_policy' tool.
CYCLE 1 RESULTS
  Before (V1):  64.0% meaningful  (64/100 sessions)
  After  (V2):  99.0% meaningful  (98/99 sessions)
Quality 99.0% meets threshold (95%) -- stopping auto-continue.

DONE  (total wall time: 12m 39s)
  Prompt version:   V2
  Golden eval set:  15 cases
Metric              V1              V2              Budget        Status
-----------         -----------     -----------     -----------   ------
Avg latency    (v)  1482.3 ms       1088.4 ms       10000 ms      PASS
Total tokens   (^)  811.4 tokens    1339.7 tokens   50000 tokens  PASS
Turn count     (=)  1 turns         1 turns         10 turns      PASS
Tool error     (=)  0.0 rate        0.0 rate        0.1 rate      PASS

The V2 prompt

When working through the The V2 prompt stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

You are a helpful company information assistant. Your primary function
is to answer employee questions about company policies by using the
available tools.

Core Directives:
1. Tool-First Approach: For EVERY user question, your first and only
   action should be to use one of the provided tools to find the answer.
2. No Answering from Memory: Do not use any general knowledge. The
   tools are the only source of truth.
3. Mandatory Tool Use: You MUST call the appropriate tool to answer the
   question. Do not state that you don't have the information or direct
   the user to HR for topics that the tools can handle.
4. Topic Inference: Carefully analyze the user's prompt to determine
   the correct topic parameter for the lookup_company_policy tool.
   The user's language may not be an exact match for the available
   topics (e.g., 'parental leave' or '401k' should be mapped to
   the 'benefits' topic).
AVAILABLE TOOLS:
- lookup_company_policy(topic: str)
  - Looks up a company policy by topic.
  - topic: The policy topic to look up. Must be one of: pto, sick_leave,
    remote_work, expenses, benefits, holidays.
- get_current_date()
  - Gets the current date.
Your goal is to successfully call the correct tool with the correct
parameters based on the user's question.

What the cycle teaches about prompt design

When working through the What the cycle teaches stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the What the cycle teaches stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Running it yourself

The Running it yourself stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

git clone https://github.com/GoogleCloudPlatform/BigQuery-Agent-Analytics-SDK.git
cd examples/agent_improvement_cycle

export PROJECT_ID=<your-project-id>

./setup.sh

./run_cycle.sh
./reset.sh

The takeaway

The The takeaway stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for f7bfa970ccb5: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 0/963: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 1/963: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 2/963: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 3/963: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.