This article is published in English.
Practical notes: The Week Three Real Security Incidents Happened to AI Agents
Operable walkthrough of Practical notes: The Week Three Real Security Incidents Happened to AI Agents: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: The Week Three Real Security Incidents Happened to AI Agents, and What Each One Actually Teaches. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Incident one: the bot that leaked private code because you said “additionally”
When working through the Incident one the bot stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# .github/workflows/agentic-issue-bot.yml
# BEFORE: org-wide token, no membership check, posts directly
name: issue-bot-vulnerable
on:
issues:
types: [assigned]
jobs:
respond:
runs-on: ubuntu-latest
steps:
- name: Fetch context and respond
env:
GH_TOKEN: ${{ secrets.ORG_WIDE_PAT }} # scoped to every repo in the org
run: |
# reads issue.body directly, treats it as trusted instruction
gh issue comment "${{ github.event.issue.number }}" \
--body "$(python3 bot_respond.py "${{ github.event.issue.body }}")"
# AFTER: repo-scoped token, membership gate, draft instead of a live post
name: issue-bot-fixed
on:
issues:
types: [assigned]
jobs:
respond:
runs-on: ubuntu-latest
steps:
- name: Check the issue author is an org member or collaborator
id: gate
env:
GH_TOKEN: ${{ secrets.REPO_SCOPED_PAT }} # scoped to this repo only
run: |
ACTOR="${{ github.event.issue.user.login }}"
ROLE=$(gh api "repos/${{ github.repository }}/collaborators/$ACTOR/permission" \
--jq '.permission' 2>/dev/null || echo "none")
if [ "$ROLE" = "none" ]; then
echo "trusted=false" >> "$GITHUB_OUTPUT"
else
echo "trusted=true" >> "$GITHUB_OUTPUT"
fi
- name: Prepare a draft response instead of posting
if: steps.gate.outputs.trusted == 'true'
env:
GH_TOKEN: ${{ secrets.REPO_SCOPED_PAT }}
run: |
python3 bot_respond.py "${{ github.event.issue.body }}" > draft_reply.md
gh issue edit "${{ github.event.issue.number }}" --add-label "needs-human-review"
# a human reads draft_reply.md and posts it manually, or via a
# second, explicitly human-triggered workflow step
Incident two: the CVE that assumed developers were the trusted side
When working through the Incident two the CVE stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
GITLAB CVE-2026-18252, AFFECTED VS PATCHED
----------------------------------------------------------
BRANCH VULNERABLE RANGE PATCHED VERSION
----------------------------------------------------------
18.x 18.9 - latest 18.x 19.1.7 (upgrade path)
19.1 19.1.0 - 19.1.6 19.1.7
19.2 19.2.0 - 19.2.4 19.2.5
19.3 19.3.0 19.3.1
----------------------------------------------------------
CVSS score: 7.3 (High)
Weakness: CWE-829, inclusion of functionality from an
untrusted control sphere
Required access: authenticated Developer role
GitLab.com SaaS / Dedicated: patched by GitLab, no action needed
Self-managed instances: patch required, urged immediately
----------------------------------------------------------
#!/usr/bin/env bash
# check_gitlab_cve.sh
# Self-hosted, no paid vulnerability scanner required.
# Compares your self-managed GitLab version against GitLab's own
# published security release JSON feed and flags known CVEs.
set -euo pipefail
CURRENT_VERSION=$(curl -s "https://your-gitlab-instance/api/v4/version" \
-H "PRIVATE-TOKEN: $GITLAB_API_TOKEN" | jq -r '.version')
echo "Running GitLab version: $CURRENT_VERSION"
# GitLab publishes security release blog posts with a predictable
# structure; for a production setup, mirror the CVE list into a
# small local file you update whenever GitLab ships a security release,
# and diff your running version against it on a cron job.
KNOWN_VULNERABLE=("18.9.0" "19.1.0" "19.1.6" "19.2.0" "19.2.4" "19.3.0")
for v in "${KNOWN_VULNERABLE[@]}"; do
if [ "$CURRENT_VERSION" = "$v" ]; then
echo "WARNING: running $CURRENT_VERSION, matches a version flagged in CVE-2026-18252 range. Patch now."
exit 1
fi
done
echo "No known match in the local CVE list. Still verify against GitLab's security release page directly."
Incident three: the deepfake that wanted compute budget, not credentials
When working through the Incident three the deepfake stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Incident three the deepfake stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
The checklist: what a small team can actually do this week
The The checklist what a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# spend_approval_gate.py
# A minimal, self-hosted out-of-band verification gate.
# No paid identity-verification vendor required, just a shared
# secret rotated on your own schedule and a callback number you
# already have on file, never one supplied in the request itself.
import hmac
import time
THRESHOLD_USD = 50_000
# Rotate this weekly; store it somewhere the approval requester
# (or an attacker impersonating them) never has access to, e.g. a
# password manager entry only finance leads can see.
CURRENT_PASSPHRASE = "harbor-quiet-tuesday"
def requires_out_of_band_check(amount_usd: float) -> bool:
return amount_usd >= THRESHOLD_USD
def verify_out_of_band(spoken_phrase: str) -> bool:
# constant-time compare so a partial match can't be timed out
return hmac.compare_digest(spoken_phrase.strip().lower(),
CURRENT_PASSPHRASE.lower())
def approve_spend(amount_usd: float, spoken_phrase: str = "") -> str:
if not requires_out_of_band_check(amount_usd):
return "approved"
if verify_out_of_band(spoken_phrase):
return "approved after out-of-band verification"
return "BLOCKED: verify via a known callback number or the current passphrase before approving"
What all three actually have in common
The What all three actually stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 9590a2274c4f: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 0/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 1/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 2/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 3/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 4/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 5 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 5/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 6 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 6/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 7 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 7/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 8 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 8/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 9 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 9/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 10 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 10/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 11 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 11/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 12 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 12/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 13 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 13/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 14 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 14/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 15 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 15/772: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.