Home / Articles / Practical notes: Building a Production-Ready Agentic Fraud Detection System

This article is published in English.

Practical notes: Building a Production-Ready Agentic Fraud Detection System

Operable walkthrough of Practical notes: Building a Production-Ready Agentic Fraud Detection System: contracts, checks, and drop-in code slots for teams shipping this pattern.

2383 words

This walkthrough rebuilds the path from raw materials to a working system for: Building a Production-Ready Agentic Fraud Detection System — Part 1: The Full Picture. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

The shape of the problem

When working through the The shape of the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

System overview

When working through the System overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Agentic Orchestration — Real Time Fraud Scoring

When working through the Agentic Orchestration Real Time stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Agent A — propensity scoring (the trained model)

When working through the Agent A propensity scoring stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Agent B — behavioral scoring (no model at all)

When working through the Agent B behavioral scoring stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Agent C — policy retrieval over pgvector

When working through the Agent C policy retrieval stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Harness engineering: the parts that don’t show up in a model card

When working through the Harness engineering the parts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

"""
SessionMemory — the public interface tying working memory (short-term) and
episodic memory (long-term, cross-session) together for context engineering.

    mem = SessionMemory()                      # new session, fresh UUID
    mem.remember("user", "...")                # auto-persists to pgvector
    mem.remember("assistant", "...")
    context = mem.build_context(next_question) # summary + recent turns + relevant past episodes
    mem.end_session()                          # finalize into episodic_memory
"""
"""
In-memory semantic cache in front of episodic recall.

Two layers:
  - exact-string embedding cache: a literal repeat query skips the OpenAI
    embeddings API call entirely.
  - semantic result cache: a near-duplicate query (different wording, same
    intent) still needs embedding to compare, but skips the Postgres/pgvector
    round-trip if it's cosine-similar enough to something already cached.

Process-local only (not shared across workers/processes) — fine for a
single running service, not a substitute for a distributed cache if this
ever runs behind multiple instances.

Must be invalidated when a new episode is saved: a cached "no good match" or
partial result set can go stale the moment the underlying corpus changes.
"""
"""
Fraud-detection event chain, built on the generic pub-sub bus in
common/pubsub.py:

    UserQueryEvent
        -> InputGuardrailService   -> InputGuardrailPassedEvent | InputGuardrailBlockedEvent
        -> OrchestrationService    -> OrchestrationCompletedEvent | OrchestrationPausedHITLEvent
        -> OutputGuardrailService  -> OutputGuardrailPassedEvent | OutputGuardrailBlockedEvent
        -> ResultPublisher         -> PipelineCompletedEvent

    HumanDecisionEvent (resumes a run paused at OrchestrationPausedHITLEvent)
        -> OrchestrationService    -> ... (same chain onward)

Each service subscribes to exactly one (or two, for resume) event type and
publishes the next event in the chain — a new listener (audit logger,
LangSmith exporter) can subscribe to any event without touching the
publishers. The guardrail/graph calls underneath are synchronous
(psycopg2, HF/OpenAI SDKs); handlers run them via asyncio.to_thread so a
slow call doesn't block the event loop for other in-flight events.
"""
@mcp.tool()
@guarded(action="score_transaction")
def score_propensity(features: dict[str, float], role: str) -> float:
    """Score one transaction's V1-V28 features for ML fraud probability (0-1). Agent A.
    Requires role: analyst or admin."""
    return score_customer_propensity(features)

@mcp.tool()
@guarded(action="score_transaction")
def score_behavior(customer_id: int, amount: float, category: str, merchant: str, role: str) -> dict:
    """Score one transaction's behavior anomaly vs its category's peer-cohort stats. Agent B.
    Requires role: analyst or admin."""
    return score_customer_behavior(customer_id, amount, category, merchant)

@mcp.tool()
@guarded(action="chat_query", free_text_arg="query")
def consult_fraud_policy(query: str, role: str, k: int = 3) -> list[dict]:
    """Semantic search over the indexed credit-card policy/handbook documents. Agent C.
    Requires role: viewer or higher. `query` is scanned for prompt injection/jailbreak and PII."""
    return consult_policy(query, k=k)


if __name__ == "__main__":
    mcp.run()
JWT_SECRET_KEY = os.environ["JWT_SECRET_KEY"]
JWT_ALGORITHM = os.environ.get("JWT_ALGORITHM", "HS256")
JWT_EXPIRE_MINUTES = int(os.environ.get("JWT_EXPIRE_MINUTES", "30"))
pwd_context = CryptContext(schemes=["bcrypt"], deprecated="auto")
oauth2_scheme = OAuth2PasswordBearer(tokenUrl="token")

def verify_password(plain_password: str, hashed_password: str) -> bool:
    return pwd_context.verify(plain_password, hashed_password)

def get_password_hash(password: str) -> str:
    return pwd_context.hash(password)

def get_user(username: str) -> dict | None:
    conn = get_connection()
    try:
        with conn.cursor() as cur:
            cur.execute("SELECT * FROM users WHERE username = %s", (username,))
            return cur.fetchone()
    finally:
        conn.close()

def authenticate_user(username: str, password: str) -> dict | None:
    user = get_user(username)
    if not user or user["disabled"]:
        return None
    if not verify_password(password, user["hashed_password"]):
        return None
    return user

Observability

When working through the Observability stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Observability stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

AWS Deployment and the frontend

The AWS Deployment and the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

What’s next

The What s next stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 5d0d59733252: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 0/721: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 1/721: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 2/721: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 3/721: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 4/721: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.