This article is published in English.
Practical notes: Numasec | The AI Agent for Cybersecurity
Operable walkthrough of Practical notes: Numasec | The AI Agent for Cybersecurity : contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “ Numasec | The AI Agent for Cybersecurity ”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
What Is Numasec?
The What Is Numasec stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Nmap → Network discovery
Nuclei → Vulnerability templates
SQLMap → SQL injection testing
FFUF → Content discovery
Nikto → Web-server checks
Trivy → Security scanning
🤖 AI Agent
│
┌──────────┼──────────┐
▼ ▼ ▼
Tools Runbooks Knowledge
│ │ │
└──────────┼──────────┘
▼
Operation
│
┌───────────┼───────────┐
▼ ▼ ▼
Findings Evidence Replay
│ │ │
└───────────┼───────────┘
▼
Report
Why AI Agents Are Interesting for Cybersecurity
The Why AI Agents Are stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Target
Scope
Tools
Observations
Findings
Evidence
Risk
Remediation
Numasec’s Security Workflow
The Numasec s Security Workflow stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Numasec s Security Workflow stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
🎯 Target
↓
📋 Scope
↓
🧭 Security Posture
↓
📖 Runbook
↓
🛠️ Local Tools
↓
🔎 Observations
↓
🚨 Findings
↓
📸 Evidence
↓
🔁 Replay / Verification
↓
📊 Report
Security Reconnaissance With AI
For the Security Reconnaissance With AI stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Domains
Subdomains
Technologies
Ports
Services
APIs
Authentication
Web applications
Cloud services
Raw Tool Output
↓
AI Interpretation
↓
Structured Observation
↓
Potential Finding
↓
Evidence
Working With Existing Security Tools
For the Working With Existing Security stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
AI replaces security tools
AI
│
├── Nmap
├── FFUF
├── Nuclei
├── Nikto
├── SQLMap
├── Trivy
└── Other authorized tools
Security Agents and Different Modes
For the Security Agents and Different stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Security Agents and Different stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Numasec
│
┌───────────┼───────────┐
▼ ▼ ▼
AppSec Pentest Research
│ │ │
▼ ▼ ▼
APIs Network CVEs
Web Systems Advisories
Runbooks: Turning Security Knowledge Into Workflows
When working through the Runbooks Turning Security Knowledge stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Web Application Assessment
1. Identify target
2. Confirm scope
3. Inspect application
4. Identify technologies
5. Map endpoints
6. Analyze authentication
7. Review APIs
8. Identify potential vulnerabilities
9. Collect evidence
10. Validate findings
11. Generate report
Findings Should Require Evidence
When working through the Findings Should Require Evidence stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Interesting response
↓
"Potential vulnerability"
Observation
↓
Candidate
↓
Verification
↓
Evidence
↓
Confirmed Finding
Observation ≠ Vulnerability
When working through the Observation Vulnerability stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Observation Vulnerability stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Potential vulnerability
Candidate
Observed
Verified
Rejected
Stale
Evidence Collection
The Evidence Collection stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
HTTP Requests
HTTP Responses
Screenshots
Tool Output
Logs
Hashes
Configuration
Reproduction Steps
Downloads/
Screenshots/
Terminal History/
Notes/
Browser Tabs/
Finding #001
│
├── Observation
├── Request
├── Response
├── Screenshot
├── Reproduction
└── Remediation
Replay and Verification
The Replay and Verification stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Initial Observation
↓
Hypothesis
↓
Controlled Test
↓
Evidence
↓
Replay
↓
Verified Finding
From Testing to Reporting
The From Testing to Reporting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The From Testing to Reporting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Executive Summary
Technical Findings
Severity
Evidence
Impact
Remediation
References
Terminal
+
Screenshots
+
Notes
+
Browser
+
Scanner Output
Getting Started
For the Getting Started stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
npm install -g numasec
numasec
Your Machine
│
▼
Local Web Application
│
▼
Numasec
│
▼
AppSec Runbook
│
▼
Findings + Evidence
The /doctor Concept
For the The doctor Concept stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Installed tools
Missing tools
Broken tools
Incorrect versions
Unavailable dependencies
Scope Is Critical
For the Scope Is Critical stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Scope Is Critical stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
AUTHORIZED TARGET
↓
DEFINED SCOPE
↓
ALLOWED TESTS
↓
CONTROLLED EXECUTION
Allowed:
example-lab.local
Not allowed:
production.example.com
third-party.example.net
unrelated infrastructure
Security Automation Needs Guardrails
When working through the Security Automation Needs Guardrails stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Scope
When working through the Scope stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Permissions
When working through the Permissions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Permissions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Evidence
The Evidence stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Logging
The Logging stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Human Oversight
The Human Oversight stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Human Oversight stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Safe Environments
For the Safe Environments stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Numasec vs Traditional Security Workflows
For the Numasec vs Traditional Security stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Terminal
+
Browser
+
Burp
+
Nmap
+
Scanner
+
Notes
+
Screenshots
+
Report
🤖 AI Agent
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Terminal Browser Tools
│ │ │
└─────────────┼─────────────┘
▼
Operation
│
┌────────┴────────┐
▼ ▼
Findings Evidence
│ │
└────────┬────────┘
▼
Report
Why This Approach Is Interesting
For the Why This Approach Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Why This Approach Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Example Authorized Workflow
When working through the Example Authorized Workflow stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
1️⃣ Define scope
↓
2️⃣ Start Numasec
↓
3️⃣ Check local tools
↓
4️⃣ Select AppSec posture
↓
5️⃣ Start appropriate runbook
↓
6️⃣ Discover application surface
↓
7️⃣ Analyze observations
↓
8️⃣ Validate interesting behavior
↓
9️⃣ Capture evidence
↓
🔟 Generate report
Architecture
When working through the Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
👨💻 Operator
│
▼
Terminal Console
│
▼
AI Security
Agent
│
┌───────────────┼────────────────┐
▼ ▼ ▼
Tools Runbooks Knowledge
│ │ │
└───────────────┼────────────────┘
▼
Cyber Operation
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Findings Evidence Replay
│ │ │
└──────────────┼──────────────┘
▼
Reports
Security Knowledge
CVE
Advisories
Package Versions
Methodologies
Tool Documentation
Vulnerability Intelligence
Software:
ExampleServer 1.2.0
CVE:
Potentially affected↓
Is the vulnerable component actually enabled?↓
Is the vulnerable configuration present?↓
Can the issue be reproduced?↓
Confirmed / Not Applicable
AI Doesn’t Replace the Security Professional
Repetition
Organization
Research
Command assistance
Data interpretation
Documentation
Workflow management
Scope
Risk
Business Impact
Exploitability
Evidence
Authorization
Remediation
Human
+
AI
+
Security Tools
+
Evidence
=
Better Security Workflow
AI = Automatic Hacker
Future of AI-Powered Security
🤖 AI Security Agent
│
┌───────────────┼────────────────┐
▼ ▼ ▼
Reconnaissance Analysis Validation
│ │ │
└───────────────┼────────────────┘
▼
Evidence
│
▼
Findings
│
▼
Remediation
│
▼
Report
Practical Learning Checklist
Local Lab
↓
CTF
↓
Authorized Test Environment
↓
Scoped Bug Bounty
↓
Professional Assessment