Home / Articles / Practical notes: Numasec | The AI Agent for Cybersecurity

This article is published in English.

Practical notes: Numasec | The AI Agent for Cybersecurity

Operable walkthrough of Practical notes: Numasec | The AI Agent for Cybersecurity : contracts, checks, and drop-in code slots for teams shipping this pattern.

3333 words

The following notes reconstruct a practical path around “ Numasec | The AI Agent for Cybersecurity ”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

What Is Numasec?

The What Is Numasec stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Nmap       → Network discovery
Nuclei     → Vulnerability templates
SQLMap     → SQL injection testing
FFUF       → Content discovery
Nikto      → Web-server checks
Trivy      → Security scanning
🤖 AI Agent
                   │
        ┌──────────┼──────────┐
        ▼          ▼          ▼
      Tools     Runbooks    Knowledge
        │          │          │
        └──────────┼──────────┘
                   ▼
               Operation
                   │
       ┌───────────┼───────────┐
       ▼           ▼           ▼
    Findings    Evidence     Replay
       │           │           │
       └───────────┼───────────┘
                   ▼
                 Report

Why AI Agents Are Interesting for Cybersecurity

The Why AI Agents Are stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Target
Scope
Tools
Observations
Findings
Evidence
Risk
Remediation

Numasec’s Security Workflow

The Numasec s Security Workflow stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Numasec s Security Workflow stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

🎯 Target
   ↓
📋 Scope
   ↓
🧭 Security Posture
   ↓
📖 Runbook
   ↓
🛠️ Local Tools
   ↓
🔎 Observations
   ↓
🚨 Findings
   ↓
📸 Evidence
   ↓
🔁 Replay / Verification
   ↓
📊 Report

Security Reconnaissance With AI

For the Security Reconnaissance With AI stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Domains
Subdomains
Technologies
Ports
Services
APIs
Authentication
Web applications
Cloud services
Raw Tool Output
      ↓
AI Interpretation
      ↓
Structured Observation
      ↓
Potential Finding
      ↓
Evidence

Working With Existing Security Tools

For the Working With Existing Security stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

AI replaces security tools
AI
 │
 ├── Nmap
 ├── FFUF
 ├── Nuclei
 ├── Nikto
 ├── SQLMap
 ├── Trivy
 └── Other authorized tools

Security Agents and Different Modes

For the Security Agents and Different stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Security Agents and Different stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Numasec
                    │
        ┌───────────┼───────────┐
        ▼           ▼           ▼
      AppSec      Pentest     Research
        │           │           │
        ▼           ▼           ▼
       APIs       Network      CVEs
       Web        Systems      Advisories

Runbooks: Turning Security Knowledge Into Workflows

When working through the Runbooks Turning Security Knowledge stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Web Application Assessment
1. Identify target
2. Confirm scope
3. Inspect application
4. Identify technologies
5. Map endpoints
6. Analyze authentication
7. Review APIs
8. Identify potential vulnerabilities
9. Collect evidence
10. Validate findings
11. Generate report

Findings Should Require Evidence

When working through the Findings Should Require Evidence stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Interesting response
        ↓
"Potential vulnerability"
Observation
     ↓
Candidate
     ↓
Verification
     ↓
Evidence
     ↓
Confirmed Finding

Observation ≠ Vulnerability

When working through the Observation Vulnerability stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Observation Vulnerability stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Potential vulnerability
Candidate
Observed
Verified
Rejected
Stale

Evidence Collection

The Evidence Collection stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

HTTP Requests
HTTP Responses
Screenshots
Tool Output
Logs
Hashes
Configuration
Reproduction Steps
Downloads/
Screenshots/
Terminal History/
Notes/
Browser Tabs/
Finding #001
   │
   ├── Observation
   ├── Request
   ├── Response
   ├── Screenshot
   ├── Reproduction
   └── Remediation

Replay and Verification

The Replay and Verification stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Initial Observation
       ↓
Hypothesis
       ↓
Controlled Test
       ↓
Evidence
       ↓
Replay
       ↓
Verified Finding

From Testing to Reporting

The From Testing to Reporting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The From Testing to Reporting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Executive Summary
Technical Findings
Severity
Evidence
Impact
Remediation
References
Terminal
+
Screenshots
+
Notes
+
Browser
+
Scanner Output

Getting Started

For the Getting Started stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

npm install -g numasec
numasec
Your Machine
     │
     ▼
Local Web Application
     │
     ▼
Numasec
     │
     ▼
AppSec Runbook
     │
     ▼
Findings + Evidence

The /doctor Concept

For the The doctor Concept stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Installed tools
Missing tools
Broken tools
Incorrect versions
Unavailable dependencies

Scope Is Critical

For the Scope Is Critical stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Scope Is Critical stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

AUTHORIZED TARGET
       ↓
DEFINED SCOPE
       ↓
ALLOWED TESTS
       ↓
CONTROLLED EXECUTION
Allowed:
example-lab.local
Not allowed:
production.example.com
third-party.example.net
unrelated infrastructure

Security Automation Needs Guardrails

When working through the Security Automation Needs Guardrails stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Scope

When working through the Scope stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Permissions

When working through the Permissions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Permissions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Evidence

The Evidence stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Logging

The Logging stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Human Oversight

The Human Oversight stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Human Oversight stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Safe Environments

For the Safe Environments stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Numasec vs Traditional Security Workflows

For the Numasec vs Traditional Security stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Terminal
   +
Browser
   +
Burp
   +
Nmap
   +
Scanner
   +
Notes
   +
Screenshots
   +
Report
🤖 AI Agent
                      │
        ┌─────────────┼─────────────┐
        ▼             ▼             ▼
     Terminal       Browser       Tools
        │             │             │
        └─────────────┼─────────────┘
                      ▼
                  Operation
                      │
             ┌────────┴────────┐
             ▼                 ▼
         Findings           Evidence
             │                 │
             └────────┬────────┘
                      ▼
                   Report

Why This Approach Is Interesting

For the Why This Approach Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Why This Approach Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Example Authorized Workflow

When working through the Example Authorized Workflow stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

1️⃣ Define scope
       ↓
2️⃣ Start Numasec
       ↓
3️⃣ Check local tools
       ↓
4️⃣ Select AppSec posture
       ↓
5️⃣ Start appropriate runbook
       ↓
6️⃣ Discover application surface
       ↓
7️⃣ Analyze observations
       ↓
8️⃣ Validate interesting behavior
       ↓
9️⃣ Capture evidence
       ↓
🔟 Generate report

Architecture

When working through the Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

👨‍💻 Operator
                       │
                       ▼
                Terminal Console
                       │
                       ▼
                  AI Security
                     Agent
                       │
       ┌───────────────┼────────────────┐
       ▼               ▼                ▼
     Tools          Runbooks         Knowledge
       │               │                │
       └───────────────┼────────────────┘
                       ▼
                  Cyber Operation
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
    Findings        Evidence        Replay
        │              │              │
        └──────────────┼──────────────┘
                       ▼
                    Reports

Security Knowledge

CVE
Advisories
Package Versions
Methodologies
Tool Documentation
Vulnerability Intelligence
Software:
ExampleServer 1.2.0
CVE:
Potentially affected↓
Is the vulnerable component actually enabled?↓
Is the vulnerable configuration present?↓
Can the issue be reproduced?↓
Confirmed / Not Applicable

AI Doesn’t Replace the Security Professional

Repetition
Organization
Research
Command assistance
Data interpretation
Documentation
Workflow management
Scope
Risk
Business Impact
Exploitability
Evidence
Authorization
Remediation
Human
  +
AI
  +
Security Tools
  +
Evidence
  =
Better Security Workflow
AI = Automatic Hacker

Future of AI-Powered Security

🤖 AI Security Agent
                       │
       ┌───────────────┼────────────────┐
       ▼               ▼                ▼
   Reconnaissance    Analysis       Validation
       │               │                │
       └───────────────┼────────────────┘
                       ▼
                   Evidence
                       │
                       ▼
                    Findings
                       │
                       ▼
                  Remediation
                       │
                       ▼
                    Report

Practical Learning Checklist

Local Lab
   ↓
CTF
   ↓
Authorized Test Environment
   ↓
Scoped Bug Bounty
   ↓
Professional Assessment

Final Thoughts

Explore Numasec

Responsible Security Notice

Operational checklist