Application security for the AI age

Use AI to pentest your code before attackers do.

CyberDuty runs your applications in isolated Docker environments and attacks them with AI test agents. Find and fix security issues before bad actors exploit them. Give your team fast, repeatable security scanning built for the AI age.

⚡ Works in Claude & Codex →
Live CyberDuty Scan — stripe/payment-api
Active Session · Agent Active
$ cyber pentest stripe-payment-api --engine all

→ Inspecting Docker runtime...
✓ Target resolved: http://stripe-payment-api:3000
✓ OWASP ZAP baseline complete
✓ Nuclei template scan complete
→ AI agent validating high-confidence findings...

MEDIUM — Session cookie missing SameSite policy
HIGH — Broken access control detected
GET /api/admin/users

Employee role → HTTP 200
Expected → HTTP 403

RELEASE GATE: BLOCKED

Your annual pentesting is pointless.

Your application changes every week. Your security test happens once a year. Everything shipped in between creates unverified attack surface.

Traditional annual pentest 12 months of engineering changes
Annual pentest
Unverified attack surface
With CyberDuty: test every meaningful release

A dashboard that's a control plane.

CyberDuty turns a fuzzy security scan into release risk, remediation status, clear ownership and continuous monitoring - everything a CTO needs.

CD
CyberDuty / Executive Security Financial API — CyberDuty Control Plane
Live security posture 18 repos protected 2 release blockers
Security posture score
↑ 6 pts this month
82
82/100
Two critical issues still block the release.
Release blockers
2 High-confidence findings
Coverage
92% Repos scanned on meaningful changes
Scans this week
147 PR, main branch, nightly
MTTR
4.6h Down 38% this month
Risk down. Coverage up.
High-risk exposure falls as CyberDuty scans more releases.
92% Continuous coverage
HIGH RISK LOW
Week 1 Week 2 Week 3 Week 4 Now
Every release tested
Open high-risk exposure Release scan coverage
High-risk findings8 → 2
Coverage44% → 92%
MTTR11h → 4.6h
Repo security posture
Know where risk sits across the software portfolio.
Repository
Coverage
Last result
Policy
stripe/payment-api Node / Express · main
Blocked
Release Gate
stripe/customer-portal Next.js · main
Watch
PR Standard
stripe/policy-service Rails API · release
Passed
Nightly Deep
Can we ship?
CyberDuty converts scan evidence into a release decision.
Release blocked

Fix 2 issues before production.

One authorization flaw is actively reproducible. One exposed debug endpoint is reachable in the Docker runtime.

View attack trace
Generate remediation
Create engineering ticket
Top open findings
Executive-ready issues with owner and remediation context.
High
Broken access control on /api/admin/users Employee role receives HTTP 200. Expected HTTP 403.
OwnerBackend
High
Debug endpoint exposes internal config /debug/config reachable in release candidate runtime.
OwnerPlatform
Med
Session cookie missing SameSite policy Observed on login flow and checkout callback.
OwnerWeb
Recent agent activity
What CyberDuty did in the last hour.
AI remediation generated for stripe/payment-api Suggested middleware fix + regression test
4m ago
Release gate failed on PR #284 Broken access control reproduced by AI agent
12m ago
Nuclei flagged exposed debug endpoint Validated against release candidate container
19m ago

Works in Claude & Codex

CyberDuty fits your Claude Code and OpenAI Codex workflow in two ways — you invoke it, or the agent does. Just install the CLI and point Claude to it.

Let Claude scan while you build.


In the terminal

Run cyber pentest directly inside a Claude Code or Codex session. The agent reads the structured output and proposes fixes before you've switched tabs.

As an MCP tool

Register CyberDuty as an MCP server. Claude invokes cyber_pentest automatically during code review — structured results come back, remediations go straight into your editor.

In your CI/CD pipeline

Drop the cyberduty-scan action into GitHub Actions (or any CI). Every PR gets pentested automatically, findings post as a check, and Claude proposes the fix before a reviewer even looks.

CLI Mode
Claude Code Terminal
$cyber pentest payment-api --port 3000
✓ ZAP baseline scan complete
✓ Nuclei template scan complete
CRITICALJWT secret hardcoded in config.js:12
HIGHSQL injection vector in /api/users
Claude ›I see 1 critical and 1 high finding. Fixing config.js:12 first…
MCP Server
Claude is invoking a tool ···
cyber_pentest({
  container: "payment-api",
  port: 3000,
  engines: ["zap","nuclei"]
})
← Tool result · 2 findings
CRITICALconfig.js:12 — hardcoded JWT secret
Claude ›Applying fix to config.js — review the diff in the editor panel →

Ship faster without betting the company on blind spots.

CyberDuty turns application security into part of the engineering workflow — giving CTOs continuous visibility without turning every developer into a security specialist.

Designed for teams that own release risk
Security gates before risky code reaches production.
Executive visibility without waiting for a yearly PDF.
AIRemediation guidance developers can act on immediately.

Catch risk before release

High-severity findings can automatically block unsafe builds before they become customer-facing incidents.

</>

Built for developers

Run locally, in CI, or continuously against connected repositories without forcing developers into a heavyweight security workflow.

EV

Evidence, not AI guesses

Findings include reproducible scanner and runtime evidence, so teams can trust what CyberDuty escalates.

Security that scales with engineering

Every significant change can get its own security review without waiting for consultants, calendar slots, or an annual report cycle.

Stop guessing. Start scanning.

Continuously test your applications with CyberDuty and give your engineering team a clear security decision before production.

Book a CyberDuty demo →