Blog

AI Penetration Testing that Doesn't Cut Corners: Why the Human Stays in the Loop

Written by UltraViolet Cyber | Sep 8, 2026, 5:32:50 PM

Eight out of ten conversations with customers right now are about speed. Most of them started as a question about testing one AI feature. They've turned into something broader: can testing move at the pace the software actually ships now?

That shift has a cause. AI coding agents and MLOps pipelines have compressed release cycles across the industry, and boards are pushing engineering to ship faster without security slowing it down. A three-week pentest was already tight. A one-week pentest is often too slow now, especially against a codebase a coding agent touches every few days.

At the same time, the software itself changed. Teams aren't only shipping faster, they're shipping different things: chatbots, agentic workflows, RAG pipelines, and now MCP servers. Each one opens a kind of attack surface most security programs haven't tested before, and few have five extra specialists on staff to cover it.

That's the coverage gap. Here's how we're closing it.

Release speed reset the testing clock

For years, pentest timelines ran on a predictable clock: a week or two of testing, a report, a retest window. That worked when the application changed slowly enough for a point-in-time assessment to still be accurate by the time the report landed.

AI-assisted development broke that assumption. When a codebase gets touched every few days by a coding agent, a pentest finished two weeks ago is already describing an application that no longer exists. Security teams are being asked to keep pace with engineering, without adding headcount at the rate engineering is adding output.

Two clocks, one gap Traditional pentest window 21 days AI-assisted release cycle 3 days A report can describe an app that already changed 7 times since testing started.

A second attack surface, same team

Teams shipping faster with AI coding agents are also building AI itself, and that's a separate problem from testing speed. Chatbots, agentic flows, retrieval-augmented generation, and MCP servers that let agents reach into other systems all introduce attack paths that didn't exist two years ago: prompt injection, data leakage through a retrieval pipeline, an agent operating outside its intended scope, an MCP server handing out access it shouldn't.

Most security teams don't have deep AI security expertise on staff already. They have the same headcount and a growing list of things that need testing for the first time. That combination, speed plus unfamiliar territory, is the coverage gap.

Why automation alone doesn't close it

Two options have existed for years. Off-the-shelf scanning tools are fast and shallow: they send a payload, read a response, flag a pattern, with no sense of what the application does for the business. Expert-led manual testing goes deeper, but it can't keep pace with how fast AI-assisted development ships code.

A newer option has entered the market: fully autonomous, agentic pentest platforms. Many are genuinely capable. They're also the reason security leaders ask the same question first: how do you trust an agent to test production systems unsupervised? What happens when it goes further than it should, misses something a person would have caught, or reports a finding that isn't real?

That's why every Solstice engagement runs with a person in the loop from start to finish.

Four ways to test Scanners Fast Shallow Manual pentest Deep Doesn't scale Autonomous agentic Fast Hard to trust Human-in-the-loop AI Fast Verified Where Solstice sits

What human-in-the-loop actually means inside Solstice

Solstice is the platform our own practitioners use to test faster. Every engagement starts with context (kickoff notes, architecture docs, whatever the client provides), which Solstice uses to draft a threat model and test plan. A consultant reviews that plan against what they actually see in the application before any testing starts.

Scope and authorization run as two separate checks. Every attack probe is checked against declared scope before it's ever sent, so nothing reaches a system that isn't in bounds. Separately, specific attack techniques only run once a consultant has confirmed they're authorized for that engagement. Solstice runs on a knowledge base built from years of our own engagements, which cuts down on the kind of confident, invented findings that make autonomous tools hard to trust in the first place. Every finding also comes with HTTP evidence and screenshots, so a consultant can check that what Solstice reports actually happened, before it goes in a report.

AI proposes. Practitioners decide.

How a Solstice engagement runs 1. Context in: kickoff notes, architecture docs 2. Threat model and test plan drafted 3. Consultant reviews the plan before testing starts 4. Scope check + authorization gate, per probe 5. Testing runs; findings logged with HTTP evidence 6. Consultant validates every finding before it's reported

What the numbers show

We ran this side by side with our own testers: the same assessments, once with a person testing alone and once with a person working with Solstice, across more than 50 assessments and dozens of customers.

Coverage held up. Solstice-assisted testing caught 65 to 100 percent of what a solo human tester found, depending on the engagement, plus findings the human missed, roughly 50 percent more overall. Speed improved too: about 20 percent faster on complex, first-time engagements. On repeat engagements, where Solstice already carries context on the application, a test that used to take seven to ten days now runs in two to three.

The person still validates every finding. Automation just gives them more ground to cover in less time.

65-100% of solo-human findings also caught with Solstice +50% more findings surfaced across the same engagements 20% faster, first-time engagement 2-3 days on repeat (was 7-10)

Where the model still needs a person

Working with an LLM on a security test shows two specific failure modes. First, it gets close to a real vulnerability and stops one step short, having decided it exhausted the obvious paths. A consultant who remembers a detail from twenty turns back can point it at the right target again. Second, it reports something as a critical finding with total confidence, then retracts it the moment someone reminds it of context it had already forgotten.

Both failure modes are why the automation ships with a person attached to it.

Testing AI with AI

The same approach now runs against the AI systems clients are shipping: chatbots, agentic flows, RAG pipelines, and the first wave of MCP server deployments. Testing a chatbot used to mean trying prompts one at a time by hand, which is slow and doesn't scale. An AI-driven approach can run hundreds of prompts in the time a person runs five.

Testing for generic jailbreaks (getting a model to describe something dangerous) checks a risk that belongs to the model provider. Testing at the application level checks something specific to the business: can this chatbot be pushed to expose account data it shouldn't reach, or call a backend tool outside its intended scope? That takes real research into how the specific application and its system prompt work, at a depth and speed a generic scanner can't match.

Where this leaves security teams

Nobody has fully settled how the pentester's role changes over the next few years, including us. A few things are already clear: testing needs to move at the speed of AI-assisted development, coverage has to extend to the AI systems teams are shipping now, and none of it works without someone accountable for what gets tested, reported, and fixed.

If your last pentest took three weeks and your engineering team ships changes in three days, that gap is worth a conversation.

How to engage

UltraViolet Cyber runs this three ways, so testing can flex with how a program actually works. À la carte for a single target with a fixed scope and window. A 3D Security Testing Subscription for teams that need to flex what gets tested, when, and how deep, without renegotiating a contract every time priorities shift. A Virtual Security Team, billed by effort instead of by engagement, for teams that want a dedicated extension of their own. All three run on the same practitioner-led, Solstice-assisted approach.

Watch the full conversation, including the Q&A on jailbreak controls, the future of the pentester's role, and how AI is starting to test AI, on demand: watch it here.