At the beginning of 2026 all the talk was around AI pentesting. How good it was, how bad it was, and the impending doom for security jobs in the industry. Every week there was a new account claiming they had the best AI pentesting setup and that they were in a position to replace human security personnel.

With all of the chaos around the topic on social media, I felt like most other security professionals. Very confused, not knowing what to trust and what to disregard as AI slop propaganda. Even worse, if you are the person in your organisation who has to decide whether one of these platforms is worth the budget, that confusion gets expensive fast.

So I did what I always do when I am confronted with an unfamiliar topic. I jumped in head first. Over the past year I worked as part of two separate teams building, benchmarking and running industry-backed AI pentesting solutions. For one of them, namely NoScope, I was heavily involved in the benchmarking, testing, and in some cases consulting on technical improvements of the solution itself.

Working on these solutions over the last year, I spent time developing a business-orientated framework. This framework is for industry leaders both technical and non-technical, who are in the same boat I was in a few months ago. I want to take you through this framework using actual war stories from being actively involved in the development of these solutions and in the end, help you develop questions you should be asking before you sign up for the next best AI product.

Let’s demystify the AI slop propaganda.

The platform behind most of the war stories below is NoScope where MWR was an early access partner. I have been part of the team building and benchmarking it. You should read everything here with that relationship in mind.

The MWR Approach

Before we get into the case studies, let’s first look at how MWR approached these systems.

Much like myself, MWR prefers a hands-on approach to unknown concepts, so we set out to understand what these AI solutions had to offer in a practical sense. Utilising an AI pentesting solution not named in this blog post, we ran a dry run against our own environment, with all the necessary IP allow-listing in place. We pointed it at our environment and watched it do its thing.

At the end of the engagement, seven vulnerabilities were reported. They ranged from informational risk to a non-exploitable medium risk finding. There was however something more interesting than the findings themselves.

AI pentesting solutions generally assign an internal ID to each candidate finding as they go. While our final report had seven vulnerabilities, the highest finding ID that reached us was 21.

We don’t have visibility into the internal triage process of this product, so while we cannot state this with certainty, we did wonder what happened with the other 14. Our guess is that some were duplicated, and some may have been out of scope or informational. Though we do wonder how many were potentially false positives discarded during verification. What we can say is that roughly two thirds of what the agent generated did not survive to the report and that somebody (or something) had to do that filtering.

That is the whole argument of this post. We never had a concern that the technology doesn’t work or doesn’t find really interesting stuff, but rather that filtering is still real work that needs to happen. This filtering is also not free. Those reading this post who are clients of MWR will be able to attest to how many of our findings are false positives. This is something to keep in mind when reviewing these solutions as well, to make sure we are comparing apples with apples.

The Makings of a Framework

Like most security professionals, I enjoy a good old methodology or mental model to help me approach a specific scenario. Same as some of my previous research endeavours, I set out to construct another model, but this time with a bit of a twist.

The framework should not only be usable by technical security personnel, because technical personnel are often not the ones making the decision on what new tools will be bought. It should be usable by the non-technical people who care far more about the business case for these solutions than the mechanics of them. As a security consultant, where the business aspect of security is always somewhere in the back of my mind when performing a test, that felt like a fairly natural way to approach it.

So how do we formulate a way to think about where these solutions play a role in a business? We start with a capability matrix. We need to understand how these solutions perform and their current shortcomings on a fundamental level.

The framework is built on two conditional categories that tie directly into how pentest engagements actually get scoped:

  1. Technical Complexity
  2. Environmental Scope

Technical Complexity

This category forces us to take a step back and really understand the technical nuance associated with our security need. It breaks down into two sub-categories:

  • Technology and context complexity: How common is the technology stack that we are testing? How critical is a business understanding of the solution to the test?
  • Deliverable requirements: Do we require specific test cases that dive deep, or do we simply want security coverage?

Take the following two examples:

Annual web application security assessment of an ASP.NET application served over Internet Information Services (IIS).

Technology and context: Standard
Deliverables: Security coverage

Versus:

User Acceptance Testing (UAT) security assessment of an internal gRPC-based API using mutual TLS (mTLS) authentication, serving a critical business process.

Technology and context: Edge case internal gRPC API serving a critical business process
Deliverables: Complex test cases required

The first is a well-trodden stack where we want confidence that the common issues are not present. The second needs someone who understands both the protocol and what the business does with it in order to ensure our testing efforts go deep enough.

Environmental Scope

This category is critical and often the final decider for how effective an AI pentesting solution will be. We need to understand the scope of what we are asking for with regards to the security assessment. Are we working with a single API endpoint, or are we expecting the pentester, human or otherwise, to test the security boundaries between multiple systems?

As we will see in the case studies, this factor is the one most often overlooked and also the one AI pentesting solutions struggle with most today.

Derived Categories

Once we have plotted those two conditional categories, the framework gives us back two answers:

  • Approach: How should we be running this test? Do we let the agent run unattended, do we direct it, or do we lead and use it for the grunt work?
  • Lead: Who owns the greater engagement? Who oversees it and how much human involvement does that realistically require?

These are the payload of the whole exercise, so it is worth being concrete about what they look like in practice.

Take our first example above, the annual ASP.NET assessment. Standard technology, security coverage deliverable, a single application in scope. The Approach would be to automate it. Let the agent run its full test case library unattended. The Lead would be the agent performing the testing itself, with a human reviewing the output before it goes anywhere near a report. The human cost here is measured in hours of review instead of days of testing.

Now take the second, the internal gRPC API over mTLS, serving a critical business process, sitting in an environment where it talks to three other internal services. Edge case technology, complex test cases required, and the interesting attack surface is the boundary between systems. The Approach would be human-led, with the agent used to cover the repetitive test cases across each service so the tester can concentrate on the boundaries. The Lead would also be the human, unambiguously. The agent is a force multiplier here and not a substitute for years of experience in finding and testing niche edge cases.

In the two examples mentioned, it can be the same human and the same AI pentesting tool, but it leads to two completely different engagements, different cost models, and different expectations to set with a client.

The Output

Put it all together and we get a matrix that looks something like this:

Four-quadrant capability matrix. Vertical axis: what kind of context, from technical to high business context. Horizontal axis: what kind of vulnerability, from known to novel. Quadrants: Automate, Verify, Direct and Lead.

Figure 1: The capability matrix.

Every one of the four quadrants has a different answer to the same simple question:

Will an AI pentest actually add value here?

War Stories

We could spend hours on the technical and business nuances on paper and I could waste your time promising that this framework works, which is boring. So let’s look at four real war stories and map them to the various quadrants to see the framework in practice.

Do note: affected parties to which these case studies are subject have either been notified prior to public release or the vulnerability disclosure embargo has been lifted as per an agreed timeline.

Commonly Identifiable Vulnerabilities

The capability matrix with the bottom-left quadrant, Automate, highlighted in orange.

Figure 2: Known vulnerability classes in a technical context.

In this quadrant, our technical and contextual complexity almost does not exist. You can almost think about this as the vulnerability scanner world of ten years ago. It is a solve space and if you are still doing everything here by hand, you are vastly out of your depth. We are not required to rabbit hole into complex systems to find security issues in the fundamental implementation of a solution. We simply want security coverage, in line with your favourite common vulnerability classification list, like the OWASP Top 10 being the usual reference point for web applications.

Approach: Automate. Let the agent run its full test case library unattended.
Lead: Agent-led. The human reviews and verifies the output.

So what do we do when the matrix drops us in this quadrant? We let the AI pentester loose and let it run.

Oh Scanner My Scanner

Let’s break this war story into three phases:

  1. I pointed NoScope at Gitea, an open-source Git solution. The agent ran for just under three hours and claimed security coverage in line with the OWASP Top 10.
  2. Looking at the test cases the agent chose to run, they were clearly simple in nature, with an odd amount of focus on authorisation and authentication flaws. It should be noted that the agent was intentionally instructed to perform baseline security tests and if you wanted, you can go a lot deeper into a solution like Gitea.
  3. A critical risk finding was raised, initially scored at CVSS 9.2 and later de-risked.

So over those three hours, the agent performed very simple test cases against an instance of Gitea and raised a critical risk finding. The finding itself was straightforward:

  1. Request an anonymous JSON Web Token (JWT). A short-lived, anonymous token the container registry provides so that public images can be pulled.
  2. Append a scope=* parameter to that request and the token comes back with administrator-level authorisation attached to it.
  3. Use that over-permissive scope to enumerate and pull private containers even though the initial idea was to limit the scope to only public containers.

For anyone who has not had to think about container registries, let’s do a quick refresher. Private containers routinely hold secrets, credentials and intellectual property. So gaining access to these is definitely not a “this vulnerability in theory could allow an attacker to theoretically…” type of situation.

Reviewing the Gitea source code, we determined the vulnerability had been present for four years without being detected, affecting more than 31,750 production instances in the wild, along with multiple downstream forks.

A news post reporting that a Gitea flaw exposes private container images without authentication, affecting all versions before 1.26.2.

Figure 3: The finding after disclosure, once it had been picked up in the press.

The agent found it because of one simple fact that this quadrant highlights very well:

AI agents don’t get tired. Set up properly, they will attempt every test case for every vulnerability class against every endpoint and they will not stop until they are done.

That tedious process would take a human days or weeks even with conventional automation, depending of course on how many coffee breaks are taken. Agents are good at repetitive, tedious work, and they do it well.

Application-Specific Vulnerabilities

The capability matrix with the bottom-right quadrant, Verify, highlighted in orange.

Figure 4: Novel vulnerabilities in a technical context.

What if there is no real critical business context to understand in order to test the application effectively, but we want to dig deep into the technical parts of it to look for novel findings that break the way something is implemented on a fundamental level?

Approach: Agent does the grunt work, human verifies the reasoning.
Lead: Shared. The agent surfaces the attack surface and the human closes the loop.

Much like the first quadrant, we can stay fairly hands-off on the tedious work, letting the agent grind through the surface area and freeing us up to focus on the genuinely technical parts. The difference is what we do with the output. In this quadrant, we need to do a decent bit more verification than simply believing the findings at face value.

The Assumption

We start this war story like the last one. I let NoScope loose on an application with no restrictions on the types of vulnerabilities we were looking for. The agent went about its day and started digging into how the application worked on a fundamental level.

An extension sandbox, more specifically one that allowed event managers to run custom workflow scripts on Alf.io, was the target for this run. The sandbox ran JavaScript but exposed a limited set of Java classes that could be invoked server side from within those scripts. That interop boundary, between the script and the Java classes behind it, is where this story comes to life.

The agent then immediately did something that Large Language Models (LLMs) are quite famous for. It made an assumption (ahem hallucinated) about the technology used to implement the sandbox. This was probably fair because the agent was running from a blackbox perspective and engine signatures could only be objectively confirmed with source code access. For those unfamiliar, there are several JavaScript engines that support this kind of Java interop, namely Nashorn, GraalJS and Mozilla Rhino, and their implementations differ meaningfully.

Working from the outside, without the source-level visibility, it found a plausible way past the script validation and, on the surface, the logic and proof-of-existence looked clean. This was generally enough to convince a technical audience familiar with the sandboxing technologies, but the exploit primitive did not show the impact required since it wasn’t yet fully exploitable. Now that the agent got this far and there was a real gap in the validation logic, I went to the source myself, past what the agent could see from the outside, and found the actual binding the sandbox had left exposed. From there, proving full impact was straightforward.

Execution output showing a success message and the command result uid=1001(alfio) gid=101(alfio).

Figure 5: Command execution on the host, confirming the sandbox escape had reached full remote code execution.

That finding became CVE-2026-35482. The full technical write-up, including the sandbox internals and the defensive takeaways, is on the NoScope blog.

This war story is less about a shortcoming and more about where the line currently sits between the human and agent. An agent working blackbox, reasoning soundly from what it could see, and a human closing the gap that only source-level access could close. That line keeps moving with every frontier model release, but for now, novel technical work, the kind that sits at the edge of what an LLM has seen before, is exactly where the second set of eyes earns its place. In this case, that’s what turned just a sandbox bypass into full RCE.

Common Vulnerabilities in Context Rich Environments

The capability matrix with the top-left quadrant, Direct, highlighted in orange.

Figure 6: Known vulnerability classes in a high business context.

Now we move into the scenarios where the application is wrapped in a lot of business and technology-specific context, all of which needs to be understood to perform an effective security assessment.

Approach: Direct the agent. Break the environment into targeted runs, one prioritised area at a time.
Lead: Human-led. The human decides where the agent points and the agent does the depth.

This is where we as humans start to play a much bigger role, because we have the ability to ingest and understand context that the agent cannot. It becomes our responsibility to structure the test and guide the agent down the right path. Give it a generic prompt in an environment like this and you will get generic results that are best suited for the bottom-left quadrant, same as your classic old vulnerability scanners.

Four Prompts, Four Runs, One Critical

This war story shows that approach adjustment clearly. The target was a cryptocurrency trading and transacting platform with a great deal of technical and business nuance. The application was massive in scope and required a lot of cross-context to really understand what was happening. Hence the following approach was applied:

  1. The first run was broad scope. I simply let NoScope loose against the application and unsurprisingly, it came back with nothing notable.
  2. The second run focused on how the application handles crypto deposits and withdrawals. Again, nothing really notable.
  3. The third run focused on the peer-to-peer transaction functionality. Once again, nothing notable.
  4. For the fourth run we directed the agent at the real-time market data ingestion functionality, which ran over a Message Queuing Telemetry Transport (MQTT) websocket connection. That is where we struck gold.

An unauthenticated MQTT publish vulnerability, which allowed any user, regardless of authentication state, to publish new market data updates to the MQTT queue. Those updates were then ingested by every single user on the platform.

The impact we proved was straightforward: we set the price of Bitcoin to 1 USD and watched the entirety of the QA environment’s market fill orders trigger, losing a lot of fictitious users a lot of fictitious money.

Platform price ticker showing Bitcoin at one US dollar, down 99.65 percent in the past 24 hours.

Figure 7: Bitcoin priced at 1 USD in the QA environment, after an unauthenticated publish to the market data queue.

The finding was ultimately rated CVSS 9.7 Critical and was patched before going to production. But three runs produced nothing notable, and the fourth produced a critical. The only variable that changed was a human deciding where to point it. The easy answer is to say that this is a shortcoming of the agent, but I think it is more intricate that this. It shows how these engagements can become somewhat of a symbiotic relationship between directing human and executing agent.

Novel Vulnerabilities in Context Rich Environments

The capability matrix with the top-right quadrant, Lead, highlighted in orange.

Figure 8: Novel vulnerabilities in a high business context.

For our last quadrant, we are in a business context rich environment where we need to understand the environment as a whole, and we are scoped to find vulnerabilities that exist in that very business context.

Approach: Human-led throughout. Use the agent for coverage inside individual applications but never for the boundaries between them.
Lead: The human. Do not outsource this quadrant.

This is where AI pentesting solutions still fall short and where humans still fundamentally dominate the pentesting scene. Not because agents cannot take in the information, but because they cannot yet hold a large amount of context and apply an intuitive mindset across all of it to find the holes.

Cross-App Privilege Escalation

This engagement came in as, paraphrasing the client’s security lead: “we have a three-app public facing e-commerce environment that we need pentested”.

To lay the land a bit, we start by authenticating to a custom Identity Provider (IDP) in the environment:

Attack flow step one. A user authenticates against a single identity provider.

Figure 9: Step one. The user authenticates against the environment’s single identity provider.

Which gives us back a JSON Web Token (JWT) that we can use to perform authenticated actions against the main e-commerce application:

Attack flow step two. The identity provider returns a JWT to the user.

Figure 10: Step two. The identity provider returns a JWT.

For the rest of this war story we need to understand one critical thing about JWTs. JWTs carry a set of standard claims, small pieces of information inside the token, that most authentication and authorisation libraries know to expect. Because they are standardised, they are effectively common knowledge and show up in any documentation you look at. The iss, sid and aud claims are examples.

The one that matters here is aud, the audience claim, which tells the receiving library which application this token is valid for. That is the standard behaviour, unless a developer deliberately extends the ingestion logic. The token our IDP handed back had a valid aud claim, but it also had a custom sid-scope claim, which is not part of the standard and not respected by any library out of the box.

Attack flow step three. The decoded JWT payload shows an aud claim and a custom sid-scope claim set to a wildcard, and the request goes to its intended target, the customer storefront.

Figure 11: Step three. The decoded token carries a standard aud claim alongside a custom sid-scope claim, and the request reaches its intended target.

Reviewing the agent’s reasoning trace, we could see that it noticed both claims. It noted that sid-scope was custom, performed its own research into JWT ingestion libraries, saw that they commonly use the aud claim to determine the intended audience, confirmed aud was present and correct, and moved on.

That was the blunder. It assumed the applications in this environment actually used the aud claim to determine scope, when, as it had itself just researched, libraries can be extended to respect whatever claim the developers choose.

When we spotted this, the test case was almost embarrassingly simple. We authenticated, took the JWT with its aud and sid-scope claims, and pointed it at the employee back office application instead of the storefront. The back office was more than happy to let us in. And once inside, the dashboard had significant authorisation issues of its own, exposing e-commerce data we should never have been able to see.

Attack flow step four. The same JWT is replayed against the back office console, which trusts the wildcard sid-scope claim and accepts it.

Figure 12: Step four. The same token, replayed against the back office console, which gates on sid-scope and accepts the wildcard.

The detail that stays with me is that the agent had every single piece of information it needed to run this test case. It found the claims, it researched the libraries, it identified the back office during recon. It just never connected the three. That is the current biggest shortcoming of security agents. They cannot reliably and consistently link large amounts of context together to produce a novel idea, because that is not how they work on a fundamental level. Ingesting and acting on large sets of security context is still a human job.

Four Simple Questions

To tie it all together, here are the questions worth asking before you commit budget to one of these new AI pentesting platforms:

  1. Which quadrant does our business need sit in? Be honest about the scope. Most organisations think they are buying quadrant one and are actually buying quadrant four. Even worse, the surface level marketing of these platforms can often blur the lines of the true quadrant they are operating in.
  2. Which quadrant does the solution claim to operate in? Most vendors will tell you they do all four. They do not. Ask for demos and PoCs and run tests cases for all the quadrants you might have a business need for.
  3. Does the solution fit a human-led workflow or does it claim to replace one? Three of the four quadrants require a human in the loop. A platform designed around the assumption that nobody is checking its work is a platform designed for one quadrant only. Trying to fit the work of other quadrants in such a solution will lead to the exact same problem you are currently facing by having thousands upon thousands of vulnerabilities and simply no way of determining where you should prioritise your remediation efforts.
  4. Does it deliver results you could put in front of a board? Proof of concepts with real business impact and a remediation report that accounts for the technology actually in use. Or does it need three technical employees to decipher and translate its output before anyone can act on it? The biggest claim of AI is an increase in speed and efficiency. This should extend for how results are presented to the various different audiences as well.

There is a South African wrinkle worth adding to question one. These platforms are priced internationally and paid for in rand. Even worse often against security budgets that are usually tighter than those of the organisations the pricing was designed for. That means the realistic local choice is often not “platform and pentest” but “platform instead of pentest”. If your business need sits in quadrant three or four, a multi-application environment, or one where the interesting risk lives in your business logic rather than your framework, that trade puts your coverage exactly where these tools are weakest. Getting question one right is critical. It is the difference between a spend that extends your team and one that replaces the part of the test that was finding the real problems.

Closing

It is worth noting that the AI pentesting scene, and AI as a whole, is moving extremely fast. A number of the shortcomings in this post may not hold a year from now. But the framework is not a bet on the technology staying still. The framework was created as a way of deciding, at any given point in time, which parts of your testing an agent can carry today. Each time you have a new test case or there is a new solution being evaluated, you will need to go through the exercise of answering these questions. Also take into consideration that what falls in quadrant two today, might be quadrant one tomorrow.

I will close with something the creator of Burp Suite, Dafydd Stuttard, said, which sums it up better than I can: “This isn’t a revolution that eliminates pentesters, it’s an evolution that empowers them”. Used properly, it is an evolution that lets us work smarter, better, faster, and with much greater precision.

If you are evaluating one of these platforms and want a second opinion on which quadrant you are actually operating in, feel free to reach out.