AI Pentesting

PCI DSS Penetration Testing Requirements: Manual vs Automated, and the PR-to-Retest Workflow

Amartya | CodeAnt AI Code Review Platform
Sonali Sood

Founding GTM, CodeAnt AI

PCI DSS v4.0.1 Requirement 11.4 requires internal and external penetration testing at least once every 12 months and after any significant change, segmentation testing on the same cycle (every six months for service providers), and retesting until every exploitable finding is fixed. All of it must follow a documented methodology your QSA can review.

The requirement is clear. Meeting it is harder when your team ships code every week and the annual pentest report describes an application that no longer exists.

This guide maps every part of Requirement 11.4 to what you actually have to do, compares manual and automated testing against it, and walks through the PR-to-retest workflow that keeps you compliant between annual engagements.

Where this connects: PCI is one of three frameworks that expect a pentest. For the SOC 2 and HIPAA side, see Compliance Penetration Testing: SOC 2, PCI DSS, HIPAA. This guide is the deep PCI dive.

What PCI DSS 4.0.1 Requirement 11.4 Actually Mandates

Start with the version note that trips teams up. In PCI DSS v3.2.1, penetration testing lived in Requirement 11.3.

In v4.0 and the current v4.0.1, it moved to Requirement 11.4, with more explicit methodology, clearer segmentation rules and mandatory retesting. v4.0.1 is the only active version, and every future-dated v4.0 requirement became mandatory on March 31, 2025.

Topic

PCI DSS v3.2.1

PCI DSS v4.0.1

Penetration testing methodology

11.3

11.4.1

Internal penetration testing

11.3.2

11.4.2

External penetration testing

11.3.1

11.4.3

Correcting and retesting findings

11.3.3

11.4.4

Segmentation testing

11.3.4

11.4.5

Service provider segmentation testing

11.3.4.1

11.4.6

Multi-tenant provider support for customer testing

Not present

11.4.7

Vulnerability scanning

11.2

11.3

If your policies still cite 11.3 for pentesting, update them before the assessment. Stale numbering signals a program nobody has reviewed.

Requirement 11.4 breaks into seven sub-requirements:

  • 11.4.1, documented methodology. A defined, documented and implemented methodology. The full list of what it must contain is below, because this is where assessments fall apart.

  • 11.4.2, internal penetration testing. At least once every 12 months and after any significant infrastructure or application upgrade or change.

  • 11.4.3, external penetration testing. Same frequency and triggers as internal, covering the CDE perimeter and internet-facing systems.

  • 11.4.4, correct and retest. Exploitable vulnerabilities are corrected according to the risk ranking you assign under Requirement 6.3.1, and testing is repeated to verify the fix. The PR-to-retest workflow exists to satisfy this one.

  • 11.4.5, segmentation testing. Where segmentation isolates the CDE, controls are tested at least once every 12 months and after any change to segmentation controls or methods.

  • 11.4.6, segmentation testing for service providers. Service providers test segmentation at least once every six months and after changes.

  • 11.4.7, multi-tenant service providers. Hosting and payment platforms that serve many customers must support their customers' external penetration testing. It became mandatory on March 31, 2025.

What the 11.4.1 Methodology Must Include

The methodology requirement is longer than it looks at first read. Under v4.0.1, it must include:

  • Industry-accepted approaches, such as NIST SP 800-115, the OWASP Web Security Testing Guide or PTES.

  • Coverage of the entire CDE perimeter and critical systems.

  • Testing from both inside and outside the network.

  • Testing to validate segmentation and any scope-reduction controls.

  • Application-layer testing that covers, at a minimum, the attack types listed in Requirement 6.2.4.

  • Network-layer testing of all components that support network functions, plus operating systems.

  • Review of threats and vulnerabilities experienced in the last 12 months.

  • A documented approach to assessing and addressing the risk from exploitable vulnerabilities found.

  • Retention of results from testing and remediation for at least 12 months.

The 6.2.4 reference carries more weight than it looks. It pulls in injection attacks, attacks on data structures, attacks on cryptography, business logic abuse (including API manipulation, XSS and CSRF), attacks on access control mechanisms, and attacks via any high-risk vulnerability you've identified.

Access control is where scanners go quiet. Broken authorization and IDOR vulnerabilities only show up when someone actually tries to reach another user's data. The 12-month threat review is the clause teams skip. Name the threats you've seen, such as a credential-stuffing wave or a skimming attempt, and show how the test targeted them.

Who Can Perform the Test

Tests can be run by a qualified internal resource or a qualified external third party. The tester must be organizationally independent of the systems under test, and doesn't need to be a QSA or an Approved Scanning Vendor (ASV).

In practice, independence means the engineer who built or manages the payment service can't be the one testing it. How to hire a penetration tester covers what qualifications to check.

What QSAs Expect to See

Your assessor looks for five artifacts:

  1. Methodology documentation following a recognized framework and mapped to 11.4.1. The PCI SSC's Penetration Testing Guidance is still the Council's reference on methodology and reporting, even though it predates v4.0 and uses the old 11.3 numbering.

  2. Scope definition with IP ranges, application URLs, authentication boundaries and anything excluded, with the reason.

  3. Exploitability proof showing working attacks and the path to cardholder data, with tester identity and independence documented.

  4. Segmentation results showing whether each out-of-scope segment could reach the CDE.

  5. Retest validation proving the fixes work, per 11.4.4, with test dates that line up with your 12-month cycle and change log.

Keep the report, remediation tickets and retest evidence together for at least 12 months. That's the retention period 11.4.1 sets.

The Significant Change Compliance Gap

11.4.2 and 11.4.3 both say "at least once every 12 months and after any significant change." Annual testing alone only satisfies the first half.

Picture a team that deploys weekly. By the next annual test, dozens of releases have shipped untested, and the QSA will ask how you know none of those changes introduced an exploitable vulnerability.

"We'll catch it next year" doesn't answer it. The standard leaves "significant" for you to define, so write a definition down, log every change that met it, and show a test for each. Closing that gap is the point of the PR-to-retest workflow below.

PCI Penetration Test vs ASV Scan vs Internal Vulnerability Scan

PCI DSS requires all three, and none substitutes for another. A Qualys or Tenable scan alone can't satisfy 11.4.

Attribute

ASV external scan (11.3.2)

Internal vulnerability scan (11.3.1)

Penetration test (11.4)

Frequency

Every three months and after significant changes

Every three months and after significant changes

Every 12 months and after significant changes

Who performs it

A PCI SSC Approved Scanning Vendor

Qualified personnel

Qualified, organizationally independent tester

What it does

Identifies known external vulnerabilities

Identifies known internal vulnerabilities, authenticated under 11.3.1.2

Attempts exploitation and chains weaknesses toward cardholder data

Tests business logic and access control

No

No

Yes, per 6.2.4

Evidence produced

ASV report with pass or fail

Scan results and remediation records

Report, remediation evidence and retest results

A scan lists what might be wrong. A pentest proves what an attacker can actually do with it, and your QSA will want evidence of real exploitation attempts.

This is the vulnerability scanning versus testing distinction written into the standard. The pressure behind it is in the breach data. Verizon's 2026 Data Breach Investigations Report found that exploitation of vulnerabilities was the most common way attackers got in, at 31% of breaches, overtaking stolen credentials for the first time in the report's history.

The Three Testing Approaches: Manual, Automated, and Hybrid

Teams usually frame this as a binary choice between a manual pentest that takes weeks to schedule and a scanner that floods the backlog with theoretical findings. There's a third option, and it's the one built for 11.4.4.

Manual penetration testing: deep, but it does not scale

Human testers are irreplaceable for certain problems. Business logic abuse needs an understanding of intent, such as a discount code validated only on the client that allows a 100% discount.

Multi-step attack chains that combine low-severity findings into a critical exploit also need a human. So do complex authentication edge cases, like OAuth misbinding in a multi-tenant system.

The trade-off is that manual testing can't keep pace with continuous delivery.

Constraint

Impact on PCI 11.4

Scheduling lead time

Engagements are booked weeks ahead, so change-driven testing lags the change

Per-engagement pricing

Each test is a separate engagement, and a retest often needs a new statement of work

Sequential execution

One tester works through scope linearly, with no parallelism across services

External-only perspective

Black box testing misses code-level context, so inside-knowledge bugs go undiscovered

The retest bottleneck is what breaks the model against 11.4.4. A finding is discovered, the fix ships, and then the confirming retest waits for another engagement slot, which a weekly release cycle can't absorb.

Automated penetration testing: fast, but watch the false positives

Automation covers the breadth well: asset discovery (subdomain enumeration, service fingerprinting), known-CVE detection, configuration drift (exposed S3 buckets, hardcoded keys), and continuous monitoring with no scheduling. The automated pentesting guide covers how it runs.

Where traditional automation fails is the exploitability gap. Scanners report vulnerabilities from version detection rather than actual exploitation, so you waste hours triaging findings that are not exploitable in your environment.

They report issues in isolation rather than constructing the attack chain that shows how reconnaissance leads to a credential, which enables a Broken Object Level Authorization flaw, which pivots to CDE access. And they struggle with authenticated coverage, the privilege boundaries and authorization bypasses where the real risk lives.

AI-driven exploitation changes this by validating exploitability through actual exploitation, constructing a working PoC and operating on a "no working exploit, no payment" model. When a report includes a curl command that actually returns cardholder data, there is no ambiguity for the QSA to resolve.

The Hybrid PR-to-Retest Workflow Keeps You Continuously Compliant

Code-aware gray box testing combines defensive code review with offensive penetration testing. Code review catches insecure patterns before merge, production testing validates whether the deployed controls actually work, unlimited retesting confirms fixes immediately without scheduling, and code intelligence guides the attack paths, analyzing authentication middleware to find excluded routes and tracing data flows from input to dangerous sink. This is the approach that keeps you compliant with 11.4.4 between annual engagements, and we cover the cadence case in full in Continuous vs Annual Penetration Testing.

Scope Failures That Break PCI Pentests

Most PCI penetration testing failures happen during scoping. Teams underestimate what is in scope, miss connected systems, and discover mid-audit that their "isolated" payment service shares infrastructure with a dozen other applications. Scope has three layers.

  • The Cardholder Data Environment is any system that stores, processes, or transmits cardholder data, but teams routinely miss the admin panel that queries payment records, the reporting dashboard that aggregates transactions, and the CI/CD runner holding production secrets.

  • The CDE perimeter is anything that can directly interact with the CDE: API gateways, authentication services, load balancers, service-mesh control planes.

  • And the third layer, systems that could impact CDE security, is where scope explodes: shared Kubernetes clusters, centralized logging that ingests CDE logs, secrets managers, developer workstations with production access, and third-party SaaS integrations.

Code-aware discovery maps this far more accurately than an outdated network diagram, because it reads the actual data flows:

@app.route('/api/checkout', methods=['POST'])
def checkout():
    card_data = request.json['payment_details']  # entry point
    token = tokenize_card(card_data)              # CDE system call
    charge = process_payment(token)               # payment processor
    log_transaction(charge)                       # logging system now in scope
    return jsonify(charge)
@app.route('/api/checkout', methods=['POST'])
def checkout():
    card_data = request.json['payment_details']  # entry point
    token = tokenize_card(card_data)              # CDE system call
    charge = process_payment(token)               # payment processor
    log_transaction(charge)                       # logging system now in scope
    return jsonify(charge)
@app.route('/api/checkout', methods=['POST'])
def checkout():
    card_data = request.json['payment_details']  # entry point
    token = tokenize_card(card_data)              # CDE system call
    charge = process_payment(token)               # payment processor
    log_transaction(charge)                       # logging system now in scope
    return jsonify(charge)

That log_transaction call means your observability platform receives payment data and is in scope. External-only testing misses it, because the log endpoint isn't publicly exposed.

Accurate scoping is how you know your real attack surface. The internal penetration testing and external penetration testing guides map the two sides that 11.4.2 and 11.4.3 require. If your CDE runs in the cloud, provider rules also apply. Cloud penetration testing for AWS, Azure and GCP covers what each allows.

Why Code-Aware Gray Box Testing Finds What Black Box Misses

External-only testing systematically misses vulnerabilities that require inside knowledge. Two walkthroughs show the difference. We cover the three testing depths in full in Black Box vs White Box vs Gray Box Penetration Testing.

Middleware exclusion to authentication bypass. A black box test discovers that /api/v1/users returns data without authentication, but cannot tell whether that is intentional. A code-aware test reads the middleware:

const publicPaths = ['/health', '/api/docs', '/api/v1/users'];
app.use((req, res, next) => {
  if (publicPaths.includes(req.path)) return next();
  return verifyJWT(req, res, next);
});
const publicPaths = ['/health', '/api/docs', '/api/v1/users'];
app.use((req, res, next) => {
  if (publicPaths.includes(req.path)) return next();
  return verifyJWT(req, res, next);
});
const publicPaths = ['/health', '/api/docs', '/api/v1/users'];
app.use((req, res, next) => {
  if (publicPaths.includes(req.path)) return next();
  return verifyJWT(req, res, next);
});

It sees that /api/v1/users is excluded from auth, that the handler accepts ?role=admin&limit=1000 with no additional check, and then proves it with a working request that returns the full admin user list with PII. The finding is not "an unauthenticated endpoint exists," it is "authentication bypass via middleware misconfiguration exposes admin enumeration," with a PoC and a direct code mapping.

Data flow tracing to confirmed SQL injection. Where a WAF might block an external probe and lead the tester to conclude "no injection found," code-aware testing reads the vulnerable path directly:

query = f"SELECT * FROM orders WHERE customer_id = {customer_id}"
return db.execute(query)  # raw string interpolation, no parameterization
query = f"SELECT * FROM orders WHERE customer_id = {customer_id}"
return db.execute(query)  # raw string interpolation, no parameterization
query = f"SELECT * FROM orders WHERE customer_id = {customer_id}"
return db.execute(query)  # raw string interpolation, no parameterization

It confirms the vulnerability exists in code, then tests whether the WAF actually blocks exploitation, proving the injection is reachable rather than assuming the WAF closed it. The same code intelligence that reviews your pull requests powers this, so the exploit agents test the exact code paths flagged during defensive analysis instead of guessing which endpoints might be vulnerable.

The Four Phases of the PR-to-Retest Workflow

The workflow turns 11.4 into a continuous loop instead of a yearly event.

Phase 1. Defensive PR Review Gates

Before code reaches production, automated gates catch vulnerabilities at the source. SAST flags injection and auth-bypass patterns, secret detection runs on every commit, and policy checks enforce input validation and cryptographic standards.

The evidence this generates supports the methodology side of 11.4.1 and the secure development requirements in 6.2. That includes PR comments with CVSS scores, commit history showing when a finding was introduced and fixed, and policy reports mapped to PCI requirements.

Phase 2. Offensive Production Testing

Once code is deployed, autonomous testing validates controls against real attacks. Reconnaissance, code-aware gray box testing and exploit validation cover Broken Object Level Authorization, IDOR, injection, SSRF and authentication bypass.

Only findings with a working proof of concept are reported, with the attack chain that shows impact on cardholder data. This produces the exploitability evidence 11.4.2 and 11.4.3 require.

Phase 3. Unlimited Retesting

This phase satisfies 11.4.4 directly. A fix is validated the moment it deploys, with no scheduling and no per-retest fee.

# Original finding
curl -X GET https://api.example.com/api/v2/users/12345 \
  -H "Authorization: Bearer victim_token"
# 200 OK, attacker accessed victim data

# After the fix, same request
curl -X GET https://api.example.com/api/v2/users/12345 \
  -H "Authorization: Bearer victim_token"
# 403 Forbidden, vulnerability closed
# Original finding
curl -X GET https://api.example.com/api/v2/users/12345 \
  -H "Authorization: Bearer victim_token"
# 200 OK, attacker accessed victim data

# After the fix, same request
curl -X GET https://api.example.com/api/v2/users/12345 \
  -H "Authorization: Bearer victim_token"
# 403 Forbidden, vulnerability closed
# Original finding
curl -X GET https://api.example.com/api/v2/users/12345 \
  -H "Authorization: Bearer victim_token"
# 200 OK, attacker accessed victim data

# After the fix, same request
curl -X GET https://api.example.com/api/v2/users/12345 \
  -H "Authorization: Bearer victim_token"
# 403 Forbidden, vulnerability closed

Validation happens within hours instead of weeks. That's the difference between meeting 11.4.4 continuously and hoping the fix held until next year's test. Penetration test retests covers what auditors look for in retest evidence.

Phase 4. Continuous Monitoring

Between annual attestations, scheduled tests and change detection keep the evidence current. Segmentation is validated at least every 12 months under 11.4.5, or every six months for service providers under 11.4.6.

Significant-change detection triggers automatic retests, so each change in your log has a matching test.

Dimension

Traditional

PR-to-retest

Testing frequency

Annual plus manual triggers

Continuous plus scheduled

Retesting speed

Weeks, waiting on a new engagement

Hours

Code awareness

Black box only

Gray box with codebase intelligence

Coverage between tests

None until the next engagement

Continuous validation

Cost model

Per engagement

Performance-based

The takeaway is simple. Traditional testing meets the annual clause, and the PR-to-retest loop meets the "after any significant change" and retest clauses too.

Making It Operational

Five steps turn the workflow into a running program.

1. Define scope with code-aware analysis and version-control the result, so the CDE boundary, segmentation controls, and connected systems are a living artifact rather than a stale diagram:

# pci-scope-snapshot.yml (version-controlled)
scope_date: "2026-03-15"
cde_components: [payment-api v2.3.1, cardholder-db (PostgreSQL 14.2)]
segmentation_controls: [VLAN 100, AWS Security Group sg-0a1b2c3d]
connected_systems

# pci-scope-snapshot.yml (version-controlled)
scope_date: "2026-03-15"
cde_components: [payment-api v2.3.1, cardholder-db (PostgreSQL 14.2)]
segmentation_controls: [VLAN 100, AWS Security Group sg-0a1b2c3d]
connected_systems

# pci-scope-snapshot.yml (version-controlled)
scope_date: "2026-03-15"
cde_components: [payment-api v2.3.1, cardholder-db (PostgreSQL 14.2)]
segmentation_controls: [VLAN 100, AWS Security Group sg-0a1b2c3d]
connected_systems

2. Integrate PR checks in CI/CD so insecure patterns are caught pre-merge and blocked on critical severity, which produces the pre-merge evidence trail 11.4.1 expects.

3. Set objective triggers for significant changes, any change to authentication or authorization, payment logic, a new endpoint touching the CDE, a cryptographic update, or an infrastructure-as-code change affecting segmentation, rather than relying on subjective judgment. This is what the "after significant change" clause in 11.4.2 and 11.4.3 actually demands.

4. Build retest SLAs by severity, critical within 24 hours, high within 72, medium within 7 days, with closure criteria that require the root cause fixed, not just the specific exploit path. This operationalizes 11.4.4.

5. Structure your audit artifacts so every finding traces from scope snapshot to report to remediation commit to retest validation, which is the exact evidence trail a QSA follows.

PCI Penetration Testing Cost

PCI penetration testing cost depends on how many CDE segments, applications and APIs are in scope, how many segmentation controls need validation, and whether retests are included. Retests are the hidden line item. If a provider charges per retest, the fix-and-verify cycle in 11.4.4 can add materially to the original quote.

See penetration testing cost in 2026 for pricing by test type. CodeAnt AI's penetration testing is free to start and charges only when it confirms a high or critical exploitable vulnerability, with unlimited retests included.

When Manual Testing Still Matters

AI-driven testing handles the bulk of PCI 11.4 obligations. Reserve manual budget for problems where human expertise adds real value, such as novel attack research, highly specialized business logic in custom payment workflows, and red team engagements that test incident response.

The efficient split runs continuous automated testing for the baseline, meaning reconnaissance, standard attack chains and unlimited retest validation. A periodic manual engagement then covers the hard problems, such as multi-stage attacks that need business context and deep validation of complex remediation. You get continuous compliance without paying twice for the same coverage.

PCI Penetration Testing for Insurers and Insurtech

Insurers fall under PCI DSS wherever they take card payments for premiums, deductibles or policy fees. Online payment portals, agent-assisted payments and embedded checkout flows all bring systems into scope.

Cyber insurance carriers ask about it too. The Corvus application asks whether you or your payment processor are PCI compliant, as covered in the cyber insurance questionnaire guide. For the insurance regulations that apply alongside PCI, see penetration testing for insurance companies. For how the same pentest report supports a cyber insurance application, see cyber insurance requirements.

Stop Marking a Checkbox. Prove the Controls Work.

PCI DSS 4.0.1 doesn't ask you to choose between manual and automated testing. It asks you to prove your security controls work, continuously, and to retest every exploitable finding until it's closed.

The PR-to-retest workflow does exactly that. Defensive review blocks vulnerabilities before merge, offensive testing validates controls in production, and unlimited retesting satisfies 11.4.4 without weeks-long delays. That's what CodeAnt AI unifies. The same code intelligence that reviews your pull requests runs the offensive test and validates the fix, so there's no gap between the defensive and offensive sides.

Every finding ships with a working proof of concept, mapped to the PCI control it affects and the exact file and line. Retests are unlimited, and you pay only when a high or critical is confirmed exploitable.

Where to Start This Week

Pull your CDE scope and check it against your code instead of your network diagram. Trace where cardholder data actually flows and note every system it touches, because that's your real 11.4 scope.

Then wire one PR gate and one post-deploy gray box scan on the highest-risk payment path, and require a working proof of concept plus a retest for anything critical. That single loop is the core of what 11.4.4 asks for, running continuously instead of once a year.

Run a free code-aware pentest →

Related Reading

FAQs

Is automated penetration testing acceptable for PCI 11.4?

What counts as a "significant change" that triggers retesting?

Do we need both internal and external penetration tests?

How fast should retesting happen after remediation?

How do PR checks relate to 11.4 versus the secure-coding requirement?

Start Your 14-Day Free Trial

AI code reviews, security and quality trusted by modern engineering teams.

Table of Content
No headings found on page

Ship clean & secure code faster

Get Pentest Report

NO CC REQUIRED