# What does penetration testing cost in 2026?

> Penetration testing cost in 2026 ranges from focused $4,000 tests to $60,000+ programs. Learn what changes a quote and when to spend less.

The useful answer is narrower than most sales pages admit: a penetration test for a small, production web product usually lands between $5,000 and $15,000 in 2026, while a complex application, cloud environment, or compliance program can run from $15,000 to $60,000 or more. A tightly bounded test can cost less. A red team exercise across people, offices, cloud accounts, and detection systems can cost far more.

Those numbers mean little until you define the target. A quote for one web application with two user roles is not comparable to a quote that includes its API, mobile clients, cloud account, internal network, and a retest. Founders often compare the totals and assume the expensive vendor has a stronger method. I have bought enough technical assessments to distrust that shortcut. Price follows tester time, scope uncertainty, reporting obligations, and the risk the firm accepts.

The sensible purchase is the smallest test that answers a business question. That might be, "Can one customer reach another customer's data?" It might be, "Will this report satisfy the security review holding up our contract?" If nobody can state the question, postpone the quote and fix the scope first.

## The 2026 price bands start with scope

Current public prices support a broad market band, but they do not create a universal rate card. Cobalt advertises a promotional autonomous web application test at $3,500 during 2026. Pentest Testing Corp publishes a $5,000 starting price for a focused web application or API test, $9,500 to $25,000 for many production SaaS products, and $18,000 to $60,000 or more for complex work. Public UK procurement documents show consultant day rates from roughly £750 to £1,400. Treat these as visible anchors from particular sellers, not an average of the whole market.

For a US buyer planning a budget, several bands are reasonable before tax, travel, emergency work, or specialist hardware. An automated vulnerability scan runs from free to about $2,500 and provides tool output with limited manual validation. A focused web or API test usually runs from $4,000 to $10,000 for one small target, a few roles, one environment, a report, and sometimes one retest.

A production SaaS application commonly calls for an $8,000 to $25,000 budget when the work covers several roles, business workflows, APIs, evidence, and a retest. A mobile application with its backend API often lands from $10,000 to $30,000. Cloud or internal network work falls in a similar $8,000 to $30,000 band, depending on accounts, hosts, identity paths, segmentation, and how far the tester may exploit a finding.

Complex or compliance driven programs start around $20,000 and can pass $60,000 when they span multiple targets, environments, formal evidence requirements, or repeated testing. A red team exercise commonly starts around $30,000 and can exceed $100,000 because it pursues an agreed goal across several control layers. None of these planning bands replaces a scope based quote.

The first row is intentionally separate. A vulnerability scanner checks known patterns at machine speed. A penetration tester uses judgment, changes approach when a control resists, tests authorization and business logic, and may combine weak findings into a meaningful attack path. Calling a scan a pentest does not make it one. A $1,500 PDF can still be useful, but only if you buy it as a scan and do not promise customers that a human tested the product.

Day rates explain much of the spread. A five day engagement does not contain five full days of attacking. The firm scopes the work, prepares accounts and tooling, performs the test, validates findings, writes the report, reviews it, holds a readout, and later checks fixes. When a quote looks surprisingly cheap, ask which of those activities disappeared.

## Vendors price the attack surface they can count

A tester cannot quote "our SaaS" with any precision. Vendors turn the product into countable units: live hosts, API routes, user roles, tenant boundaries, mobile platforms, cloud accounts, identity providers, integrations, and critical workflows. Every unit adds paths to inspect and combinations to attempt.

User roles matter more than many founders expect. An application with owner and member roles has four obvious authorization directions: member to owner, owner to member, one member to another, and one tenant to another. Add support impersonation, reseller administration, or delegated billing and the matrix grows. Ten nearly identical screens may take less work than one complicated import workflow that parses files, calls third parties, and changes permissions.

API counts also mislead when teams submit raw route totals. Fifty generated CRUD routes over the same authorization layer can be simpler than six endpoints that move money, invite users, exchange OAuth tokens, and accept webhooks. Give vendors an endpoint inventory, but mark the operations that cross trust boundaries. Good testers price the thinking, not the number in an OpenAPI file.

Ask every bidder to estimate testing days against the same scope sheet. This compact artifact is enough for an initial quote:

> **Target:** staging web application configured like production and its public API  
> **Roles:** anonymous, member, workspace owner, support administrator  
> **Critical workflows:** signup, invitation, password reset, file import, billing change, account deletion  
> **Tenant model:** shared application and database, tenant ID enforced in the application layer  
> **Excluded:** denial of service, social engineering, production customer data, payment pages operated by another provider  
> **Deliverables:** technical report, management summary, evidence for each confirmed finding, readout, one retest within 30 days

Send the same page to three firms. If their day estimates differ sharply, ask each one to state what it will sample and what it will cover exhaustively. That conversation reveals more than a certification list.

## Access changes depth, time, and price

Black box, gray box, and white box describe how much information the tester receives. Buyers often treat black box work as the most realistic because an outside attacker starts without credentials. For a small product with a limited budget, that logic usually wastes paid hours on discovery your team could provide in an afternoon.

A black box tester receives little beyond a target and the rules. This approach measures the exposed path from an outsider's position, but it spends time finding hosts, mapping routes, registering accounts, and guessing how the system works. It fits an external perimeter review or a test of what an unauthenticated attacker can discover. It provides weak coverage of paid features, administrative functions, and cross tenant authorization unless the tester can reach them.

A gray box tester gets test accounts, role descriptions, API documentation, and often an architecture briefing. This is the best default for a startup application. The tester still attacks the running system but reaches important states quickly. More of the fee goes toward permission failures, session handling, workflow abuse, injection, file processing, and exploit chains.

White box work adds source code, configuration, infrastructure definitions, or direct engineer access. It can find issues that dynamic testing may miss, and it lets the tester trace suspicious behavior to its cause. It may cost more if the contract adds formal code review across a large repository. It may cost less when selective source access replaces hours of blind probing. Ask whether source assistance is included in the test or sold as a separate review.

Do not confuse access level with tester independence. Giving a firm two customer accounts and an architecture diagram does not coach it toward a passing result. It lets the firm spend its time on the risky parts. Keep one unauthenticated phase in the rules if outsider visibility matters, then provide credentials for the rest.

## Scope uncertainty is an expensive line item

Vendors charge for what they can see and add a buffer for what they cannot. A founder who says "the API is small" while omitting three GraphQL services, a support console, and a legacy upload host creates uncertainty. The bidder either raises the price, narrows the contract, or absorbs the surprise and rushes the last days. None produces good coverage.

Several quote drivers have little to do with raw application size. A rushed start can require a different tester or overtime. Testing production may require narrow windows, engineers available during the test, rate limits, and rollback plans. Sensitive data can add handling controls. A customer deadline may require a specific report format, tester biography, independence statement, or evidence mapped to a standard. Each obligation consumes time.

Retesting deserves its own line. Some firms include one verification round for findings fixed within 30 or 60 days. Others sell retesting by the day, include it only in an annual subscription, or verify only high severity items. "Free retest" can still have limits on timing, eligible findings, and changed functionality. Put the window and coverage in the contract.

Cloud policy can also change the plan. AWS currently permits testing against listed customer services without prior approval, but it prohibits several activities, including denial of service tests, and requires approval for command and control activity. It also forbids customers from testing AWS services themselves. Other providers and software vendors have their own rules. The NIST SP 800-115 definition of rules of engagement is useful here: agree on detailed constraints and authority before testing begins. A vendor that has not asked where it may attack has not finished scoping.

Stability affects the quote even when it stays off the proposal. If developers deploy breaking changes every day, evidence becomes stale and the tester repeats setup. Freeze the relevant build, seed known data, keep test accounts alive, and name an engineer who can answer blocking questions. These small acts buy more useful attack time without negotiating the rate.

## A credible quote specifies the work you will receive

The total is the least informative part of a penetration testing proposal. A credible quote connects money to targets, effort, method, evidence, and limits. If two proposals do not define those items the same way, their totals cannot be compared.

At minimum, expect the statement of work to name the targets and environments, testing dates, access level, included roles, excluded techniques, expected tester days, reporting format, communication process, retest terms, and rules for a scope change. It should say whether the tester will validate scanner results manually and whether the final report distinguishes confirmed vulnerabilities from unverified observations.

Method names help only when the proposal says how the firm will use them. The OWASP Web Security Testing Guide assigns identifiers to concrete test scenarios, such as identity, authorization, session, input validation, and API tests. Ask which WSTG areas apply to your product and where the tester will go beyond them for business logic. A claim to "cover OWASP" is vague because OWASP publishes several different projects and none turns a time bounded test into exhaustive proof.

Read a redacted sample report before signing. You want reproducible evidence: affected target, preconditions, request or action, observed response, impact, and a fix that fits the cause. Severity should reflect exploitability and business impact rather than the scanner's default label. Screenshots alone are weak evidence when a request and response can show the actual authorization failure.

Compare two offers. Offer A costs $6,000 and promises a "web application pentest" for one generic user, automated and manual testing, and a PDF report. It says nothing about effort or retesting. Offer B costs $10,500 and names the web application, documented public API, four supplied roles across two tenants, six tester days, report review, technical and management reporting, a live readout, and one retest within 45 days.

Offer A is not automatically bad. It may answer a narrow external question. Offer B buys a defined authorization matrix and closure after fixes. Choose based on the question you need answered, then make the cheaper vendor write that same scope if it claims the work is equivalent.

Insurance, references, and certifications reduce procurement risk, but they do not rescue a vague scope. Ask who will actually test, how the firm reviews that person's findings, and what happens if the assigned tester becomes unavailable. A famous logo on the proposal matters less than a senior reviewer who catches a false conclusion before it reaches your customer.

## Compliance pays for evidence, not immunity

A compliance driven pentest costs more when the buyer needs proof that matches a control, an auditor's expectations, or a customer's questionnaire. That premium buys documentation and process. It does not buy a secure product, and a clean report does not prove that no exploitable bug exists.

PCI DSS is specific. Version 4.0.1 places penetration testing in Requirement 11.4 and distinguishes external and internal testing, segmentation testing where applicable, remediation, and repeated testing. Scope and frequency depend on the entity and requirement. If payment card obligations drive the purchase, give the tester your cardholder data environment and segmentation scope rather than asking for a generic SaaS test. Confirm the required assessor evidence before the engagement, not after the report arrives.

SOC 2 works differently. The AICPA Trust Services Criteria do not impose one universal annual pentest on every company. Your controls, risk assessment, auditor, and customer commitments determine what evidence you need. Many buyers and auditors still expect an independent report because it is easy to review. Ask the auditor whether it needs a pentest, a vulnerability assessment, proof of remediation, a date inside the audit period, or all of these. The answer can change the scope by thousands of dollars.

ISO 27001 also uses a risk based management system rather than a single priceable pentest recipe for every organization. A sales representative who says a premium package is "required for ISO" owes you the exact control, your stated control implementation, and the auditor's evidence request. Do not buy a larger engagement from a loose compliance claim.

The OWASP guide makes a point many reports bury: a penetration test finds a representative sample of possible security risks. I agree, and I would make that limitation explicit in any customer conversation. The report shows that qualified people tested a defined system during a defined window. Your development practices determine what happens the next morning.

## Early products should test one dangerous slice

An early product rarely needs the broadest test it can imagine. It needs a focused test around the failure that could kill a deal, expose customer data, move money, or hand an attacker privileged access. A narrow manual engagement on that slice is more honest than a cheap claim that the whole platform was tested.

For a multitenant SaaS product, the first slice is often identity and tenant separation. Include signup, login, password reset, invitation, role change, object access, export, and deletion across at least two tenants. For a marketplace, include price and payout changes, refund paths, seller identity, and webhook handling. For an AI product, include who may supply instructions, which data can enter context, what tools the model may call, and whether one user's data can appear in another user's result.

Set the boundary in business language before translating it into URLs. "Test whether a member of Workspace A can view or change Workspace B's records through the browser or public API" is a useful objective. "Test app.example" is a hostname. The objective tells the tester where to spend judgment when time runs short.

A focused test can exclude low consequence marketing pages, denial of service, broad cloud configuration, employee phishing, and third-party systems you cannot authorize. State those exclusions in the report so nobody later treats the result as broader than it was. Add a second scope only when it addresses a separate near term decision.

This approach works only after basic hygiene. Do not pay a senior tester to report default credentials, public debug endpoints, missing dependency patches, or a storage bucket your own tooling already flags. Run routine checks, fix the obvious issues, and give the tester the residual risk. The test then has a better chance of finding a broken authorization rule or workflow exploit that a machine will not understand.

Avoid the popular "one annual pentest covers security" recommendation. It persists because procurement likes a dated report and vendors like a repeatable sale. A release made a week later can invalidate a finding, while untouched dangerous code can remain vulnerable between annual visits. Use the report for independent challenge and evidence, then keep cheaper controls running with each change.

## Cheaper alternatives answer different questions

There is no cheaper direct replacement for a skilled person trying to exploit your product. There are cheaper ways to answer narrower questions, and an early team should use them before and between manual tests.

Dependency scanning finds known vulnerable packages for little or no direct cost, but it misses custom code and most business logic. Static analysis finds suspicious source patterns, while dynamic scanners probe a running application for known web weaknesses. Both range from free tools to paid platforms. Neither reliably reasons through tenant abuse or a complex sequence of valid actions.

A cloud configuration review compares accounts and services with a chosen baseline. It can expose dangerous identity or storage settings, but it does not test application logic. Threat modeling uses team time or a specialist workshop to locate boundaries where data, money, or authority can cross. It directs later testing but provides no proof that a live exploit works.

A focused code review asks whether one sensitive implementation is correct. It sees detail that an outside tester may miss but says little about code outside the selection. A bug bounty invites researchers to submit eligible findings over time and charges platform fees, rewards, and triage effort. It does not promise predictable coverage or a finished report by a customer deadline.

These controls compound. A threat model identifies the payment webhook and tenant export as dangerous. Static analysis and code review inspect their implementation. Automated tests assert authorization on every build. A manual tester then tries to bypass the whole chain in a deployed environment. Paying for each layer to repeat a generic injection scan wastes money.

An automated platform can be a sensible bridge when a customer wants recent external evidence and the budget cannot support a broad manual engagement. Read the deliverable before buying. Determine whether a human validates findings, whether the report calls itself a penetration test, what target and workflows it covers, and whether your customer or auditor accepts it. The low price is irrelevant if procurement rejects the report.

Internal testing can also help, especially before product market fit. An engineer who did not build the feature can abuse role changes, alter object identifiers, replay webhooks, and inspect storage permissions. Document the scope and evidence. Independence may not satisfy a customer or standard, and familiarity can hide assumptions, but internal work removes easy findings before an external test.

A Team & AI Audit from oleg.is is not a security assessment; I use it to identify engineering operating savings that can fund focused specialist work without pretending the two services are interchangeable. Cutting the wrong security scope to protect an inefficient delivery process is poor arithmetic.

## Buy when the result changes a decision

The right time to purchase a test is when its result will change a release, contract, risk acceptance, or architecture decision. "Security is important" is true but too vague to schedule a useful engagement.

Buy before a high consequence launch when the application will hold sensitive customer data, move funds, administer infrastructure, or expose powerful automation. Buy when an enterprise contract explicitly requires an independent report and the revenue justifies the work. Buy after a major change to identity, tenant isolation, payment flow, cloud boundary, or externally reachable architecture. Buy after remediation when you need an independent party to confirm that an exploit path is closed.

Wait when the product changes so quickly that the tested build will disappear before the report arrives, unless the risky architecture itself is stable. Wait if the team has not fixed obvious scanner findings. Wait if no one can provide working accounts, seed data, an endpoint inventory, or a responsible engineer during the test. The vendor will spend expensive time reconstructing your product and may still miss the important states.

Do not wait merely because the product is small. A tiny service that signs transactions or can read every customer's documents has a small codebase and a large consequence. Do not buy merely because the product is large. A broad test of an unstable, low sensitivity prototype can produce a long report that changes nothing.

Use a simple decision record with four fields: the decision the report supports, the loss scenario you care about, the exact target, and the date by which evidence matters. Add the buyer or auditor who will consume the report. If the expected contract is worth $20,000 and a compliant test costs $25,000, the deal alone does not justify it. If the same report supports several contracts and tests a risk you already need to address, the economics change.

Schedule remediation capacity before testing. A report with ten confirmed findings creates engineering work, clarification meetings, deployments, and retesting. If the team cannot touch the results for three months, move the test closer to the repair window unless a deadline prevents it. Fresh evidence and an available developer shorten the path from finding to verified fix.

## Cut the quote by removing uncertainty

The cleanest negotiation tactic is better preparation. Vendors lower prices when they can estimate effort and believe the environment will cooperate. Asking for a discount without changing scope usually removes tester time invisibly.

Provide a one page architecture diagram, host and API inventory, role matrix, two accounts per relevant tenant, seeded test data, and a short description of every critical workflow. Mark components run by third parties and state who can authorize testing. Share known findings and previous reports. Hiding them does not create an independent test; it pays the tester to rediscover old work.

Then ask the bidder to separate optional scope. Price the web application, public API, cloud account, mobile client, and retest as explicit components. You may learn that the API is already exercised through the web workflows, or that the internal network has no bearing on the customer request. You may also learn that excluding the support console leaves the most powerful role untested. Make those choices in daylight.

Offer a stable staging environment that matches production security controls without containing real customer data. Give the tester a direct technical contact and agree on response times for blocked accounts or unclear behavior. Allow safe evidence collection. If production must be tested, define rate limits, forbidden actions, monitoring contacts, and stop conditions.

Ask for fewer management artifacts if nobody needs them. A board presentation, several report formats, custom risk mapping, and repeated readouts all consume senior time. Do not cut the technical reproduction steps or review process. The engineers fixing the issue need exact evidence, and the reviewer protects you from weak findings.

Finally, reserve scope changes for genuine discoveries. Adding a mobile application on day three is a new test, not a clarification. A fair contract states the rate or approval process for extra work. That protects both sides and prevents the quiet squeeze in which a tester races through an enlarged target to preserve a fixed fee.

The price you should accept is the lowest one that still funds enough competent time to answer your stated question and produce evidence your reader will accept. Anything cheaper is a different product. Anything broader should justify itself with a second decision you actually need to make.
