
Pentesting Resources & Buyer Guides
Structured evaluation tools that help organisations compare methodology, reporting quality, provider type, and commercial terms before selecting a penetration testing vendor.
Pentesting resources and buyer guides are structured evaluation tools that help organisations compare methodology, reporting quality, provider type, and commercial terms before selecting a penetration testing vendor.
This guide is for buyers who already know they need a penetration test and now need a defensible way to shortlist providers. Use it to:
- separate manual, expert-led testing from scanner output
- weight criteria to your actual driver
- run vendor conversations that expose weak scope, weak reports, and weak post-engagement support
It is not a tools list. It is a procurement framework you can take into scoping calls, RFPs, and internal sign-off.
What Separates a Real Pentest from an Automated Scan
A penetration test is a manual, expert-driven security assessment that attempts to exploit vulnerabilities to understand real-world business impact, not a report generated by automated scanning tools.
A vulnerability scan is automated, matches known weakness signatures, and produces a list. A penetration test is authorised, human-led work that chains findings, attempts exploitation, and explains what an attacker could actually do to the business.
Vulnerability Scanning
Automated tooling against known signatures. Fast, broad, and useful as inventory — but it produces a list of potential issues, not proof of impact.
Penetration Testing
Manual testing plus targeted tooling, directed by an experienced tester. Exploited or validated findings, attack paths, and business impact.
Red teaming is a different service: an objective-led simulation of a determined adversary, not a scoped vulnerability assessment.
Do not buy a scan labelled as a pentest, and do not buy a pentest if what you actually need is a red team.
| Factor | Vulnerability scan | Penetration test |
|---|---|---|
| Method | Automated tooling against known signatures | Manual testing plus targeted tooling, directed by an experienced tester |
| Output | A list of potential issues, often CVSS-scored | Exploited or validated findings, attack paths, and business impact |
| Depth | Known CVEs and configuration checks | Logic flaws, access-control failures, privilege escalation, multi-step chains |
| Business context | Rarely present | Required: what the finding means for data, operations, and reputation |
| Remediation value | Generic patch or config advice | Proof-of-concept steps and specific fix guidance |
Treat scanner-only work as inventory, not assurance. Automated scans miss business-logic errors, broken authorisation, and chained paths that never appear as a single high CVSS row.
That is the commercial risk: a cheap scan sold as a pentest can still generate an audit-looking PDF while leaving the path to data or admin access untested.
The most reliable way to test whether a vendor does genuine manual testing is to request a redacted report sample. Look for:
- proof-of-concept steps
- exploitation chains
- business context
Ask whether the sample shows how two or three medium issues became a high-impact path. A CVSS list with no attack narrative is a scan dressed as a pentest.
The Core Evaluation Criteria: What to Assess Before You Shortlist Any Vendor
Assess every pentesting vendor against seven criteria before they enter a shortlist: methodology, scope, reporting, retesting, tester qualifications, timeline, and data handling.
Score vendors on evidence, not brochures. A defensible shortlist is built from written methodology, a redacted report, named testers, and clear commercial boundaries.
Prepare a draft scope before you ask for quotes. Vague requirements produce vague proposals, and vague proposals hide exclusions.
Methodology and quality
Methodology quality is proven by a written, manual-first process that includes attack path analysis, not by marketing claims about ethical hacking.
Ask for a sample methodology document. A strong document states how testers combine reconnaissance, manual exploitation, and attack-path analysis, and where automation is used only as a starting point. Ask directly how they separate manual testing from scanning, and how they report chains of issues rather than isolated CVEs.
Attack path analysis is the difference between a list of weaknesses and an explanation of how an attacker would actually move.
Low or medium findings often matter only when they combine: an information leak, a weak access-control check, and a privilege-escalation bug can become a business-critical path. Ask vendors how they identify, validate, and write up those chains. If they can only talk about individual CVSS rows, they are not testing the way attackers work.
Scope and approach
A thorough scope defines the attack surface, test types, exclusions, and boundaries before any testing window opens.
A well-defined scope names systems, applications, APIs, and networks in play, the test types (for example external, internal, web application), what is explicitly out of bounds, and how newly discovered assets are handled. It should also state authentication levels, test accounts, environments, and whether cloud, APIs, and third-party integrations are in or out.
Vague scope is the usual reason critical systems go untested while the report still looks complete.
Read the exclusions as carefully as the inclusions. If the proposal omits identity providers, admin interfaces, or the APIs that actually move data, the engagement will not answer the question you think you bought.
Reporting and remediation
A high-quality pentest report includes an executive summary, methodology, findings with CVSS scores and business context, proof-of-concept steps, remediation guidance, and a clear risk rating.
Request a redacted report sample during sales. Score it against the checklist below. If the sample cannot be used by both a CISO and an engineer, the live report will not be usable either.
CVSS is a useful severity input. It is not a substitute for exploitability, asset criticality, or business impact. A high finding that cannot be reached in your architecture is not the same as a medium finding that leads straight to customer data.
Report Quality Scorecard
- Executive summary that a non-specialist can act on
- Business context for each significant finding
- CVSS scores used with explanation, not as the only ranking method
- Proof-of-concept steps that show how the issue was exploited or validated
- Attack paths or chained findings, not only isolated issues
- Remediation guidance specific enough to assign to an owner
- Clear risk ratings that reflect exploitability and impact, not scanner defaults
- Language that is precise and readable, without unexplained jargon
Mark each item present, partial, or absent on the sample. A vendor who cannot share a sample, or whose sample fails more than two of these points, should not reach the shortlist.
Retesting policy
A usable retesting policy states whether retesting is included in the original price, what is in scope for retest, and how quickly it can be scheduled.
Get this in writing before award. Confirm whether retest covers only the listed findings, related controls, or a full re-run, and whether a revised report is included. An extra charge or an undefined wait after remediation is a common budget surprise and delays closure of audit findings.
Tester qualifications
Tester quality is assessed by the named people assigned to your engagement, their certifications, and their years of hands-on testing, not by a generic team of experts line.
Ask who will actually test. Look for OSCP, OSWE, CREST, or equivalent credentials, plus evidence of relevant stack and industry experience. Ask whether those named people can be substituted without your agreement.
You are hiring the testers. The company brand does not sit at the keyboard.
Timeline and process
A realistic timeline is set by scope and attack surface and should cover a scoping call, a defined testing window, reporting, and retesting.
Small, well-bounded web application tests can complete in days of testing plus reporting time. Multi-system internal and external work often needs several weeks from kick-off to debrief.
Ask for the full sequence, including how findings are communicated during the window if a critical issue appears, and ask about lead time to start. Strong testers book ahead, and a short remaining compliance deadline is not a reason to accept the first available scanner-led slot.
Data handling and confidentiality
Data handling must be explicit: storage location, access controls, NDA terms, retention, and UK GDPR obligations.
Pentesting exposes credentials, architecture, and sometimes production data. Ask where evidence is stored, who can access it, how long it is kept, and how it is destroyed. Ask whether screenshots, logs, and payloads leave your environment, and whether subprocessors or platforms will hold copies.
For UK and EU buyers, GDPR processing terms, lawful basis, and data-processing agreements are procurement items, not footnotes. If you need UK or EU residency for evidence, say so before scoping, not after the first screenshot is taken.
Choose Your Evaluation Priority: Match Your Criteria to Your Context
Match your pentesting evaluation criteria to one primary driver: compliance, risk reduction, security maturity, or incident response.
A compliance-driven pentest and a risk-reduction pentest are not the same engagement. A test can satisfy an auditor’s evidence request and still miss the attack path that matters to the business.
Decide which driver you are optimising for before you compare vendors. UK and EU programmes often have a compliance clock (GDPR evidence, DORA, PCI DSS, ISO 27001) running at the same time as a real risk need. Name which one governs scope if they conflict.
Self-assessment: pick the statement that is most true, then weight your shortlist to that driver.
Compliance-driven
You need evidence for GDPR, DORA, PCI DSS, ISO 27001, or a similar audit. Prioritise clear scope documentation, report structure, regulatory familiarity, and a deliverable an auditor can file without translation.
Risk reduction
You want exploitable issues found and fixed before an attacker uses them. Prioritise manual depth, attack path analysis, proof-of-concept quality, and usable remediation support.
Security maturity
You already test regularly and want continuous or advanced coverage, including APIs and AI/LLM applications. Prioritise PtaaS-style cadence, specialist skills, and repeatable scoping.
Incident response
You have had a breach or suspicious activity and need a focused assessment quickly. Prioritise rapid start times, tight confidentiality, and testers who have worked in breach or forensics-adjacent contexts.
If two drivers apply, name a primary and a secondary. Do not let a compliance deadline silently strip manual testing out of the scope.
An audit-shaped report that never attempted exploitation can still leave you exposed, and it can still fail you later when a real incident shows the path the test never walked.
Consultancy vs. PtaaS vs. Crowdsourced Testing: Which Provider Type Fits Your Needs?
Choose consultancy, PtaaS, or crowdsourced testing by matching testing depth, reporting, engagement model, and data-handling needs to your organisation, not by which model is currently fashionable.
You can use more than one model across a year. Each individual engagement still needs one accountable owner, one scope, and one report standard. Do not mix community testers, a portal, and a consultancy brand in the same contract unless you know who signs the findings and who holds your data.
Consultancy
A consultancy pentest is a project-based, human-led engagement that delivers deep manual testing, named testers, and a formal point-in-time report.
This model fits compliance-driven work, complex architectures, and buyers who need a single accountable team. Senior testers and narrative reporting are the usual strengths. It also fits poorly documented estates where scoping itself is part of the value, because a named team can sit with your architects and still change the test plan when the surface is larger than expected.
When to avoid: a full consultancy project can be the wrong fit for simple, high-frequency regression testing where you already have a mature internal security function and only need recurring coverage of a stable, well-understood surface.
PtaaS (Penetration Testing as a Service)
Penetration Testing as a Service (PtaaS) is a platform or subscription model that delivers recurring tests, faster turnaround, and often self-service scoping.
This model fits security-mature teams that need ongoing testing rather than a single annual snapshot. The value is cadence and operational speed, provided the underlying testing is still manual where it matters. Ask whether the subscription buys tester hours and attack-path work, or mainly scanner cycles with a human summary.
When to avoid: PtaaS is a poor fit when you need deep manual testing of a complex, poorly documented architecture, or when legal and data-handling rules forbid platform-based evidence storage.
Crowdsourced testing
Crowdsourced testing uses a platform to connect your attack surface with a community of testers under a broad-coverage or pay-for-results model.
This model can fit large web and mobile surfaces where diverse skill sets and volume of eyes matter. It is weaker where you need one named owner, tightly controlled data, or a single coherent report for an auditor. Breadth is not the same as a joined-up attack path written for your business.
When to avoid: crowdsourced testing is a poor fit for strict data-handling regimes, CNI-adjacent environments, and buyers who need a single accountable point of contact from scoping through retest.
| Provider type | Testing depth | Reporting | Engagement model | Data handling | Best fit |
|---|---|---|---|---|---|
| Consultancy | High manual depth on a defined scope | Formal, narrative, audit-friendly | Project, point-in-time | Usually contracted under NDA with a single firm | Compliance, complex systems, named accountability |
| PtaaS | Variable; strong when the platform still funds manual work | Often portal-based, faster, sometimes thinner narrative | Subscription or retainer, recurring | Evidence often lives on a vendor platform | Mature teams needing continuous coverage |
| Crowdsourced | Broad coverage, uneven depth per tester | Finding-led; less consistent as a single document | Community / pay-for-results | More people see more of the surface | Large web/mobile surfaces with flexible data rules |
Decision box
- If your primary driver is compliance, start with consultancy.
- If it is security maturity and recurring coverage, start with PtaaS.
- If it is breadth across a large application estate and data sharing is acceptable, crowdsourced testing can be in the mix.
- If the driver is incident response, choose the type that can start fastest with named, experienced testers and tight confidentiality.
Evaluating AI and LLM Pentesting Capabilities
AI and LLM pentesting is specialised testing of model-backed applications for prompt injection, data leakage, model manipulation, training-data poisoning, and insecure output handling, not a relabelled web application test.
Traditional web application testing still matters for the surrounding app, APIs, and access control. It does not automatically cover model behaviour, prompt-injection paths, or data flow through retrieval and tool-calling pipelines.
Testers need to understand how prompts, context windows, retrieved documents, plugins, and downstream actions combine. Treat AI testing claims as unproven until the vendor shows a dedicated method and AI-specific findings.
Ask these questions and keep the answers on file:
- Do you have a dedicated AI/LLM testing methodology, written down?
- Which frameworks do you use, for example the OWASP Top 10 for LLM Applications?
- Can you provide redacted examples of AI-specific findings from past engagements?
- How are testers trained on AI security, and who would be assigned?
- How do you test data leakage, insecure output handling, and tool or plugin abuse in the actual pipeline?
- Will the test cover the model integration, the retrieval layer, and the application that acts on model output, or only the chat UI?
AI pentesting is still an emerging field. Vendors can claim capability without demonstrated experience. A generic web methodology plus a promise to look at the chatbot is not enough. Case studies, redacted AI findings, and a named methodology are the credibility filter. Scepticism is the correct default.
Pentesting Pricing Models: How to Compare Costs Without Getting Burned
Pentesting is typically sold as fixed-scope, hourly, or subscription/retainer pricing, and each model changes depth, flexibility, and cost risk.
| Model | How it works | Advantage | Risk | Best for |
|---|---|---|---|---|
| Fixed-scope | A set price for a defined surface and test type | Budget certainty | Scope may be tightened to protect margin; depth can suffer | Compliance-driven, well-defined engagements |
| Hourly | Fees track time spent | Can follow unexpected attack paths | Total cost is harder to predict; weak process can inflate hours | Exploratory or poorly documented environments |
| Subscription / retainer | Recurring fee for ongoing testing | Continuous coverage and faster re-entry | Can be excessive for a one-off audit need | Mature programmes with repeated testing |
Do not treat price as a quality proxy. Expensive vendors can still deliver scanner-like reports. Lower-cost vendors can still deliver excellent manual work.
Compare what is included: scoping calls, report revisions, retesting, and post-engagement access to testers. Ask for the extras tariff in writing: out-of-scope assets, extra days, rush reporting, and retest rounds.
The cheapest fixed-scope quote is often the most dangerous, because it usually implies a tightly constrained surface. Ask what is excluded from the fixed scope, not only what is listed. Ask how the vendor behaves if the surface is larger than the questionnaire suggested. If a critical path sits outside the price, you have bought a certificate, not a test.
Questions to Ask Every Pentesting Vendor Before You Sign
Ask every pentesting vendor the questions below and judge the answers against the quality signals, not against confidence in the sales call.
If a vendor hesitates to provide a redacted report sample, treat that as a red flag and stop the process until a sample appears.
Take notes against each answer. A fluent sales narrative that cannot be matched to a document is not evidence.
1. Methodology
“Can you share your methodology document?”
Good answer: a detailed written methodology covering manual testing, attack path analysis, and business context, with a clear statement of where automation is used.
2. Tester qualifications
“Who will be assigned to my engagement, and what are their qualifications?”
Good answer: named testers, relevant certifications such as OSCP, OSWE, or CREST, and years of hands-on work on similar systems. A claim about a team of experts is not an answer.
3. Reporting
“Can you share a redacted report sample?”
Good answer: immediate willingness to share, and a sample that contains an executive summary, business context, proof-of-concept steps, and remediation guidance.
4. Retesting
“Is retesting included in the price, and how quickly can it be scheduled?”
Good answer: a written policy, a defined window, and a clear statement of cost if retesting is extra.
5. Scope
“How do you define scope, and what happens if you find something outside it?”
Good answer: a collaborative scoping process, written boundaries, and a defined path for scope expansion rather than silent omission.
6. Data handling
“How do you handle our data during and after the engagement?”
Good answer: storage location, access controls, NDA terms, retention, destruction, and GDPR processing explained without hedging.
7. Post-engagement support
“What happens after the report is delivered?”
Good answer: remediation walkthrough, a channel to the testers for clarification, and a stated retest and revised-report process.
8. Stack and sector fit
“Have you tested systems like ours, and can the assigned testers show that experience?”
Good answer: specific technology or industry examples from the people who will do the work, not a company-wide claim that someone somewhere has seen a similar stack.
How to Prepare for Your Pentest: A Pre-Engagement Checklist
Prepare for a pentest by lining up stakeholders, documentation, attack-surface boundaries, environments, communications, and rules of engagement before the testing window starts.
Start vendor evaluation early. Capable testers book ahead, and a proper scoping call takes calendar time. Internal delay is as common as vendor delay. Do not wait until the audit date is immovable and then accept whatever slot is left.
- Identify key stakeholders. Include security, IT, application owners, compliance, and legal. Name a single internal owner for vendor questions during the test.
- Gather documentation. Prepare network diagrams, application architecture, API documentation, user roles and privileges, and previous test reports where they exist.
- Define the attack surface. List systems, applications, APIs, and networks in scope, and write explicit out-of-scope items. Ambiguity here becomes missed coverage later.
- Prepare test environments. Confirm whether staging is representative, whether production will be tested, and whether there are availability or change-freeze constraints. Unrepresentative staging produces findings you cannot trust.
- Provision access. Ready test accounts, MFA exceptions if agreed, VPN or jump-host access, and a named contact who can unblock testers the same day.
- Set expectations with the team. Decide who is told about the test, how findings will be escalated, and how blue-team or SOC alerts will be handled so testers are not treated as live attackers by accident.
- Clarify rules of engagement. State what is permitted around social engineering, denial-of-service testing, data access, and out-of-hours work.
- Schedule the scoping call. Take the documentation, in-scope list, constraints, and compliance driver into that call so the proposal matches the real surface.
What Happens After the Report? Post-Engagement Support and Retesting
Post-engagement support is the work after delivery: remediation guidance, tester access for questions, retesting of fixes, and a revised report where agreed.
The report is the start of remediation, not the end of the engagement. A vendor’s willingness to explain findings and retest fixes is a quality signal that sales decks do not show.
Engineers will have questions about exploit conditions, false-positive risk, and fix order. If the testers disappear after PDF delivery, you will spend longer arguing about the findings than fixing them.
A pentest is a point-in-time assessment. It identifies weaknesses at a specific moment. It does not guarantee that the estate stays secure, and it does not replace patching, design review, or ongoing testing.
Clarify retesting terms before you sign. Some vendors include a retest in the original price. Others charge extra and queue it behind new projects. Confirm whether the retest validates the specific fix or only re-runs a scan against the same host.
Ask before award:
- Does the report include actionable remediation steps, and is there a debrief call?
- Is retesting included, how is it scoped, and how quickly can it be scheduled?
- What is the process for confirming that a fix actually closed the issue?
- Can engineers contact the testers through a defined channel?
- Will you issue a revised report after retesting?
Frequently Asked Questions
What is the difference between a pentest and a vulnerability scan?
A vulnerability scan is an automated check for known weaknesses. A penetration test is a manual, expert-driven assessment that attempts exploitation to show real-world impact, attack paths, and business risk.
If there is no attempt to exploit or chain issues, it is not a pentest.
How long does a pentest take?
Duration is set by scope and attack surface.
A focused web application test may need only a few days of testing plus reporting time, while internal, external, and hybrid scopes often run to several weeks from scoping call through report and debrief.
Add vendor lead time and retest time when you plan an audit date.
What should a pentest report include?
A usable report includes:
- an executive summary
- methodology
- findings with CVSS scores and business context
- proof-of-concept steps
- remediation guidance
- a clear risk rating
If those elements are missing, the report will not support either engineers or auditors.
How do I know if a vendor is overpromising?
Overpromising shows up as:
- unnamed testers
- no redacted report sample
- a methodology that is only a scanner workflow
- a price that only works if large parts of the surface are excluded
Ask for names, a sample report, written retesting terms, and a list of exclusions. Refusal or delay on the sample is enough reason to walk away.
What is the difference between external, internal, and web application pentests?
An external test assesses internet-facing systems as an unauthenticated or lightly authenticated outsider.
An internal test assumes a foothold inside the network and looks at lateral movement and privilege escalation.
A web application test targets application logic, authentication, session handling, and authorisation in the app itself.
Many programmes need more than one of these, mapped to the actual asset, rather than a single label that sounds complete.
Should I choose a specialist pentesting firm or a generalist security consultancy?
Choose a specialist pentesting firm when the outcome you need is deep, manual testing and a high-quality exploit-led report.
Choose a generalist security consultancy when pentesting is only one workstream inside a wider programme and you accept that testing depth may be thinner.
For a high-assurance test of a critical application or network, specialist delivery is the safer default.
Next Steps: Turn Your Evaluation into a Vendor Shortlist
Turn this evaluation into a shortlist by locking scope, provider type, methodology evidence, and pricing model before you request quotes.
- Define the attack surface, test types, exclusions, and primary driver (compliance, risk reduction, maturity, or incident response).
- Select the provider type that fits that driver: consultancy, PtaaS, or crowdsourced.
- Assess methodology and reporting using a written method document and the Report Quality Scorecard on a redacted sample.
- Compare pricing models on inclusions and exclusions, especially retesting, not on headline fee alone.
- Request quotes from the shortlist and compare proposals against the same questions in this guide.
A thorough evaluation is the practical way to avoid paying for a scan labelled as a pentest. When the research is done, compare live service options and ask for a proposal that you can score against the criteria above.
Ready to shortlist a pentesting vendor?
Explore our pentesting services or request a quote once your scope and evaluation notes are ready. TEST. FIND. FIX. PROTECT.