Dunicot A cybersecurity consultancy and advisory firm.

Methodology · Reporting · 7 min read

The eight checks a finding should survive before you report it

Every finding that reaches a client has cleared eight checks first, because a report’s credibility is spent by its first false positive.

The first false positive in a penetration test report costs your client an afternoon. The third costs the report its authority, and once engineers learn to discount a document, the real findings in it stop getting fixed too. That is the actual damage, and it is self-inflicted.

Every medium-severity finding and above in our reports passes eight checks before it is written down. A finding that fails any one of them gets investigated further or dropped. It never gets reported with a hedge like “may be exploitable” or “appears to allow”, because those phrases are how a tester transfers their own uncertainty onto the reader.

1. Fresh reproduction

Reproduce it from a clean session with no prior state: new browser profile, new login, nothing cached and nothing left over from the experiment that found it.

A surprising number of apparent findings are artefacts of the tester’s own accumulated session state: a token from an earlier step, a cookie set by a previous test, a cached response. If it only works in the session where you found it, you have found something about your session, not about the application.

2. Confirm from the victim’s side

For anything involving another user (IDOR, cross-tenant access, privilege escalation, stored XSS), confirm the impact from the second account’s own session, not by inference from the attacker’s response.

An endpoint returning HTTP 200 and a body is not proof that the victim’s data was exposed. It might be an empty template, a placeholder object, or the attacker’s own record echoed back. Log in as the victim and verify what changed or leaked.

3. Isolate the parameter

Change one thing. If the payload alters three values at once, you do not yet know which one causes the behaviour, and neither will the engineer who has to fix it.

Parameter isolation is also what turns a finding into an actionable one: “this endpoint is vulnerable” sends a team hunting; “this endpoint trusts the tenant_id in the body instead of the session” is a fix.

4. Verify tool output by hand

Scanner and automated tooling output is a lead, never a finding. Every candidate gets manually reproduced before it is written down.

This is the single biggest quality difference between reports in this market. A tool that reports a SQL injection based on a timing differential has observed a timing differential, which is also what a slow query, a cold cache, a retry, or a shared-tenancy neighbour produces.

5. Rule out environment noise

Reproduce at least three times. For anything timing-based (race conditions, time-based injection, resource exhaustion), time it three times and compare the distribution, not a single sample.

Staging environments are noisy: cold starts, autoscaling, shared databases, background jobs. A 4-second delay that appears once is noise. A 4-second delay that appears on 3 of 3 attempts with a payload and 0 of 3 without one is a finding.

6. Demonstrate real impact

Read actual data, perform an actual action, or obtain an actual session. “An attacker could potentially” is not impact; it is a hypothesis about impact.

This is also where severity gets honest. Plenty of technically valid findings turn out to reach nothing of value once you follow them, and they should be reported as what they are rather than inflated to justify the engagement.

7. Rule out the proxy

Confirm the response came from the application rather than from a WAF, a CDN, a load balancer or an error page. A 403 from a WAF and a 403 from the authorisation layer look identical in a proxy history and mean opposite things.

The quickest discriminator is usually a differential: send the same payload unauthenticated and authenticated, or with a benign variation, and see whether the boundary tracks the application’s logic or the edge’s ruleset.

8. Check it is not a duplicate

Check the finding against what has already been reported for that asset: by you, by a previous engagement, or in the client’s own tracker if you have access.

Re-reporting a known issue as new is a small thing that reads as carelessness, and it makes the real findings around it harder to trust.

Why publish the gate

Two reasons. First, it is a claim you can hold us to: if a finding in one of our reports fails one of these, that is a defect we want to hear about.

Second, it is portable. Nothing here is proprietary: any team can apply the same eight checks to their own internal findings, their bug bounty triage, or the reports they receive from other vendors. Applying check four alone to an incoming vendor report tends to be educational.

In short

Point 1
A report’s credibility is spent by its first false positive, not earned by its finding count.
Point 2
Tool output is a lead; only manual reproduction makes it a finding.
Point 3
Impact means data read or an action performed, not “could potentially”.
Point 4
Timing findings need three timed attempts, not one.

Want this applied to your stack?

Everything written here comes out of delivered engagements. Describe the platform and the deadline.