Dunicot A cybersecurity consultancy and advisory firm.

API security · Data protection · 6 min read

Your API returns more than your interface shows

The interface shows a name. The response contains the record. Nothing is broken, nothing alerts, and the data is gone.

Open the network tab on almost any product and look at what comes back from the endpoint behind a user list. The page renders a name and an avatar. The response frequently contains the email address, the phone number, the internal identifier, the role, the account status, the last login timestamp, the password-reset state and a few fields whose purpose nobody currently remembers.

Nothing is broken. No control failed. The endpoint was asked for a user and it returned a user: the whole object, because serialising the whole object is what the framework does by default and nobody wrote the line that stops it.

Why this is worse than it looks

It is silent. There is no exploit to detect, no anomaly to alert on, and no payload in a log. Retrieving it is indistinguishable from normal use, because it is normal use.

It scales. An endpoint that over-returns one record over-returns every record, and anything with pagination hands over the data set at whatever rate the API allows.

It defeats the interface as a control. Teams reason about exposure in terms of what the product displays, and the product displays a name. The regulator, meanwhile, reasons about what was accessible.

Where it hides

The obvious list endpoints are usually the least interesting, because someone has looked at them. The exposure sits in the paths nobody demos.

  • Search and autocomplete, typeahead endpoints that return full objects so the client can render one field.
  • Exports, CSV and PDF generation that runs on a different code path from the request that produced the view.
  • Error responses, validation failures that echo the submitted or conflicting record back.
  • Embedded relations: an order that includes the customer, which includes the account, which includes the owner.
  • Webhooks, event payloads sent to third-party integrations, carrying far more than the event needed.
  • GraphQL, where the client chooses the fields, so field-level authorisation is the only control there is.
  • Mobile APIs, built for one client, so nobody expected the response to be read by a person.
  • Audit and activity feeds, which by design describe what other people did.

How to find it in your own product

You do not need a penetration test to start. Take the ten most-used endpoints, call each with a normal account, and diff the response against what the interface renders. Every field in the gap is a decision nobody made.

Then ask, per field: does the client need this to render? If no, it should not be in the response. That question resolves the overwhelming majority of cases, and it is faster than any tooling.

The fix, and the anti-fix

The fix is explicit output contracts: a response schema per endpoint listing the fields it returns, so adding a column to a table never silently adds a field to an API. Serialisers, view models, DTOs, GraphQL field resolvers with authorisation. The mechanism matters less than the default being deny.

The anti-fix is filtering in the client. Hiding a field in the interface changes nothing about what crossed the network, and it produces the exact false confidence this class depends on.

Under GDPR Article 25 this is data protection by design in the most literal sense available: the default is what determines whether the data was exposed, and the default is currently “return everything”.

In short

Point 1
Diff every API response against what the interface renders; the gap is unmanaged exposure.
Point 2
Exports, search, errors, webhooks and mobile endpoints are where it hides.
Point 3
Fix with explicit per-endpoint output schemas, deny by default.
Point 4
Filtering in the client changes nothing; the data already crossed the network.

Want this applied to your stack?

Everything written here comes out of delivered engagements. Describe the platform and the deadline.