Dunicot A cybersecurity consultancy and advisory firm.

Attack scenario · Government and public sector

Dependency confusion: how one package name becomes code in your build

Dependency confusion lets an attacker register your internal package name on npm or PyPI, so the next build installs theirs and runs it on the build agent.

Critical (conditional) Dependency confusion (public-registry package substitution)

At a glance

Class
Dependency substitution, commonly called dependency confusion
Severity
Critical where resolver preference, install-time or build-time code execution, and a build identity that can sign, publish or reach production are all observed. High where execution is confirmed but signing is isolated and credentials are job-scoped and short-lived. Medium where only an unclaimed internal name and a merging or unmapped resolver configuration are proven, with no demonstrated execution path.
OWASP mapping
OWASP Top 10 A08:2021 Software and Data Integrity Failures; OWASP Top 10 CI/CD Security Risks CICD-SEC-3 Dependency Chain Abuse
CWE
CWE-427 uncontrolled search path element for the index search order variant, with CWE-1357 reliance on insufficiently trustworthy component for the trust decision and CWE-494 download of code without integrity check where no lockfile hash or origin pinning is enforced
Reachable from
A free public registry account plus a build that resolves the name from a public index. No access to the organisation is needed, but the organisation's resolver configuration decides the outcome.
Who is affected
Confirmed exposure: build agents and the credentials readable from them. Conditional exposure: signed release artefacts and downstream consumers, where the compromised stage can reach signing or publishing.
Usual trigger
An internal package name visible in a public build log, lockfile or JavaScript bundle

What it is

Dependency confusion, also called dependency substitution, is a supply chain attack in which someone publishes a package to a public registry under the name of one of your internal packages, and your build installs theirs instead of yours. The substitution happens inside the dependency resolver, and it takes one of three forms. Where a private index and a public index are both consulted and their results merged, as with pip and an extra index, NuGet without package source mapping, or a proxying registry that merges upstream and local answers, the resolver gathers candidates from both and installs the highest compatible version, so a public package published at an arbitrarily higher number than the internal library wins. Where npm, yarn or pnpm resolve, nothing is merged: each name is requested from exactly one registry, so the failure is a routing failure, an unscoped or unmapped internal name fetched from the public default, and any version satisfying the range is enough. Where Maven or Gradle resolve, repositories are tried in declared order and the first that can serve the coordinates answers, so ordering and missing content filtering decide.

Winning that resolution runs code only where installation can execute. Node lifecycle hooks run where a package declares them and scripts are not disabled, and a Python source distribution runs code while it is built; with wheels only, or install scripts disabled, the substituted package waits to be imported by the build or test stage instead. Where execution is available, the code lands in an expensive place. A build agent is normally the machine that legitimately holds registry publish tokens, cloud deployment roles, signing material and the source of the next release, so a single install converts a free public registry account into execution inside the release process.

In a government or public sector setting the blast radius widens, because internal platform libraries are deliberately shared: one common authentication or logging package is consumed by several departments, by outsourced delivery teams, and by citizen-facing services the attacker never has to touch directly. The organisation is not defending a dependency here, it is defending the integrity of everything it signs.

How the attack unfolds

Mechanism, not a recipe. Reproduction detail stays in client reports, because a page that hands a reader a working attack is a liability.

Where internal package names leak: build logs, lockfiles and JS bundles

Dependency confusion starts with a name, not a secret. Verbose installer output printed into a public build log, a lockfile or requirements file committed to a public repository, a published package whose manifest lists internal dependencies, a JavaScript bundle or source map containing internal module paths, and layers of a publicly pushed container image all carry dependency names in plain text. Each of those tells an attacker which names your builds will ask for.

The name is claimed on the public registry

Internal names are normally unregistered publicly, so the name goes to whoever asks for it first, and registration needs no relationship with the organisation and leaves no trace in its own systems. What the attacker publishes depends on which resolution failure they are aiming at. Against a merged-index resolver they publish at an arbitrarily high version, and that decides the outcome only where the dependency is declared with an open or wide range, is newly added, or is resolved without a lockfile.

Against npm routing the version is irrelevant, because the contest is over which registry is asked, so any version satisfying the range suffices. Where the dependency is pinned or tightly bounded, a high version is not a compatible candidate and loses outright.

Why the resolver picks the public package

On the next build that refreshes dependencies, the behaviour follows the package manager. pip has no concept of index priority, so with an extra index configured it collects candidates from every index and installs the highest compatible version regardless of origin, and NuGet behaves the same way until package source mapping is configured. A private registry that proxies an upstream and merges the candidate lists reproduces that same contest while appearing isolated.

npm, yarn and pnpm ask one registry per name, so the exposure is an internal name that sits outside any scope mapped to the private registry and is therefore fetched from the public default. Maven and Gradle accept the first repository in declared order that can serve the coordinates. In each case the resolver holds no instruction that this particular name belongs to one index only.

Installation executes code on the build agent

Package installation is not always a file copy. Node lifecycle hooks run where the package declares them and scripts have not been disabled, and a Python source distribution runs code while it is built, both with the privileges of the account the pipeline runs as. Where that is the case, the attacker has execution inside the build environment before any test, scan or review stage has looked at the artefact.

Where installation cannot execute, because only wheels are installed or install scripts are off, the substituted code waits until the build or test stage imports it.

Install-time code reads pipeline secrets and rewrites the build output

The runner's environment, mounted secrets, federated cloud identity token and registry publish credentials are all readable from the attacker's install script, which is now running as the pipeline account. Credential theft is not the worst outcome. Alteration of the build output is, and reaching it takes one more step: because the hook runs during dependency installation, before anything is compiled, it has to plant persistence inside the job first, by patching a build tool or a sibling dependency in the workspace, injecting a configuration or plugin hook, or placing a wrapper earlier on PATH.

That persistence then executes during the build stage and modifies the bundle, or simply rewrites the dependency tree so the compromise reaches the output, while the sources in the repository stay unchanged and the change is invisible to anyone reading the code.

The backdoored build is signed and distributed downstream

Where signing and publishing run in the same job, or in a stage the build stage can influence, a build altered by a substituted dependency is still signed by the organisation's own legitimate identity, because the alteration happens before the signing and publishing stages. Downstream departments, suppliers and citizen-facing services that check the signature alone will accept it, and internal libraries fan out to many consumers from a single release. Provenance does not conceal the attacker: attestation produced by a trusted builder outside the job records the resolved materials and their digests, so the substituted package appears in the attestation and an unexpected name or digest is detectable.

That is why material-level verification downstream catches this and signature-only verification does not.

Business impact

The immediate loss in a dependency confusion compromise is credential material with reach far beyond the build. A pipeline identity typically holds a cloud deployment role, object storage access, database migration rights and a publish token for the private registry, which means a single install can read production configuration, reach datastores holding citizen records, and push a further malicious version of any internal package the token covers. Where builds use short-lived federated tokens the window is narrower but still sufficient, because the attacker is executing while the token is valid. This is the point where a dependency problem stops being a development concern and becomes a live data exposure.

Recovery costs more than the compromise. Once execution inside the build process is established, every artefact produced after the first malicious install is suspect, and nobody can say which release that was without build logs detailed enough to show when resolution changed. Remediation therefore means rotating every secret the runner could read, rebuilding and re-signing releases from verified sources, and re-verifying deployments across each downstream consumer, while the service itself may need to be held on a known-good older version. For public sector platforms with no alternative channel, holding a citizen-facing service on an old release is itself a service failure, and supplier contracts rarely say who pays for the re-verification of a shared internal library.

The control that failed here is integrity, and integrity is what assurance frameworks assume rather than test. Public administration entities in the EU fall within the NIS2 reporting regime, subject to how each member state has transposed it, and there the early warning for a significant incident is due within 24 hours of the entity becoming aware of it, with data protection obligations attaching separately if personal data was reachable from the pipeline. Public procurement commonly asks suppliers for software bills of materials and build provenance, and a compromise of this shape satisfies both on paper unless the recorded materials in the attestation are actually verified, because the malicious code arrives inside a correctly signed artefact produced by the organisation's own pipeline. An organisation that cannot state which builds are clean cannot answer an auditor, a regulator or a downstream department, and that inability tends to outlast the technical cleanup.

How it is found

What a tester looks for
SignalHow it is confirmed
Internal dependency names that are unclaimed and claimableWe inventory every dependency name from manifests, lockfiles, vendored directories, JavaScript bundles and container image layers, then record three independent facts for each name: that it is absent from the public index, that it is claimable under that registry's naming and namespace rules, and that it is reachable from the indexes the build is actually configured with. A 404 is not proof of claimability. npm refuses names too similar to existing packages and a name inside a scope cannot be claimed while the organisation owns the scope, PyPI normalises names so that foo.bar, foo_bar and foo-bar are one name and refuses names under its own naming policy, and a name on a registry the build never queries is irrelevant either way. We establish claimability from registry policy and namespace ownership, never inferred from a 404 and never by test-publishing. Vulnerability scanners do not settle this class, because a freshly published attacker package has no advisory filed against it; the deciding property is name ownership, not version age.
Resolver configuration that merges public and private indexesWe read the resolver configuration rather than assume it: npm scope-to-registry mappings and default registry fallthrough, pip index and extra-index settings, NuGet package source mapping, and Maven or Gradle repository ordering and content filters. Secure code review covers the committed pipeline and project configuration, which is where effective CI behaviour is set, and it covers each lockfile entry's resolved origin as well as its integrity hash, because a lockfile generated on a developer machine that resolved the name from a public index is internally consistent, carries valid hashes for the attacker's artefact, and will be installed faithfully by CI.
Where the names are escaping toWe search public repositories, published package manifests, job logs and artefacts, source maps and pushed image layers for internal package names and private registry hostnames. This separates theoretical exposure from a name an attacker can already read, and it shows which leak source needs closing as well as which name needs claiming.
Whether a substituted package would execute on install and reach the internetWe prove resolver preference against a controlled stand-in: we stand up a local index that serves the internal name at a higher version, place it in the configuration slot the public index occupies in the real pipeline configuration, and observe which index is queried and served using resolver verbose output and the egress log. We then examine the agent for whether install scripts run by default, whether source distributions are built rather than wheels installed, and whether outbound access is unrestricted. Confirmation is the observed resolution and the observed egress path, not the configuration file alone. Testing for this class never involves publishing to a public registry under a name the client does not own.
Whether a substitution has already happenedWe check whether any inventoried internal name is already registered publicly, by whom, and when relative to build history, diff the resolved origins recorded in existing lockfiles and build logs against the private index, review runner egress and private registry proxy logs for upstream fetches of internal names, and rebuild from verified sources to compare digests against the artefacts currently published. An internal name already held by an unrelated account is treated as a suspected compromise, not as a precondition.
The privilege the build agent would hand overWe enumerate the runner's readable environment, mounted secrets, federated cloud role and publish token scope to measure what one install would obtain. Cloud penetration testing follows that identity to its limits, establishing whether the pipeline role can reach production data, alter infrastructure, or sign and publish on the organisation's behalf.

How it is fixed

Controls that hold
ControlWhat makes it hold
Claim the names and namespaces now, then fix the resolutionPrevention rests on making private names unresolvable from public indexes, and the fastest move toward that is defensive registration. Register your organisation's scope or namespace on every public registry your builds resolve from, publish placeholder packages for the internal names already exposed, and monitor those registries for new registrations matching your name patterns. This closes the window in hours rather than sprints, and it is a stopgap: the structural controls below are what keep holding as new packages are added.
Make private names unresolvable from public indexes, per ecosystemThe control differs by package manager, and treating one version of it as universal leaves whole estates untouched. For npm, place internal packages inside an organisation scope, register that scope publicly under your own account so a client missing the mapping cannot resolve it from whoever else claimed it, and map the scope exclusively to the private registry. For NuGet, combine prefix reservation with package source mapping. For Maven and Gradle, use a dedicated groupId with repository content filtering. For pip and PyPI there is no namespace ownership and no per-package index pinning, because index-url and extra-index-url are global, so for Python the single authoritative index is not a secondary measure, it is the primary control.
Treat the private index as authoritative, not additionalUse a single index with an explicit allowlist of upstream packages instead of adding a public index alongside the private one, and where the tooling supports origin pinning, such as package source mapping, state which source is permitted to serve which pattern. A proxying registry must refuse upstream resolution for the organisation's entire name pattern, prefix or groupId, including names that do not exist privately yet, because existence-based exclusion leaves the new or not-yet-pushed internal package exposed and that is precisely the name an attacker claims. The build should fail closed when a name matching that pattern cannot be served by the private index.
Pin resolution and verify integrity on every buildCommit lockfiles with integrity hashes and install through the strict path that refuses to update them, and use hash-checking mode for Python, which removes version arbitrage as a mechanism. Two conditions carry this control: the pipeline must fail closed on lockfile drift, and dependency changes must pass human review, since a lockfile that CI quietly regenerates protects nothing. Integrity hashes pin content and not provenance, so distribute resolver configuration, meaning .npmrc, pip.conf and nuget.config, to developer machines as a managed artefact, and check each lockfile entry's resolved origin in CI.
Remove install-time code execution from the buildDisable install scripts by default, prefer prebuilt distributions over source distributions, and allowlist the small set of packages that genuinely need a build step. This does not stop substitution, but it means a substituted package has to wait to be imported by code under review rather than executing the moment it is fetched.
Reduce and separate what a build can doKeep runners ephemeral, scope credentials to the single job with a short lifetime, and sign in a separate stage under a separate identity the build stage cannot reach. Restricting outbound traffic raises the cost of exfiltration without preventing it, because the public registry the build legitimately needs is itself a serviceable channel, so the real reduction here is credential scope and lifetime plus signing isolation. Generate provenance attestation in a stage the build cannot influence, and enforce verification at the consumer including the resolved materials, since unverified provenance is an audit artefact rather than a control.

Questions

What is a dependency confusion attack?

Dependency confusion, also called dependency substitution, is a supply chain attack where someone publishes a package to a public registry using the name of one of your internal packages, and your build installs theirs instead of yours. The resolver makes the substitution, either by preferring a higher version across merged indexes or by requesting an unscoped internal name from the public registry, and where installation can run code, through npm lifecycle scripts or a Python source distribution that has to be built, that package then executes with the build agent's privileges.

Is dependency confusion the same thing as typosquatting?

No. Typosquatting needs a developer to misspell a real public package name, so it depends on human error. Dependency confusion needs no error at all: the name in your manifest is correct, and the resolver itself reaches the public package because nothing in the configuration says that this particular name must only ever be served by your private registry.

Does using a private registry protect us from dependency confusion?

Not on its own. What matters is whether the private index is authoritative for your names or merely one more place the resolver looks. If both indexes can answer for a name, the highest version wins, and a private registry that proxies an upstream and merges the candidate lists produces exactly that situation while appearing to be isolated.

Why does pip choose the public package over our internal index?

Because pip has no concept of index priority. With an extra index configured it collects candidate versions from every index and installs the highest compatible version regardless of which index supplied it, so an internal package at its real version number loses to a public package of the same name published at an arbitrarily higher one, with no warning that the origin changed. Pinned requirements with hash checking remove that arbitrage, and a single authoritative index with upstream allowlisting removes the second index entirely.

Should we register our internal package names on the public registry?

Yes, as a stopgap. Claiming placeholder packages under your internal names on each public registry you resolve from removes the free name an attacker needs, and it is worth doing straight away for names that are already exposed. It is not the fix: scoping internal packages to a namespace only your private registry serves is what keeps holding as new packages are added.

Can dependency confusion be tested without publishing a package to npm or PyPI?

Yes, and publishing is not what testing requires. Dependency confusion testing is a resolution and exposure exercise: we inventory internal dependency names from manifests, lockfiles, bundles and image layers, check whether each one is both unclaimed and claimable publicly, then reproduce resolution in an isolated container against a controlled local index to show which index actually wins under your real configuration. We never publish to a public registry under a name you do not own.

Which package managers are affected by dependency confusion?

All of the major ones, by different mechanisms. pip gathers candidates from every configured index and installs the highest compatible version, so an extra index settles nothing about origin. NuGet behaves the same way until package source mapping is configured. npm, yarn and pnpm ask a single registry per name, so the exposure there is an internal name sitting outside any scope mapped to your private registry. Maven and Gradle take the first repository in declared order that can serve the coordinates, which makes ordering and content filtering the control. Go is the outlier, because a module path embeds a domain you control, and that removes the free name this class depends on.

Services: where this is tested

Test for this on your stack

Describe the platform and the deadline. A fixed quote follows a short scoping call.