Prevent Private URLs From Leaking Into Static Site Output
Stop internal hostnames, staging URLs, and private network details from leaking into static-site HTML, JavaScript, feeds, source maps, and metadata.
Table of Contents9 sections

Static site generators make deployment simple because the final output is just files. That simplicity can hide a dangerous assumption: if a value is only used during the build, it cannot reach the browser.
It can.
A private hostname, staging API URL, internal bucket name, or localhost callback can leak into generated HTML, JavaScript, source maps, feeds, or metadata even when the original environment variable never appears in source control. The right defense is not “keep .env out of Git.” It is to treat the generated dist/ directory as a security boundary and verify it before deployment.
Why build-time secrets and private URLs are different problems
A secret is valuable because someone can use it. A private URL may contain no credential at all, yet still reveal infrastructure names, internal routing, non-public environments, or assumptions about your network.
That distinction matters because secret scanners are optimized for tokens, keys, and recognizable credential formats. They may correctly report a clean build while https://staging-api.internal.example is sitting in a rendered page.
Static generation also changes where the leak happens. Consider a build step that reads an environment variable:
const apiBase = import.meta.env.API_BASE_URL;
If that value is interpolated into a component, serialized into client state, written into a JSON payload, or included in a canonical URL, it becomes public output. Removing the environment variable after the build does nothing. The bytes have already been generated.
Define what must never appear in public output
The strongest guard starts with an explicit denylist of private patterns. Do not rely on developers remembering every place a URL might be rendered.
Useful categories include:
- localhost and loopback addresses;
- private DNS suffixes such as
.internalor.local; - staging and development hostnames that are not intended for readers;
- RFC 1918 IPv4 ranges when they should never be exposed;
- internal service names used by containers or private networks;
- private repository, vault, or filesystem paths.
Keep this list focused. A rule that blocks every occurrence of words like dev or test will create enough false positives that people eventually stop trusting it.
For a public technical site, the policy should answer a simple question: “Could a stranger learn something about a non-public system from these generated files?”
Scan the artifact, not only the source tree
Source scanning is useful, but the deployable artifact is the final truth.
A minimal Node.js verifier can recursively inspect generated text files and fail the build when it finds forbidden patterns:
import { readFile, readdir } from 'node:fs/promises';
import path from 'node:path';
const blocked = [
/localhost(?::\d+)?/i,
/127\.0\.0\.1/,
/https?:\/\/[^\s"'<>]+\.internal\b/i,
/https?:\/\/10\.\d+\.\d+\.\d+/,
];
async function walk(dir) {
for (const entry of await readdir(dir, { withFileTypes: true })) {
const file = path.join(dir, entry.name);
if (entry.isDirectory()) await walk(file);
else if (/\.(html|js|css|json|xml|txt|map)$/i.test(file)) {
const content = await readFile(file, 'utf8');
for (const rule of blocked) {
if (rule.test(content)) {
throw new Error(`Private output detected in ${file}: ${rule}`);
}
}
}
}
}
await walk('dist');
The important design choice is where this command runs. Put it after the production build and before the deployment step. If the scan fails, nothing should be promoted.
This pattern is similar to other fail-closed delivery checks. The deployment pipeline should prove that the artifact is acceptable rather than assuming the build command produced something safe.
Watch the places developers forget to inspect
Generated HTML is only one output surface. Private values often escape through less obvious files.
Source maps can contain original source strings. RSS or Atom feeds can inherit canonical URLs from site configuration. Open Graph metadata can expose an asset host. JSON-LD can serialize configuration into structured data. A generated search index can copy full page text. Client-side hydration payloads can include server-side values that looked harmless during development.
Even error pages deserve inspection. A fallback page generated with a development base URL is still public if it ships in dist/.
The practical rule is to scan every text-based artifact that can be deployed, not only index.html.
Separate public configuration from server-only configuration
A verifier catches mistakes, but configuration design should make those mistakes harder to create.
Use separate namespaces for values that are intentionally public and values that must remain build-only or server-only. Frameworks often provide a convention for public environment variables; treat that prefix as an explicit publication decision rather than a convenient way to make configuration available everywhere.
For example:
PUBLIC_SITE_ORIGIN=https://example.com
INTERNAL_PREVIEW_ORIGIN=https://preview.internal.example
Only the first value should ever be imported by browser-facing modules. The second can be used by a private validation step, but it should never cross the rendering boundary.
If a component needs data from a private service, fetch that data during the build and render the safe result. Do not render the private service address into the page.
Make the check deterministic and explainable
A security guard that occasionally fails without telling developers why will be bypassed.
When the verifier finds a match, report:
- the file that contains it;
- the rule that matched;
- a short, redacted preview if that is safe;
- the remediation path.
Avoid printing a complete matched credential or sensitive URL into CI logs. The checker should reveal enough context to fix the build without turning the log into a second leak.
It also helps to version the rules in the repository. A change to the denylist then becomes reviewable code, not an undocumented setting hidden in a deployment dashboard.
Test the guard with a deliberate failure
A verifier is only useful if you know it can stop a bad artifact.
Add a fixture or test that writes a known private hostname into a temporary build directory and asserts that verification fails. Then test a normal public URL and assert that it passes.
This catches a surprisingly common failure mode: the script exists, CI reports it as successful, but the command is scanning the wrong directory or file extensions.
For the same reason, run the verifier against the exact directory that the hosting provider deploys. If Cloudflare Pages, Netlify, or another platform publishes dist/, scanning a different staging folder gives false confidence.
Treat deployment as a proof chain
A reliable static-site release can be modeled as a short proof chain:
source
-> validated content
-> production build
-> generated artifact scan
-> deployment
-> live smoke test
Each arrow should preserve an invariant. Content validation proves the source contract. The build produces deterministic output. The artifact scan proves that private patterns are absent. Deployment promotes those exact bytes. A live smoke test proves the expected public page is reachable.
This is stronger than adding another secret scanner because it verifies the boundary that actually matters: what users can download.
The practical takeaway
Keeping .env files out of Git is necessary, but it is not sufficient for static sites. Build-time configuration can become public whenever it is rendered, serialized, indexed, or copied into deployment artifacts.
Define the private patterns your site must never expose, scan the complete generated output, fail closed before deployment, and test the guard with a known bad fixture. Once dist/ becomes an explicit security boundary, private URL leaks stop being an accidental side effect of static generation and become a condition your pipeline can prove false.
Continue Exploring
You Might Also Like

Subdomain Routing with Cloudflare Pages Middleware
Learn how to host branded subdomains on a single static site deployment using Cloudflare Pages edge middleware to prevent duplicate content penalties and handle clean routing.

Safe Multi-Environment Database Orchestration
Learn how to manage PostgreSQL databases across local, staging, and production tiers securely without schema drift or credential leaks.

Post-Deploy Sanity Checks: Verify Production Without Re-Running Your Test Suite
A practical guide to designing small post-deploy sanity checks that verify the live release, critical dependencies, and rollback signals without duplicating CI.