Security headers on 4,688 small-business websites: 49.7% met none of 7 criteria

Hacker News by 30 min read 14x views
Security headers on 4,688 small-business websites: 49.7% met none of 7 criteria

Share Post

A 2026 study of 7,040 directory-listed U.S. local-business websites: which safety headers they send, and which they get wrong.

By RACKCRUNCH Team · Scans 2026-09-24 · Also accessible as a PDF (security-headers-study-2026.pdf, 400 KB, SHA-256 c8fa6a281799ff1f86c79b62f5ba5b52535a4ca9d9b6bdce69d61f5bfaeb6d91).

Why we looked

RACKCRUNCH has a liberated Security Header Check. You provision it a URL, it makes one request, and it tells you which reply headers are set, which are set badly, and which are missing.

Checking one location at a period made us inquisitive concerning the bigger picture. We didn't average banks or big tech. We meant the plumber, the law office, the car lot, the pizza place: businesses that may deficiency dedicated web-security staff.

We needed a catalog of those businesses, and there isn't a spotless one. So we used the closest item we could find: the community Curlie web directory, the human-edited successor to DMOZ. We drew 7,040 directory rows at random from its US local "Business and Economy" categories, as of the 2026-02-02 snapshot. They portray 7,022 distinctive first registrable domains. Think of it as an SMB-oriented directory sample. We did not inspect how big any of these businesses are. Some volition be larger than "small", and the directory tilts toward businesses established adequate to get listed. When this study says "sites", it method these directory-listed sites, not all small endeavor in the US.

On 2026-09-24 we ran two HTTPS request-chain scans of all sampled URL, following at most three redirects; all reported estimates arrive from the second scan. We peruse the reply headers and never the pages themselves.

We expected low adoption. That is approximately what we got. The surprises came from how low, and from one figure we had to obtain apart.

Which headers sites send

5,642 of the 7,040 sampled rows gave us a usable HTTPS response. 4,701 of those were HTTP-200 responses, and they came from 4,688 distinctive final registrable domains. Those 4,688 domains are our chief basis (called "dedup-200" in the data files): one HTTP-200 reply per final registrable domain, from the second scan. The definitions box at the top of this study spells out all base. Every fig in this study is on that basis unless it says otherwise.1

We started alongside the simplest question: which of seven study-defined explicit-header criteria does all location meet? These are criteria under our rubric, not a safety grade, and a location can fairly depart several of them out. Three need a term up front. A missing Referrer-Policy gets the browser's safe default, so we don't treat it as a failure; the item goes lone to sites that set a safe value themselves. Permissions-Policy is motionless experimental and unevenly supported throughout browsers, and our criterion counts an definitive declaration, not safety gained. Cross-Origin-Opener-Policy depends on context: it matters most for sites that open pop-ups or grip cross-origin windows, and it can interrupt sign-in and fee flows.

No definitive criterion in our rubric was near-universal.

HSTS was the most commonly observed header, current on 43.8% of sites. The most commonly passed criterion in our seven-item rubric was X-Content-Type-Options: nosniff, at 39.7%; lone 12.3% met our stronger HSTS criterion of at smallest one twelvemonth affirmative includeSubDomains. We call that study-defined powerful HSTS, and it is measured on the final reply presenter only. If a location redirects to www, includeSubDomains there covers hosts under www, not the naked domain or brother or sister hosts specified as shop. Of the 578 strong-HSTS passes, 260 (45.0%) came from a www final host, 316 from the naked domain and 2 from another subdomain. Thus, for nearly fractional of strong-HSTS passes, the observed guideline was on a www presenter and does not established safety of the naked domain. 31.2% dispatch a clickjacking header alongside a recognized restrictive value under our parser. 8.0% dispatch an definitive non-wildcard Permissions-Policy declaration, and 1.1% dispatch a recognized non-default Cross-Origin-Opener-Policy value.

86.6% of sites don't dispatch a Referrer-Policy at all, and as above, that is small bad than it sounds. When the header is missing, contemporary browsers autumn rear to strict-origin-when-cross-origin, which is a sensible default. The additional concerning cases are definitive feeble bequest values.

And one location in five gives item away. 20.5% [19.3-21.7] of sites matched our version-token rule: a numeral in any of Server, X-Powered-By, X-AspNet-Version or X-AspNetMvc-Version, specified as Apache 2.4.x, Microsoft-IIS/10.0 or nginx 1.x. For 27 sites, the lone leak was an ASP.NET type header. That can aid version-targeted reconnaissance, but it is a low-severity finding: hiding a type figure patches nothing.

Here is the complete array on the chief base. Intervals quantify sampling doubt under the stated example model, not example bias, parser misclassification, CDN variation, or selective phase attribution.

The one invalid Referrer-Policy is a location that sends "Referrer-Policy: yes".

Put the seven criteria together and the header figure follows. Among the distinctive final registrable domains that returned HTTP-200, 49.7% met none of seven study-defined explicit-header criteria [95% CI 48.3-51.2], or 2,331 of 4,688.2 This is an acceptance measure, not an evaluation of the proportion of insecure websites.

The identical measures on all usable response, including error and bot-challenge pages, are in the sensitivity analyses at the end. Such responses may indicate CDN, WAF, hosting-platform, or use error-page configurations fairly than the headers on the site's average homepage, which is why they are not the chief base.

The CSP figure that cut apart

In our archetypal draft, the CSP outcome looked nearly respectable. Roughly one location in six "passed". That seemed high. So in a revised inspection we stopped counting CSPs and started study them, alongside a parser implementing the documented subset of CSP Level 3 described in the rubric.

About one location in five, 21.2% [20.0-22.4], sends an enforced CSP. That appears fine until you appearance at what the policies say. Most are framing rules (frame-ancestors, which is anti-clickjacking, not anti-script) or upgrade-insecure-requests. The three most average exact policies were frame-ancestors 'self' (170 sites), the Shopify phase default block-all-mixed-content; frame-ancestors 'none'; upgrade-insecure-requests; (167), and naked upgrade-insecure-requests (130); together they document for 47.1% of the 992 sites sending an enforced Content-Security-Policy. None of the three has a manuscript directive.3

CSP is an crucial defense-in-depth authority against injected scripts. Here is how the policies interrupt downward by what they really do for scripts.

What undoes the rest? A guideline can autumn into additional than one row:

That leaves eight distinctive HTTP-200 domains, 0.17% [0.1-0.3], that passed our header-only script-CSP rule. This does not established that the guideline is unbypassable or that nonces are caller and correctly applied in leaf markup. We did not test nonce freshness, whether nonces equivalent the scripts in the page, or whether allowlisted manuscript hosts can be abused, for example through JSONP endpoints. We did peruse all eight policies in complete by hand to inspect the parser's study of the header. None of the eight sends script-src-attr, so inline event handlers autumn rear to the identical restricted script-src. Two are policies issued by a hosted location builder fairly than written by the location owner. One allows no scripts at all. Two motionless authorize subdomain-wildcard hosts as manuscript sources, which our rubric reports separately and does not figure against them.

Eight. Most observed CSPs consisted chiefly of framing or mixed-content directives.

We got one item wrong

A outline of this study had a nice instant in it: six sites out of 4,673 met all eight rubric criteria. We liked that line. It was wrong.

Our archetypal rubric gave CSP frame-ancestors credit twice, formerly as clickjacking safety and again as a passing CSP. The second inspection continue fixed that. It additionally made the CSP inspect necessitate a script-restricting guideline that passes our header-only rule, and it stopped counting no-op values akin COOP unsafe-none or a wildcard-only Permissions-Policy. Each of the six misplaced the double-counted point.

The archetypal outline stated six sites met all eight rubric criteria; the corrected rubric says one, and eight additional met seven. The one got there on a CSP that uses strict-dynamic alongside a nonce, which our before parser misread.

A few of those nine sites portion header patterns accordant alongside a reused configuration: one developer's or vendor's checklist applied to multiple sites. We are not naming any of them. Nine sites is too few to assistance a story, and headers change. Meeting all eight criteria additionally says nothing concerning the remainder of a site: it can motionless have grave use flaws.

The explicit-header acceptance index

Adding up the criteria gives a sole figure per site, which we call the explicit-header acceptance index: the seven criteria complete affirmative an eighth, type hygiene (no type in the application banner). It mixes dissimilar things, from broadly helpful hardening to context-dependent controls and a low-severity banner check, all valued equally. So treat it as a summary of adoption. The per-control array complete is the chief result.

Explicit-header acceptance index, criteria met out of 8

The median is 1 and the average is 1.80. The 0 row is smaller than the "none of seven" header since many sites encounter the eighth criterion, type hygiene, merely by not announcing their software. We haven't validated the indicator as a measure of how safe a location is.

Header-identified server and hosting labels

On the second scan we kept the complete Server header, not fair the version-leaking ones. Between Server tokens and platform-specific CSPs, we could nexus a server or hosting tag to 3,481 of the 4,701 HTTP-200 responses (74.0%). Our before labeling managed 15.0%. The biggest sole group, 1,357 sites (28.9%), is one we call Cloudflare-edge: the Server header says exactly "cloudflare". That tells you the reply came through Cloudflare's network. It says nothing concerning the server rearward it.

These numbers are descriptive. Labels lone be for servers that acknowledge themselves, so the tagged collection is selective, and nothing current tells you what would happen if a stated endeavor moved hosts. Some labels arrive from platform-specific CSPs, which are among the headers being measured, so the difference is partially built from its own outcomes. Each reply gets one tag under fixed rules, and the archetypal regulation that matches wins, so no location carries two labels. Builder and phase signatures arrive first, since they name whoever controls the configuration. Server tokens arrive after. Matching is case-sensitive, as the header was sent. The command is GoDaddy builder, Flywheel, Pagely, Shopify, Wild Apricot, Cloudflare-edge, Microsoft IIS, Apache, nginx, OpenResty and Amazon S3. The exact rules are listed following the rubric below and in the data dictionary.

We made many comparisons current and in the field division and applied no multiplicity correction. Differences between labels or categories may indicate phase composition, category assignment, reply propensity or which servers tag themselves, fairly than item concerning the businesses themselves.

Subgroup tables keep reply rows since field and province pertain to the sampled listing; de-duplicating final destinations would necessitate deciding which sampled listing's attributes to keep.

Share gathering none of seven criteria, by header-identified label

HTTP-200 basis (n=4,701). Wild Apricot (6 sites) and Pagely (3) are too small to report. Intervals on the average acceptance indicator are omitted for readability. † Recomputed under the final rubric (X-Frame-Options copy normalization and COOP re-run); these cells moved from the before draft. Cloudflare-edge is a new row; it was unlabeled in before drafts.

Every Shopify-labeled location had a recognized clickjacking header and nosniff, and none met study-defined powerful HSTS. Every GoDaddy-builder location met powerful HSTS and had a recognized clickjacking header, and none had nosniff. Sites tagged Apache, nginx or IIS average between 0.58 and 1.80 on the acceptance index, and between 49.8% and 72.1% of them encounter none of the seven criteria. Cloudflare-edge sites average 1.72, and 59.6% encounter none of the seven.

Policies that passed our header-only script-CSP regulation are rare under all label. Of the eight, 3 are Cloudflare-edge, 2 nginx, 1 Apache and 1 IIS, and 1 has no label. Shopify, GoDaddy builder, Wix and Squarespace sites have none.

By benevolent of business

Directory categories provision us a coarse way to divided the example by category of business. Here are the nine categories alongside at smallest 80 HTTP-200 responses.

HTTP-200 basis (n=4,701). Categories under 80 responses are remaining out. Intervals on the average acceptance indicator are omitted for readability. † Recomputed under the final rubric; these cells moved from the before draft.

The identical multiplicity caveat applies here. Shopping has the lowest observed none-of-seven share: 35.5% encounter none of the seven criteria, against 56.2% for genuine estate. That is a descriptive, exploratory gap. Shopping additionally leads on recognized clickjacking headers (48.1%), nosniff (56.2%) and enforced CSP (33.3%). That fits the tag image above, since hosted shop builders lean to container those headers. But this array can't inform you why.

Real asset has the lowest observed average indicator (1.49 out of 8). It additionally has the lowest clickjacking charge (22.6%).

The CSP pillar needs the identical alert as before. Enforced CSPs display up everywhere, but lone eight policies passed our header-only script-CSP regulation throughout the complete chief base; sector-specific counts were too small to interpret.

A recognized non-default COOP value is too rare to difference by sector. The one elimination is restaurants and bars, at 4.4% [2.8-6.9], and equal there it looks akin it comes from phase defaults.

The cobbler's children

Then there is the row we had been half-dreading.

Computers & Internet is the Curlie category whose listings are concerning technology. 58.0% [51.4-64.2] of sites listed there encounter none of the seven criteria. That is the second-highest portion of any category, rearward lawful services at 59.3%. Their strong-HSTS rate, 15.5%, is the highest in the table, but lone by a little, and the intervals overlap.

Sites listed in Curlie's Computers & Internet category did not display higher header-criterion attainment. Category affiliation does not established the business's services or personnel expertise. The duration is additionally wide, and it overlaps most of the table.

By state

We cut the results by state. Nine states have at smallest 150 HTTP-200 responses, from 170 to 570 each, and the portion gathering none of the seven criteria ranges from 42.9% [36.0-50.0] in New York to 56.2% [49.2-63.0] in Florida. State estimates varied descriptively. We did not prespecify or power the study for province comparisons, and we did not behavior multiplicity-adjusted province tests, so we do not position states or infer state-level differences. A Pearson chi-square test of the 9-by-2 array (state by none-of-seven) established no detectable heterogeneity (chi-square 8.77, 8 degrees of freedom, p = 0.36; the per-state counts (state-counts-2026-09-24.tsv, 2 KB, SHA-256 f757ac8d93c5f3c51060202ce51393315f5cb1b8e65fb209d6da899b2d16b03f) are adequate to rerun it), and the unadjusted New York-Florida difference does not last a correction for the 36 imaginable pairs. The state-level counts rearward this test are in the released inspection output.

The flat under everything

There is one additional number, and it sits under all the others.

1,398 of the 7,040 sampled rows, 19.9% [18.9-20.8], gave us no usable HTTPS reply at all. 737 unsuccessful at the TLS or association stage. 215 redirected to plain HTTP. 204 timed out, and 192 had DNS that doesn't resolve. The remainder were redirect loops, DNS errors and a fistful of odd cases.

That is its own problem, concerning preparedness and TLS fairly than headers. We can't say item concerning the header posture of those sites, since we never got that far. What we can say is that for approximately one sampled row in five, our scanner did not get a usable HTTPS reply under our scan rules.

What a location owner can do, in order

The fixes are mostly short. They are not all risk-free. Roll them out in stages and test as you go.

  1. Deploy HSTS in stages. Begin alongside a abbreviated max-age, verify all affected services, afterward addition toward a lengthy duration. Add includeSubDomains lone following confirming that all current and anticipated subdomain supports HTTPS, since browsers volition refuse plain-HTTP connections to any subdomain that doesn't. Consider preload lone following gathering its requirements and accepting its operational consequences. Set the header on the naked domain as fine as www.
  2. Add clickjacking protection: X-Frame-Options: SAMEORIGIN, or CSP frame-ancestors 'self'. Check archetypal that nothing lawful embeds your pages (booking widgets, partner sites, your own apps). This volition interrupt them.
  3. Add X-Content-Type-Options: nosniff. One line, and mostly low-risk, but test MIME-dependent downloads and bequest behavior.
  4. Check your Referrer-Policy. If it is missing, the browser default is reasonable. If it is set to a bequest value akin no-referrer-when-downgrade, alter it to strict-origin-when-cross-origin.
  5. Patch your server software. Removing type numbers from Server, X-Powered-By and the ASP.NET type headers is a low-severity cleanup: it hides the issue and doesn't fix it. Keeping the application current matters more.
  6. Start a Content-Security-Policy in report-only mode, afterward move toward a nonce- or hash-based strict CSP following testing. Strict-CSP direction additionally recommends object-src 'none' and base-uri 'none', which our rubric does not require, so passing our regulation is not the identical as a production-quality strict CSP. The eight header sets that passed our regulation do not established implementation norm or nonce freshness.
  7. Be careful alongside Cross-Origin-Opener-Policy. same-origin can interrupt pop-up sign-in and fee flows. Test those before turning it on.

On hosted platforms, part of this is already done for you, and part isn't. Check which headers your phase leaves out.

How we measured

Population. Domains listed in US locality "Business and Economy" categories in the Curlie snapshot (the records in the curlie-rdf-all.tar.gz archive dated 2026-02-02, downloaded 2026-09-24 from https://curlie.org/download): 110,444 distinctive domains. Categories covering government, healthcare, education, financial services and armed forces were excluded, as were listing directories and social platforms. Business size was not verified. Curlie is human-edited and runs dead-link cleanup, so the pond skews toward established businesses alongside maintained listings.

Sample. 7,040 rows drawn at random alongside a fixed seed. These shield 7,022 distinctive first registrable domains; 18 rows shared an first registrable domain alongside another row.

Scan. We ran two HTTPS request-chain scans of all sampled URL, following at most three redirects; all reported estimates arrive from the second scan. Method: HTTPS GET. The reply build was never read; the association was closed following the reply headers. TLS verification was on, and lone community IP addresses were contacted. Timeouts were 3 s for DNS, 3 s per hop and 8 s total, alongside 24 concurrent workers. The scanner identified itself as RACKCRUNCH-header-check/1.0 alongside a communication URL for opt-out. Scans ran from Google Cloud (Oregon, US) at 12:26-12:31 EDT (first) and 13:16-13:21 EDT (second) on 2026-09-24. For scan-two requests that produced a final HTTP response, the scanner retained the complete final-response header obstacle internally (see Measurement ethics and data release).

Drift between scans. All 7,040 sampled URLs were scanned the two times. Status transitions from the archetypal scan to the second:

The archetypal scan recorded lone per-check values truncated at 400 characters, so a complete difference under the final rubric is not possible. Of the 7,040 sampled rows, 4,608 returned HTTP-200 in the two scans. Under the first rubric, 13 of those rows (0.28%) changed at smallest one check, and the largest alter in any header prevalence between scans was 0.25 points (X-Content-Type-Options, all-usable base). Those are old-rubric diagnostics, not evidence that the corrected classifications are stable. As a bounded check, we applied the final rubric to the 4,484 of those rows (97.3%) whose scan-1 inspect values were complete. 7 rows changed at smallest one indicator item, 0.16% [0.08-0.32]: clickjacking 2, Permissions-Policy 3, X-Content-Type-Options 3, type hygiene 2 (some rows changed additional than one). The 124 excluded rows are the ones alongside the longest CSPs, so this supports short-term stability but doesn't demonstrate it.

Disposition. 5,642 usable HTTPS responses: position 200: 4,701; 403: 577; 202: 245; 404: 59; 500: 11; 307: 9; 429: 8; 401: 5; 503: 4; 405: 4; 520: 3; 526: 3; 400: 3; 406: 2; 521: 2; 502, 525, 410, 423, 523 and 530: 1 each. 1,398 unusable (each nonaccomplishment category is defined in the data dictionary): TLS or nexus nonaccomplishment 737, redirect to non-HTTPS 215, timeout 204, DNS does not determine 192, too many redirects 29, DNS lookup unsuccessful 13, non-public location 4, invalid redirect 4. Responses alongside HTTP position 202 (245 responses), frequently generated by bot-protection layers, were classified as reachable but excluded from the HTTP-200 analysis.

Duplicates and parked pages. Usable responses resolved to 5,614 distinctive final registrable domains; HTTP-200 responses to 4,688, the chief base. Among usable responses, lone 15 final registrable domains were reached from additional than one sampled row (43 responses in total, at most 9 for any one final registrable domain). Deduplication moves estimates by 0.3 points or less. 9 HTTP-200 responses (0.19%) landed on parked, expired-domain or phase pages. These were identified from the final hostname following redirects (for example expireddomains.com or forsale.godaddy.com), not from leaf content. They remain in the basis since their headers are what the scanner received.

Intervals. Wilson 95% confidence intervals throughout (Wilson, 1927). The random diagram was of directory rows, and the chief basis is distinctive final registrable domains following redirects, so the intervals on that basis are approximate. They quantify sampling doubt under the stated example model, not example bias, parser misclassification, CDN variation, or selective phase attribution. Label, field and province cuts are exploratory.

Rubric

The eight criteria of the explicit-header acceptance index, as used in this report. The complete rubric code is in the reproduction pack (repro-pack-2026-09-24-v2.5.1.tar.gz, 239 KB, SHA-256 9e00c1a41ac412f3354eecb4686ba929aeaed359ad2e859e52c312e51cfd30b3), and all accepted, weak, malformed and no-op value is defined in the data dictionary (data-dictionary.md, 7 KB, SHA-256 76469bc6e27a98df28c9a422c15da1a2b56d2bf05a22c7be08bbf6737177f6a2).

The "none of seven" header uses the archetypal seven items and leaves out type hygiene.

Hosting labels. Rules are checked in this command and the archetypal equivalent wins. Matching is case-sensitive.

  1. GoDaddy builder: CSP contains "godaddy.com", or Server starts alongside "DPS"
  2. Flywheel: Server starts alongside "Flywheel"
  3. Pagely: Server starts alongside "Pagely"
  4. Shopify: CSP starts exactly alongside "block-all-mixed-content; frame-ancestors 'none'; upgrade-insecure-requests"
  5. Wild Apricot: CSP contains "wildapricot"
  6. Cloudflare-edge: Server is exactly "cloudflare"
  7. Microsoft IIS: Server starts alongside "Microsoft-IIS"
  8. Apache: Server starts alongside "Apache"
  9. nginx: Server starts alongside "nginx"
  10. OpenResty: Server starts alongside "openresty"
  11. Amazon S3: Server starts alongside "AmazonS3"

Anything alternatively gets no label.

Caveats

  • One page, one moment. Headers can differ by page, CDN border or login state. We peruse the residence leaf only.
  • Sites rearward a CDN or hosted phase frequently display the platform's headers, not the owner's configuration. That is genuine safety for the site, and it additionally method several passes are inherited.
  • We did not test the http:// to https:// redirect. Our checker does, but these scans followed one petition sequence per site.
  • The example comes from a human-edited directory. It skews toward businesses established adequate to be listed, and endeavor size was not verified.
  • Labels shield the 74.0% of HTTP-200 responses that acknowledge a server or hosting phase in their headers. The comparisons depict those sites. They don't evaluation what a phase causes.
  • The clickjacking fig includes 116 sites credited since we treat identical repeated X-Frame-Options values, joined by our scanner, as one value (see Rubric). Under the current HTML handling model, identical repeated values specified as SAMEORIGIN, SAMEORIGIN keep same-origin framing protection; older browser implementations have differed. Counting those as ineffective would provision 1,347 (28.7%).
  • This is a ray scan, not a safety audit of any idiosyncratic business.

Sensitivity analyses

All usable responses. Every usable HTTPS response, including error and bot-challenge pages (n=5,642). On this base, 48.4% [47.1-49.7] encounter none of the seven criteria (2,733 of 5,642).

The COOP and CSP figures current need a caveat. Most of the 352 same-origin COOP values arrive from CDN difficulty pages, and so do the CSPs: 325 of the usable responses (5.8%) transport nonce/hash-strict CSPs, but these are mostly CDN bot-challenge interstitials (Cloudflare difficulty pages transport nonce CSPs), not deployments by the sites; 316 of the 318 script-src-attr uses on this basis sit on those Cloudflare responses. On the chief basis that category is 5 sites. The remainder of the mix: no CSP 4,301 (76.2%), no manuscript directive 829 (14.7%), allowlist-based 177 (3.1%), strict-dynamic alongside anchor 8 (0.1%), invalid 2.

Base comparison. The identical measures throughout bases:

COOP drops from 6.6% to 1.1% on the HTTP-200 base, which excludes most identified difficulty responses, and Referrer-Policy and Permissions-Policy autumn by concerning a third. Deduplication barely moves anything. The none-of-seven portion is concerning fractional on all three bases.

Measurement ethics and data release

We scanned lone community residence pages complete HTTPS, one petition sequence per sampled URL per scan, following at most three redirects. We never peruse leaf bodies, submitted forms, logged in or probed for vulnerabilities. The scanner contacted lone community IP addresses, ran at most 24 concurrent connections throughout the entire example alongside abbreviated timeouts, and identified itself as RACKCRUNCH-header-check/1.0 alongside a communication URL for opt-out. We have received no opt-out requests to date.

For all scan-two petition that produced a final HTTP response, the scanner retained the complete final-response header obstacle internally. The lone bounds was a cap of 65,536 bytes on the entire block, applied at grasp and before any classification; the largest captured obstacle was 15,362 bytes, so no scan-two final-response header obstacle was truncated by the capture-size cap. Headers from redirect hops were not kept. The archetypal scan kept lone per-check values, cut at 400 characters following classification, which is why drift between scans can lone be bounded. The released dataset carries category labels, not header values, so readers can recompute all fig in this study from it but cannot re-grade the headers themselves. The scans were unauthenticated GET requests alongside no cookies sent, so the data holds nothing beyond what any visitor's browser would receive.

The released dataset (scan2-2026-09-24-deidentified.jsonl, 4.1 MB, SHA-256 329a1a910b4ae4b005fdd864fe9b5370df7e30bf01fc14b790baf7664d00428d) has exactly one row for all of the 7,040 sampled rows, and nothing more. Rows without a usable reply transport lone their error category and HTTP status. Each row lists the sector, state, whether the reply was usable, the error class, HTTP status, figure of redirects, a category tag for all graded header (HSTS, CSP, X-Frame-Options, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, COOP and version-token disclosure) and the per-criterion results. It holds no header values, no scan times and no raw header map, so no cookies, CSP nonces, reporting endpoints or type strings depart our side.

The dataset is de-identified. Every endeavor appears lone under a stable pseudonymous ID, and neither this study nor the dataset names any location as having feeble or missing headers. We chose this flat since the item of the study is the general picture, and naming small businesses alongside feeble headers would add hazard for them without making that image any clearer. If you think your endeavor power be in the sample, run your own domain through the liberated RACKCRUNCH Security Header Check (rackcrunch.com/security-headers) to see what the scanner saw. We volition verify a particular domain's row lone to a verified owner of that domain.

Technical sources

Data and reproducibility

Everything rearward this study is published. The de-identified dataset has one row for all of the 7,040 sampled homepages, and the data dictionary explains all site and value in it. The reproduction pack holds the dataset, the sampling, scan and inspection scripts alongside the random seed, the rubric and CSP parser, and the tests; its README (README-2026-09-24-v2.5.1.md, 7 KB, SHA-256 6fa74c65096a2981f784d2f7c0b168e90e66c91272f0768042af9d3454415c7c) walks through all stage and lists what cannot be rerun. The identical records are accessible as a git bundle (repro-pack-2026-09-24-v2.5.1.bundle, 238 KB, SHA-256 250306e16e674c9e7e95c5f913ca2016fa63d39fb07a3350d5c436575bf7d303), tagged release-2026-09-24-5-1 (commit f96a488c7922df11ca0c83e78f0ef739f7ef4bca).

One command regenerates all array in this study as final-tables.md (final-tables.md, 20 KB, SHA-256 07c5a2aea843a4c6ef03c8cfe36848074cb367a724c63c822d8d3f61e422ed90):

pip instal -r requirements.txt && python3 reanalysis4.py data/ final-tables.md

The lone dependency exterior the Python norm archive is pinned in requirements.txt (requirements.txt, 1 KB, SHA-256 7d5f5a94ac1690db632a2f6b2e51aaf74a8ba332ee248aa6299ec6e713120946). The two charts are drawn from those tables: the adoption-index diagram from the allocation array in division 3, and the tag diagram from the phase array in division 13. There is no distinct diagram code. The parser tests, including the aureate corpus, run with:

python3 test_csp3.py && python3 -m unittest test_golden

Software versions used: Python 3.10.12; scipy 1.15.3, pinned in requirements.txt (used lone for the province chi-square; everything alternatively is the Python norm library). Rerunning the scan itself is not part of the command: it needs live network admission and the directory snapshot, and results alter complete time.

To inspect a download, difference its SHA-256 (shown in small imprint following all link) alongside sha256sum FILE on Linux or shasum -a 256 FILE on macOS. Inside the pack, MANIFEST.sha256 (MANIFEST.sha256, 1 KB, SHA-256 e6c533998c3675c01fce0c5fe160df6c182c8996a570522a30081be6341bcf3b) lists the hash of all file; inspect them all alongside sha256sum -c MANIFEST.sha256.

Files inner the reproduction pack
  • .gitignore 1 KB, SHA-256 862263fa1f46c20f0d1e4dac5ffcc75abd55c08211b2c3864c5f8764b9d87793
  • MANIFEST.sha256 1 KB, SHA-256 e6c533998c3675c01fce0c5fe160df6c182c8996a570522a30081be6341bcf3b
  • README.md 7 KB, SHA-256 6fa74c65096a2981f784d2f7c0b168e90e66c91272f0768042af9d3454415c7c
  • csp3.py 10 KB, SHA-256 253713b8ec414ce9d93e6e180254d67c3409ffae5e94d5675315ff0e52eadb1f
  • data-dictionary.md 7 KB, SHA-256 76469bc6e27a98df28c9a422c15da1a2b56d2bf05a22c7be08bbf6737177f6a2
  • data/scan2-2026-09-24-deidentified.jsonl 4.1 MB, SHA-256 329a1a910b4ae4b005fdd864fe9b5370df7e30bf01fc14b790baf7664d00428d
  • data/state-counts-2026-09-24.tsv 2 KB, SHA-256 f757ac8d93c5f3c51060202ce51393315f5cb1b8e65fb209d6da899b2d16b03f
  • final-tables.md 20 KB, SHA-256 07c5a2aea843a4c6ef03c8cfe36848074cb367a724c63c822d8d3f61e422ed90
  • reanalysis4.py 28 KB, SHA-256 79d759a2372f57b1c15b0e5136d9e3935c6c08fce464106accf13405c0090c16
  • requirements.txt 1 KB, SHA-256 7d5f5a94ac1690db632a2f6b2e51aaf74a8ba332ee248aa6299ec6e713120946
  • rubric3.py 7 KB, SHA-256 3bc61fc02e75e45a96147907cfc604579c461bc74dadc6e597b13823cf025517
  • scripts/build-deidentified-dataset.py 5 KB, SHA-256 3d86b0d84af020d12a3131ceb91a4ae642d85ac9b486cfeebe631640f498e4c6
  • scripts/sample-from-curlie.py 3 KB, SHA-256 2c19a1af2c5b1dcb94bb897d28937f1f3c3807b2efae584eae47e5444c60ae16
  • scripts/scan.py 9 KB, SHA-256 7ea184650c9373829c51f59bd0144a4f916b33e5f33eb65d5d7c5996d41ed206
  • scripts/scan2.py 9 KB, SHA-256 657b04b409926eb1a4f26829e10174927c31be93ef34ae3f60e1a70542bea711
  • test_csp3.py 7 KB, SHA-256 f0f94db40395f76460616de801d67f1e852ea67da38a63265d0d1ae80cfc0d4b
  • test_golden.py 29 KB, SHA-256 d53f3c3b8ba0bdddb7bb7f3cc751ff1a1b89aec8f0d5fa7ae41ebc1f32ebf800

Re-identification caveat. Rows transport lone a example row figure and a keyed pseudonym, and the key is not released. But anyone who reruns the sampling manuscript alongside its fixed kernel (42) against the identical directory snapshot can rebuild the example and equivalent rows to sites. That comes alongside reproducible sampling. No location is named anyplace in the release; the eight qualifying Content-Security-Policy examples from the manual audit appear lone under pseudonyms.

Version former and errata

  • 2026-09-24 (v7.4.2). The PDF's COOP note now matches this page; no numbers changed.
  • 2026-09-24 (v7.4.1), group v2.5.1. Added the redirect figure to the catalog of released fields; no numbers changed.
  • 2026-09-24 (v7.4), group v2.5. Removed the eight named qualifying sites from the data release, and confirmed the study names none of them, so the de-identification assertion holds; no numbers changed.
  • v7.3, group v2.4 (2026-09-24). The names for the inspection bases were made accordant throughout this page, the PDF, the README, the data dictionary and the generated tables, and the definitions box was added. No figure changed.
  • v7.2 (2026-09-24). Byline and PDF heading metadata corrected. No figure changed.
  • v7.1 (2026-09-24). First community release.

How to cite

RACKCRUNCH Team. Security Headers on Directory-Listed U.S. Local-Business Websites, study type 7.4.2, reproduction group v2.5.1. RACKCRUNCH, 2026-09-24. https://rackcrunch.com/security-headers-2026

Artifact hashes

SHA-256 of all document published alongside this study. They are computed at build period from the exact records served.

  • MANIFEST.sha256 1 KB, SHA-256 e6c533998c3675c01fce0c5fe160df6c182c8996a570522a30081be6341bcf3b
  • README-2026-09-24-v2.5.1.md 7 KB, SHA-256 6fa74c65096a2981f784d2f7c0b168e90e66c91272f0768042af9d3454415c7c
  • data-dictionary.md 7 KB, SHA-256 76469bc6e27a98df28c9a422c15da1a2b56d2bf05a22c7be08bbf6737177f6a2
  • final-tables.md 20 KB, SHA-256 07c5a2aea843a4c6ef03c8cfe36848074cb367a724c63c822d8d3f61e422ed90
  • repro-pack-2026-09-24-v2.5.1.bundle 238 KB, SHA-256 250306e16e674c9e7e95c5f913ca2016fa63d39fb07a3350d5c436575bf7d303
  • repro-pack-2026-09-24-v2.5.1.tar.gz 239 KB, SHA-256 9e00c1a41ac412f3354eecb4686ba929aeaed359ad2e859e52c312e51cfd30b3
  • requirements.txt 1 KB, SHA-256 7d5f5a94ac1690db632a2f6b2e51aaf74a8ba332ee248aa6299ec6e713120946
  • scan2-2026-09-24-deidentified.jsonl 4.1 MB, SHA-256 329a1a910b4ae4b005fdd864fe9b5370df7e30bf01fc14b790baf7664d00428d
  • security-headers-study-2026.pdf 400 KB, SHA-256 c8fa6a281799ff1f86c79b62f5ba5b52535a4ca9d9b6bdce69d61f5bfaeb6d91
  • state-counts-2026-09-24.tsv 2 KB, SHA-256 f757ac8d93c5f3c51060202ce51393315f5cb1b8e65fb209d6da899b2d16b03f

Sample drawn from the Curlie directory (https://curlie.org), used under the Creative Commons Attribution 3.0 Unported License. Sampling frame: the records in the curlie-rdf-all.tar.gz archive dated 2026-02-02, downloaded 2026-09-24 from https://curlie.org/download.

With satisfied from Curlie.org - the largest human-edited directory of the web. Contribute by submitting a website or becoming an editor. Curlie data is licensed under the Creative Commons Attribution 3.0 Unported License.

Questions and opt-out

Write to [email protected] alongside questions, corrections or an opt-out request.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads