Benchmarks
The state of small-business websites, measured
From 5,297 free scans of publicly reachable business sites, measured 2026-09-07. Every share carries its sample size and a 95% interval.
These numbers come from Go Voltic's own scanner: 5,297 four-page scans of publicly reachable business sites, 4,368 of them in the 30 days ending 2026-09-07. Each row is the share of scanned sites that raised a finding. No site is named here and none can be identified: this page publishes aggregates only. A scan that could not load a page is excluded from every denominator and counted separately, because a site we could not read is not a site with a score of zero.
The eight most common findings
| Finding | Share of sites | What it means |
|---|---|---|
| No skip-to-content link | 87.2%95% CI 86.3% to 88.1% | Keyboard users must tab through the whole navigation on every page. |
| No article or breadcrumb markup | 77.5%95% CI 76.4% to 78.6% | Content pages carry no Article or BreadcrumbList markup for machines. |
| Content only appears after scripts run | 77%95% CI 75.8% to 78.1% | The page source carries little readable text, so crawlers that do not run scripts see almost nothing. |
| No terms linked | 75.1%95% CI 73.9% to 76.2% | No linked terms governing the service. |
| No privacy policy linked | 70.5%95% CI 69.2% to 71.7% | No linked statement of what happens to visitor data. |
| No main landmark | 63.6%95% CI 62.3% to 64.9% | No main element, so assistive technology cannot jump to the content. |
| H1 missing or duplicated | 58.4%95% CI 57.1% to 59.7% | The page has no H1, or several, instead of one clear top heading. |
| No social proof above the fold | 56%95% CI 54.6% to 57.3% | Nothing near the top shows other people trust the business. |
Share is the fraction of the 5,297 scanned sites that raised the finding. The interval is a 95% Wilson interval.
Can AI read and repeat these sites?
The single loudest number in the corpus is the gap between two findings. Almost nobody blocks AI crawlers on purpose: robots.txt turns them away on 0.9% of sites. But 77% of sites serve pages whose text barely exists until scripts run, and a crawler that does not run scripts reads almost nothing. The lock on the door is rare. The empty room behind it is the norm.
| Finding | Share of sites | What it means |
|---|---|---|
| Content only appears after scripts run | 77%95% CI 75.8% to 78.1% | The page source carries little readable text, so crawlers that do not run scripts see almost nothing. |
| No article or breadcrumb markup | 77.5%95% CI 76.4% to 78.6% | Content pages carry no Article or BreadcrumbList markup for machines. |
| No business entity in structured data | 44.6%95% CI 43.3% to 46% | Structured data exists but never states the organization behind the site. |
| No structured data | 41.5%95% CI 40.2% to 42.8% | Nothing on the page tells machines what the business is in machine terms. |
| No plain summary for machines to lift | 29.7%95% CI 28.5% to 31% | Title or description missing, so an assistant has no clean answer to quote. |
| Page regions not marked semantically | 52.4%95% CI 51% to 53.7% | No header, nav, main or footer landmarks, so machines parse a soup of divs. |
| Home page thin on content | 21%95% CI 19.9% to 22.1% | Too little substantive text for a machine to learn what the business does. |
| AI crawlers blocked in robots.txt | 0.9%95% CI 0.7% to 1.2% | The robots file explicitly turns away AI assistants’ crawlers. |
Trust signals
| Finding | Share of sites | What it means |
|---|---|---|
| No terms linked | 75.1%95% CI 73.9% to 76.2% | No linked terms governing the service. |
| No privacy policy linked | 70.5%95% CI 69.2% to 71.7% | No linked statement of what happens to visitor data. |
| No social proof above the fold | 56%95% CI 54.6% to 57.3% | Nothing near the top shows other people trust the business. |
| No obvious next step | 43.4%95% CI 42.1% to 44.8% | The page never tells a ready visitor what to do next. |
| No security headers | 44.9%95% CI 43.6% to 46.3% | None of the standard browser protection headers are set. |
| Tracking runs with no consent notice | 41.7%95% CI 40.4% to 43.1% | Analytics or ad scripts fire before any consent choice exists. |
| No visible way to reach the business | 30.8%95% CI 29.6% to 32.1% | No phone, address or contact route a visitor can find. |
| Insecure resources on a secure page | 30.9%95% CI 29.6% to 32.1% | An HTTPS page loading HTTP resources, which browsers flag or block. |
Accessibility
| Finding | Share of sites | What it means |
|---|---|---|
| No skip-to-content link | 87.2%95% CI 86.3% to 88.1% | Keyboard users must tab through the whole navigation on every page. |
| No main landmark | 63.6%95% CI 62.3% to 64.9% | No main element, so assistive technology cannot jump to the content. |
| Images missing alt text | 17.9%95% CI 16.8% to 18.9% | Images carry no text alternative for screen readers or image search. |
| Form fields without labels | 17.9%95% CI 16.9% to 19% | Inputs a screen reader announces as nothing at all. |
| Pinch-zoom disabled | 11.5%95% CI 10.7% to 12.4% | The viewport tag blocks zooming, which some visitors need to read at all. |
| Screen readers cannot tell the language | 10.7%95% CI 9.9% to 11.5% | Assistive technology has to guess how to pronounce the page. |
| No mobile viewport tag | 9.5%95% CI 8.7% to 10.3% | The page renders desktop-sized on phones. |
| Frames without titles | 24.8%95% CI 23.7% to 26% | Embedded frames a screen reader can only announce as frame. |
Speed
| Finding | Share of sites | What it means |
|---|---|---|
| Render-blocking scripts | 47.8%95% CI 46.5% to 49.1% | External scripts load without defer or async and stop the page from drawing. |
| Heavy HTML document | 42.9%95% CI 41.6% to 44.2% | The page document itself is large before a single image loads. |
| Images all load at once | 38.3%95% CI 37% to 39.6% | No lazy loading, so offscreen images compete with the visible page. |
| Many separate stylesheets | 36.7%95% CI 35.4% to 38% | Each stylesheet is a separate blocking download. |
| Images shift the layout | 32.6%95% CI 31.3% to 33.9% | Images without declared dimensions move the page as they arrive. |
| Slow server response | 21.9%95% CI 20.8% to 23% | The server takes too long to send the first byte. |
| Compression not confirmed | 9.3%95% CI 8.5% to 10.1% | The server did not confirm gzip or brotli on the document. |
What is mostly solved
Some problems the industry talks about are nearly gone in this population. HTTPS is missing on 0.3% of sites. Deliberate AI-crawler blocking sits at 0.9%. The common failures above are quieter than either and far more widespread.
By sector
Sector pages draw on a different population: deep scans from our research sweep, sector-labelled, up to 40 pages per site. Deep and four-page scans measure at different depths, so the two populations are never pooled into one number.
Get the data
The whole table behind this page, every finding with its share, sample size and interval, for the free population and for each sector cut. Two files, regenerated with the page: latest.csv and latest.json (measured 2026-09-07).
A dated edition is written the first time the page is built in a month and never changes afterwards, so a number you cite stays where you found it.
How to cite
Go Voltic (2026). The state of small-business websites, measured: 2026-09 edition, 5,297 free scans, measured 2026-09-07. https://go-voltic.com/benchmarks
These files are published under the Creative Commons Attribution 4.0 license (CC BY 4.0): reuse and adapt them for any purpose, with attribution to Go Voltic and a link to this page. Keep the sample size with the share. No scanned site is identified in any of these files.
Method
Population: every publicly reachable business site scanned by Go Voltic's free four-page scan that loaded at least one page (5,297 scans at 2026-09-07), of which 4,368 ran in the 30 days ending 2026-09-07. Each finding is a deterministic check described on its own row; the checks are the same ones the paid reports grade. Shares are per scanned site. Intervals are 95% Wilson intervals. Sites that refused the scan are excluded from denominators and tallied separately. Figures regenerate from the live measurement store; the date above is the measurement date, not a publication date. More detail: our methodology.
See where your own site stands. Run the free scan
Citing these numbers is welcome: see how to cite and get the data.