Measuring the Citation Gap
428 small and mid-sized business websites across four countries, measured as a retrieval client sees them — not as a browse…
The Source Code of Technical Reports
428 small and mid-sized business websites across four countries, measured as a retrieval client sees them — not as a browse…
The plain-English guide to getting your videos found, read and cited by Google, Bing and AI answer engines — the search surfaces …
The plain-English guide to getting your business found, read and recommended by ChatGPT, Perplexity, Google AI and Claude — witho…
A lawyer’s cheat sheet for out-engineering the legal market. Stop the hidden Technical Tax™ on your Google Ads — low mobile Quali…
Halt the Technical Tax on your business website in 48 hours. The hidden premium you pay on every Google Ads click because of slow…
428 small and mid-sized business websites across four countries, measured as a retrieval client sees them — not as a browser renders them.
Retrieval-grounded generative systems can fetch web sources and construct answers that expose citations. Existing measurement of web quality addresses the rendered document: accessibility audits execute JavaScript, structured-data surveys sample crawler corpora, and consent studies read robots.txt in isolation. None measures what a retrieval client actually receives when it requests an ordinary business homepage and is not a browser.
This paper specifies a protocol for that measurement and applies it across the United States, the United Kingdom, Canada and the Netherlands, in home services, legal, medical and dental, and a mixed small-business group. Each domain is requested under up to four conditions with every attempt recorded, so refusals are reported rather than converted into missing data, and seven binary preconditions are observed with the supplying observer recorded per value.
The Citation Gap is then defined operationally: sites that satisfy the access preconditions — reachable, permitted, indexable — but do not satisfy all four extraction-readiness preconditions the protocol specifies. Of 349 access-eligible sites, 187 (53.6%) fall below that threshold. Its structure is unexpectedly shallow: 114 of the 187 fail exactly one extraction precondition, most often a top-level heading (45 sites), structured data (38), or a meta description (31), and none fails solely on non-JavaScript content.
Adoption of llms.txt is 20% overall and concentrates among sites already outside the gap (29% versus 14%). Blocking of AI crawlers in robots.txt is rare — 13 of 382 measurable sites — and the observed access failures occur instead at the network edge. An independent local implementation of Google’s agentic accessibility-tree audit agrees with it on 55% of sites, so that measure is reported as implementation-specific.
The sample is designed to exercise the protocol across heterogeneous business websites, not to estimate population prevalence. No engine behaviour is observed and no causal claim is made.
By Vikas — vSourceCode.com
© 2026 vSourceCode.com · Technical Tax™ · Viewport Attention Zone™ · vKernel™ · Business Vir™ · Concepts by vsourcecode.com
By Vikas — vSourceCode.com
© 2026 vSourceCode.com · Technical Tax™ · Viewport Attention Zone™ · vKernel™ · Business Vir™ · Concepts by vsourcecode.com
By Vikas — vSourceCode.com
© 2026 vSourceCode.com · Technical Tax™ · Viewport Attention Zone™ · vKernel™ · Business Vir™ · Concepts by vsourcecode.com
By Vikas — vSourceCode.com
© 2026 vSourceCode.com · Technical Tax™ · Viewport Attention Zone™ · vKernel™ · Business Vir™ · Concepts by vsourcecode.com
By Vikas — vSourceCode.com
© 2026 vSourceCode.com · Technical Tax™ · Viewport Attention Zone™ · vKernel™ · Business Vir™ · Concepts by vsourcecode.com
Share your site and Vikas will find the Technical Tax. No pitch. Just the numbers.