About the crawler
It is building a free, public search index of official government websites so that people can find what their governments publish. This page explains what it does, and how to slow it down or turn it off.
GovCrawlBot/1.0 (+https://crawl.govglance.org/about)GovCrawlBot are applied
first, and User-agent: * otherwise. Crawl-delay is honoured
up to a cap.ETag or Last-Modified, it will send
If-None-Match and take a 304 for an answer.# In your robots.txt
User-agent: GovCrawlBot
Crawl-delay: 10
User-agent: GovCrawlBot
Disallow: /search
Disallow: /calendar
User-agent: GovCrawlBot
Disallow: /
Changes take effect within a day, which is how long robots.txt is cached. If you need something removed sooner, or a URL dropped from the index, email us and we will handle it by hand.
403 that looks identical to a transient failure, so the crawler
keeps retrying on its normal schedule. A robots.txt rule is understood, respected, and
costs your server one request a day instead.
Government information is public by law and hard to find in practice. It is spread across tens of thousands of separate websites, most with a search box that only covers that one site, and commercial search engines index them unevenly and rank them against everything else on the web.
This index covers government domains and nothing else, which means a search for a form, an ordinance, or a council agenda returns the government's own page rather than a copy of it. Everything indexed is already public; nothing behind a login is touched.