Includoc, home

IncludocBot

IncludocBot finds documents on websites whose owners verified them in Includoc, and fetches public files people ask us to check. How to control it.

Last updated

IncludocBot is the web crawler and file fetcher for Includoc, a document accessibility checker. If you've seen it in your server logs, this page explains what it does, how it behaves and how to control it.

How it identifies itself

IncludocBot sends this user agent with every request:

IncludocBot/1.0 (+https://includoc.com/bot)

What it does

IncludocBot does two jobs.

1. Website inventories for verified owners

Organizations that use Includoc can build an inventory of every document on their own website. IncludocBot crawls a site only after someone has added it to Includoc and verified that their organization owns it, with a DNS record, a meta tag or an email to the domain's administrator.

On a verified site, IncludocBot:

  • Reads your sitemap first, then follows links to find more pages
  • Stays within the verified domain and never follows links to other sites
  • Looks for links to PDF, Word, PowerPoint and Excel files
  • Records each file's address, the pages that link to it, its size, type and last-modified date
  • Downloads the documents it finds, so the site's owner can check them for accessibility
  • Visits again on the owner's schedule, weekly or daily, and uses standard headers (ETag and Last-Modified) to skip files that haven't changed

On sites that build their links with JavaScript, IncludocBot may load pages in a headless browser to find those links. It still identifies itself as IncludocBot and follows the same rules.

2. Single files people ask us to check

When someone pastes the address of a public document into our checker, IncludocBot fetches that one file so we can check it. It doesn't follow links from the file or crawl the rest of your site. Size and file-type limits apply, and it only fetches files that are publicly available without signing in.

What it doesn't do

  • It doesn't fill in forms, submit data or sign in to anything.
  • It doesn't try to get past passwords, paywalls, CAPTCHAs or other access controls.
  • It doesn't follow links off a verified domain.
  • It reads web pages only to find links to documents. It doesn't keep copies of your pages.
  • Nothing it fetches is used to train AI models.

How fast it crawls

IncludocBot makes at most one request per second to any host.

Controlling IncludocBot with robots.txt

IncludocBot follows the rules in your site's robots.txt file, for website inventories and for single-file checks. If your robots.txt can't be read because of a server error, a single file that someone asked us to check is still fetched once.

To block IncludocBot from your whole site, add this to your robots.txt:

User-agent: IncludocBot
Disallow: /

To keep it out of one folder:

User-agent: IncludocBot
Disallow: /internal/

If your robots.txt has no group for IncludocBot, it follows your rules for all crawlers (the User-agent: * group).

If you're the verified owner and want IncludocBot to reach a folder your robots.txt closes to other crawlers, add an Allow rule for IncludocBot:

User-agent: IncludocBot
Allow: /documents/

Blocking IncludocBot also stops your own organization's Includoc inventory from finding documents, so check with your web team before you add a rule.

Report a problem

If IncludocBot is causing trouble on your site, or you think it's crawling a site it shouldn't, email support@includoc.com. Tell us your domain and roughly when you saw the requests. We'll look into it and can stop crawling your site.

To report a security issue, email security@includoc.com.

  • Trust and security

    How Includoc stores, protects and deletes your documents: US storage, encryption, short retention, isolated processing and no AI training on your files.

  • About Includoc

    Includoc helps public agencies, colleges and consultants check, fix and prove document accessibility, with honest AI and human review. Not an overlay.