Sembrelio

SembrelioBot

SembrelioBot is the web crawler of Sembrelio, a research prototype for an Executive MBA thesis.

Status: active. SembrelioBot fetches pages only for websites submitted for an assessment.

Purpose

SembrelioBot fetches public pages of a website only after a person who states that they are authorised to do so has submitted the website for an assessment. The text of the pages is used to extract facts about the business, such as services, location and contact details. These facts are compared with answers from AI services. SembrelioBot does not crawl the web at large and does not collect content for training AI models.

What SembrelioBot requests

First robots.txt, then the start page, sitemap.xml (and sitemaps listed in robots.txt) and llms.txt, then further pages of the same website, with pages such as legal notice, contact and services first. Only the submitted website is fetched, with and without "www"; other domains, subdomains and linked websites are not. SembrelioBot does not run JavaScript. The page as delivered, its text and a Markdown version are stored for the assessment.

User agent

SembrelioBot identifies itself with this user agent:

SembrelioBot/0.1 (+https://sembrelio.com/bot)

robots.txt

SembrelioBot reads robots.txt before fetching pages and follows the rules for the token SembrelioBot, or the rules for all user agents (*) if there are no specific rules. If robots.txt cannot be read because of a server error or a timeout, SembrelioBot fetches no pages. To block SembrelioBot from your whole website, add:

User-agent: SembrelioBot
Disallow: /

Limits

At most 30 pages per website and assessment, at most one request per second (slower if robots.txt sets a Crawl-delay), with a timeout of 20 seconds per request. SembrelioBot does not keep cookies, does not log in, submit forms or bypass paywalls.

Contact

Questions or problems with SembrelioBot: bot@sembrelio.com