Crawler policy

How this crawler behaves, and how to have your site removed.

Requests identify themselves:

MovesWatch/1.0 (+https://moves.watch/crawler)

The User-Agent is always truthful. Where a site blocks datacenter TLS fingerprints and its robots.txt permits the path, the TLS handshake may be made to look like a browser's. The User-Agent above does not change.

How each signal is read

What the site doesHow it is readWhat happens
robots.txt disallows the pathA machine-readable refusalNever fetched. No override exists
robots.txt is silentNo stated objectionFetched
Datacenter IP or TLS blockA blunt filter, no stated intentMay be worked around, per above
403 with a challenge pageAn active refusalStops, recorded as blocked
CAPTCHA or JS challengeAn explicit refusalNever solved. Recorded as blocked
Login wallNot public dataStops, recorded as requiring auth
An email from the site ownerA refusalRemoved permanently

robots.txt and a CAPTCHA state an intent. A fingerprint check does not, so it is read as a filter rather than as a refusal.

Rate

One request at a time per site, at least two seconds apart. Retry-After is honoured. Three consecutive failures stop the crawl of that page. robots.txt is re-read every 24 hours. Reads of the Internet Archive are slower still: one request per second on a single connection, off-peak.

What gets published

WhatPublicSigned in
Extracted prices, plans, and featuresYesYes
The changed lines around an edit (25 at most)YesYes
The full captured pageNoYes, rate-limited, not indexed
Link to your page and to the Internet Archive copyYesYes

Prices and plan names are facts, and facts are reportable. Marketing prose on the same page stays yours, which is why the public view carries the changed lines and a link to the original rather than the page itself.

Removal and correction

RequestResponse
Stop crawling usHonoured permanently, and recorded
Remove a specific captureThe stored page is deleted; the record that it existed, and when, remains
Remove all historyConsidered case by case. Extracted facts are reportage and are not removed by default
Something here is wrongRe-fetched and re-read, with the correction published alongside the original reading

Write to crawler@moves.watch. A person answers within five business days.

Personal data

Careers pages are the only surface that carries personal data in any volume, and the useful signal there is aggregate: how many roles of what kind, where, and when. Names, direct email addresses, phone numbers, and links to individual profiles are stripped before anything is written to disk, so none of them are stored at any point. Careers captures are kept for 24 months.