Crawler Verification
Crawler verification process
A service name in a User-Agent does not establish that the request came from that service. Verification checks DNS, published IP ranges, ASN information, and behavioral patterns against the claimed crawler identity.
Why verification matters
- Distinguish impersonation
- Improve analytics accuracy
- Understand resource use
- Evaluate access against security requirements
- Improve delivery to legitimate crawlers
| Method | What it checks |
|---|---|
| Reverse DNS | Hostname and organization domain associated with the IP |
| IP ranges | Membership in published crawler address ranges |
| ASN | Source network ownership |
| Heuristics | Characteristic signatures and access patterns |
A shared-cloud ASN, User-Agent, or behavioral pattern alone cannot establish service authenticity. The list below describes verification methods; it does not imply equal verification confidence for every crawler.
Platform-specific verification
Platforms verifiable through published information
- OAI-SearchBot: IP range verification + ASN verification
Partially verified platforms
Some platforms use common cloud provider IPs or don't publish verification methods, making complete verification challenging:
- You.com
- Bytedance
- Yahoo
- OpenClaw
- DeepSeek
- Baidu
- BaiduSpider: Reverse DNS verification
- Huawei
- Yandex
- Gemini
Stay updated
Review verification methods as platforms emerge, crawler infrastructure changes, new methods become available, and security requirements evolve. Do not treat a fixed list as sufficient for future traffic.
Important note about data updates
Changes to identification methods can change historical and current classifications. When traffic shifts substantially, distinguish actual volume changes from verification or classification updates.