# robots.txt for DataDoe — AI-first Amazon MCP data layer # Updated: 2026-09-14 # ============================================ # Sitemap discovery # ============================================ # The main sitemap is appended automatically by Webflow # (Site settings > SEO > "Remove sitemap.xml from robots.txt" = Off). # Do not add it here or it will appear twice. Sitemap: https://www.datadoe.com/hub/sitemap.xml # ============================================ # Default — allow all good crawlers # ============================================ User-agent: * Allow: / Disallow: /401 Disallow: /404 Disallow: /api/ Disallow: /*?ref= Disallow: /*?utm_ Disallow: /*?fbclid= Disallow: /*?gclid= # ============================================ # Search engines # ============================================ User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: / User-agent: Googlebot-News Allow: / User-agent: Bingbot Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / User-agent: Applebot Allow: / User-agent: YandexBot Allow: / # ============================================ # AI / LLM crawlers — explicit welcome # ============================================ # OpenAI User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Anthropic User-agent: ClaudeBot Allow: / User-agent: anthropic-ai Allow: / User-agent: Claude-Web Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / # Google AI User-agent: Google-Extended Allow: / User-agent: GoogleOther Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Apple Intelligence User-agent: Applebot-Extended Allow: / # Meta AI User-agent: meta-externalagent Allow: / User-agent: FacebookBot Allow: / # Mistral User-agent: MistralAI-User Allow: / # Cohere User-agent: cohere-ai Allow: / User-agent: cohere-training-data-crawler Allow: / # Common Crawl (used by many LLMs for training) User-agent: CCBot Allow: / # Allen AI User-agent: AI2Bot Allow: / User-agent: AI2Bot-Dolma Allow: / # You.com User-agent: YouBot Allow: / # Diffbot (used by Bing, Anthropic, others) User-agent: Diffbot Allow: / # Amazon AI (Alexa, Rufus) User-agent: Amazonbot Allow: / # Bytedance / TikTok / Doubao User-agent: Bytespider Allow: / # DeepSeek User-agent: DeepSeekBot Allow: / # Kagi search User-agent: KagiBot Allow: / # Phind User-agent: PhindBot Allow: / # Brave Search User-agent: BraveBot Allow: / # Neeva / Snowflake search User-agent: NeevaBot Allow: / # Archive (LLMs cite archived versions) User-agent: ia_archiver Allow: / User-agent: archive.org_bot Allow: / # ============================================ # SEO research tools (let them index DataDoe) # ============================================ User-agent: AhrefsBot Allow: / User-agent: SemrushBot Allow: / User-agent: rogerbot Allow: / User-agent: MJ12bot Allow: / User-agent: dotbot Allow: / # ============================================ # Block resource-heavy junk # ============================================ User-agent: PetalBot Disallow: / User-agent: ZoominfoBot Disallow: / Sitemap: https://www.datadoe.com/sitemap.xml