User-agent: * # --- Allow public sections --- Allow: / Allow: /logbook$ # --- Block sensitive or user-only areas --- Disallow: /admin/ Disallow: /dashboard Disallow: /profile Disallow: /logbook/ Disallow: /lists Disallow: /plans Disallow: /settings # Auth flows already noindex via meta robots, but listing # them here saves crawl budget on uninteresting pages. Disallow: /login Disallow: /register Disallow: /forgot-password Disallow: /reset-password Disallow: /verify-email # Scraper honeypot (HoneypotController). Disallowed so any crawler # that honours robots.txt never follows the hidden link; only # clients ignoring these rules walk into it. Disallow: /properties/ # --- Unreleased: embeddable widgets --- # /widgets also carries meta robots noindex. /embed/ is the JSON # the widget script fetches plus widget.js itself — no crawler has # any reason to want either, and blocking them now costs nothing # because nothing is embedded anywhere yet. Note the backlink in # the snippet is static markup in the HOST page, so blocking # widget.js here does not weaken it when this does launch. Disallow: /widgets Disallow: /embed/ # --- Block API & data endpoints from all crawlers --- Disallow: /api/ Disallow: /dive-map/pins Disallow: /dive-map/region-search Disallow: /dive-map/community-spots # /posts/* is POST-only (like / comment / delete). Googlebot hitting # the URL with GET returns 405 and clutters GSC as a 4xx issue. Disallow: /posts/ # /dive-sites/{slug}/lists is PATCH-only (sync save-to-list state), # /dive-sites/{slug}/save is POST/DELETE-only. The URLs are exposed # in the page source via route() helpers inside JS fetches, which # link previewers and bots probe with GET. We catch those server- # side and 301 to the canonical site URL, but disallowing here # saves crawl budget and keeps GSC clean. Disallow: /dive-sites/*/lists Disallow: /dive-sites/*/save # --- Prevent duplicate content from query parameters --- Disallow: /*?lat= Disallow: /*?lng= Disallow: /*?region= Disallow: /*?query= Disallow: /*?search= # --- Block AI/LLM TRAINING crawlers --- # Answer-engine/citation agents (ChatGPT-User, OAI-SearchBot, # PerplexityBot, Claude-Web) are deliberately ALLOWED: they fetch # pages to answer a live user question with a link back, and they # are the audience for /llms.txt and /llms-full.txt. They fall # through to the User-agent: * rules above. Keep this list in # sync with app/Http/Middleware/BlockBots.php, which enforces it. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Anthropic-AI Disallow: / User-agent: Google-Extended Disallow: / User-agent: GoogleOther Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: Diffbot Disallow: / User-agent: cohere-ai Disallow: / # --- Block SEO/backlink crawlers --- # These are not search engines. They crawl to resell backlink and # keyword data to whoever pays, and they send no visitors: over 17 # days of August 2026 they took 7,568 requests and were served # 7,491 pages between them, against zero referred traffic. # # PetalBot is Huawei Petal Search — a real engine, but with no # meaningful Australian audience for a Sydney dive site, and the # second-heaviest crawler on the whole site at 2,717 requests. # # Amazonbot is deliberately NOT here. It feeds Alexa answers, which # is a plausible way somebody asks whether a beach is diveable. # Bingbot, Googlebot, Applebot and facebookexternalhit stay too: # real search, Siri/Spotlight (which matters with an iOS app), and # link previews when a diver shares a site to a group chat. User-agent: PetalBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / # Seen as SERankingBacklinksBot; the token is the stable part. User-agent: SERanking Disallow: / User-agent: DataForSeoBot Disallow: / # --- Sitemap for crawlers --- Sitemap: https://vizzbud.com/sitemap.xml