# If the Joomla site is installed within a folder # eg www.example.com/joomla/ then the robots.txt file # MUST be moved to the site root # eg www.example.com/robots.txt # AND the joomla folder name MUST be prefixed to all of the # paths. # eg the Disallow rule for the /administrator/ folder MUST # be changed to read # Disallow: /joomla/administrator/ # # For more information about the robots.txt standard, see: # https://www.robotstxt.org/orig.html User-agent: * Disallow: /administrator/ Disallow: /api/ Disallow: /bin/ Disallow: /cache/ Disallow: /cli/ Disallow: /components/ Disallow: /includes/ Disallow: /installation/ Disallow: /language/ Disallow: /layouts/ Disallow: /libraries/ Disallow: /logs/ Disallow: /modules/ Disallow: /plugins/ Disallow: /tmp/ # ========================================================================= # Blocco crawler AI / LLM (training e scraping). Disallow totale. # Nota: Google-Extended / Applebot-Extended NON bloccano la Ricerca, # solo l'uso dei contenuti per addestramento AI. # meta-webindexer = crawler dell'indice web di Meta AI (aggiunto 2026-07-15: # il 14-15/07 ha saturato il server strisciando le pagine profonde). # ========================================================================= User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: ClaudeBot User-agent: Claude-Web User-agent: anthropic-ai User-agent: Claude-SearchBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Bytespider User-agent: Amazonbot User-agent: Meta-ExternalAgent User-agent: Meta-ExternalFetcher User-agent: meta-webindexer User-agent: FacebookBot User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Diffbot User-agent: Omgilibot User-agent: Omgili User-agent: ImagesiftBot User-agent: YouBot User-agent: Timpibot User-agent: DuckAssistBot User-agent: PetalBot User-agent: AI2Bot Disallow: / # ========================================================================= # Applebot (ricerca Siri/Spotlight): resta AMMESSO sulle pagine di sintesi # (anno, mese), ma NON sulle pagine profonde a 6+ segmenti (giorno, # settimana, elenco mese): quello spazio URL e' quasi infinito (24 lingue x # paesi x anni x mesi x giorni) e il crawl massivo del 14-15/07/2026 ha # saturato il server. Il gruppo dedicato sostituisce per Applebot il gruppo # "*": i Disallow di base sono replicati qui. # ========================================================================= User-agent: Applebot Crawl-delay: 30 Disallow: /administrator/ Disallow: /api/ Disallow: /bin/ Disallow: /cache/ Disallow: /cli/ Disallow: /components/ Disallow: /includes/ Disallow: /installation/ Disallow: /language/ Disallow: /layouts/ Disallow: /libraries/ Disallow: /logs/ Disallow: /modules/ Disallow: /plugins/ Disallow: /tmp/ Disallow: /*/*/*/*/*/* Sitemap: https://www.onlinecalendar.pro/calpro/calpro-sitemap.xml