Skip to content

Robots (SEO & GEO)

The platform includes configurable Search Engine Optimization (SEO) and Generative Engine Optimization (GEO) controls. SEO manages how traditional search engines like Google and Bing index your site. GEO controls how AI systems like ChatGPT, Claude, and Gemini discover and use your content.

All robot controls are configured in Admin > Site Config > Robots and are restricted to Super admins.

Search Engine Robots

Standard directives from the Robots Exclusion Protocol, honored by all major search engines. These control the meta robots tag and the /robots.txt file.

SettingStandardDefaultEffect
Allow IndexingRobots Exclusion ProtocolOnWhen off, search engines will not add your pages to their results (noindex)
Allow Link FollowingRobots Exclusion ProtocolOnWhen off, search engines will not follow links on your pages (nofollow)

AI & LLM Controls

These settings control whether AI systems can crawl and use your site content for training or generating responses. The platform uses a defense-in-depth approach with multiple enforcement layers.

Settings

SettingStandardDefaultEffect
Allow AI BotsRobots Exclusion ProtocolOnMaster toggle for AI bots reading your content. When off, all known AI bots are blocked via robots.txt and the page opts out with noai + noimageai — your store stops appearing in answers from ChatGPT, Claude, Perplexity, and Google AI Overviews
Allow AI IndexingExperimentalOnWhen off, adds noai meta tag to opt out of AI training and retrieval. Not all providers honor this.
Allow AI Image UseExperimentalOnWhen off, adds noimageai meta tag to opt out of AI image training. Not all providers honor this.
Blocked AI BotsRobots Exclusion ProtocolEmptyBlock specific bots by user-agent name, even when the master toggle is on
llms.txt ContentEmerging (llmstxt.org)Empty (auto-generated)Custom content for the /llms.txt endpoint. Leave empty to auto-generate from site config

Enforcement Layers

AI bot controls are enforced at multiple levels for maximum coverage:

  1. robots.txt — Dynamic, database-driven rules. Primary read-side enforcement for well-behaved crawlers.
  2. Meta robots tags — index/follow plus the noai and noimageai directives in the HTML <meta> tag. Database-driven, and the layer that reaches an AI bot which ignores robots.txt.
  3. llms.txt — Structured site description endpoint that AI systems can consume.

Known AI Bots

The platform recognizes these AI bot user-agents for per-agent robots.txt blocking. Note that most vendors run more than one crawler: a training bot that collects content to build models, and a separate search or user-initiated bot that fetches pages at the moment someone asks a question. Blocking only the training bot leaves your store readable by the assistant itself, which is why both appear below.

BotOrganizationTypeWhat it does
GPTBotOpenAITrainingTrains OpenAI foundation models
OAI-SearchBotOpenAISearch indexIndexes for ChatGPT search — NOT covered by blocking GPTBot
ChatGPT-UserOpenAIUser-initiated fetchFetches a page when a ChatGPT user asks about it (OpenAI says robots.txt may not apply)
ClaudeBotAnthropicTrainingTrains Anthropic models
Claude-SearchBotAnthropicSearch indexIndexes for Claude's search — NOT covered by blocking ClaudeBot
Claude-UserAnthropicUser-initiated fetchFetches a page when a Claude user asks about it
anthropic-aiAnthropicTrainingLegacy token, kept so older saved config keeps working
Google-ExtendedGoogleTrainingGemini training and grounding only — NOT Google Search, NOT AI Overviews
Google-CloudVertexBotGoogleTrainingCrawls sites for Vertex AI grounding
PerplexityBotPerplexitySearch indexIndexes for Perplexity answers
Perplexity-UserPerplexityUser-initiated fetchFetches a page when a Perplexity user asks about it — Perplexity says it generally IGNORES robots.txt
Meta-ExternalAgentMetaTrainingTrains Meta AI models
Meta-WebIndexerMetaSearch indexBuilds Meta AI's search index
Meta-ExternalFetcherMetaUser-initiated fetchFetches a page for a Meta AI request — Meta says it MAY bypass robots.txt
FacebookBotMetaTrainingLegacy Meta crawler token
Applebot-ExtendedAppleTrainingApple Intelligence training — does NOT affect Siri or Spotlight search
CCBotCommon CrawlOpen datasetOpen dataset that many other models train on
BytespiderByteDanceTrainingByteDance / TikTok model training
AmazonbotAmazonTrainingAlexa and Amazon AI services
Google-Extended only controls Gemini AI training and grounding. Blocking it does not affect your store's Google Search rankings or indexing — and it does not remove you from Google AI Overviews, which are generated from Google's ordinary search index. The only setting that changes what AI Overviews can use is Allow Indexing, and turning that off removes your store from Google Search entirely.
The user-initiated fetch bots are the exception to all of this. When a person pastes your link into an assistant, several vendors fetch the page regardless of robots.txt: OpenAI says the rules "may not apply" to ChatGPT-User, Perplexity says Perplexity-User "generally ignores" them, and Meta says Meta-ExternalFetcher "may bypass" them. Anthropic documents Claude-User as honoring them. So blocking AI crawlers reliably stops your content being ingested and indexed; it does not guarantee nobody can pull up your store through an assistant.

llms.txt

The /llms.txt endpoint provides a structured text description of your site that AI systems can read. This is an emerging standard (see llmstxt.org) that helps AI chatbots accurately answer questions about your business.

If you leave the llms.txt content field empty, the platform auto-generates content from your site name, description, location, and contact information. For best results, write custom content that includes:

  • A clear description of what your business does
  • Your location and service area
  • Products and services offered
  • Contact information
  • FAQ section answering common customer questions
The FAQ section is particularly valuable for GEO. AI chatbots answer user questions — so structuring your content as Q&A pairs maps directly to how your business appears in AI-generated responses.