Robots (SEO & GEO)
The platform includes configurable Search Engine Optimization (SEO) and Generative Engine Optimization (GEO) controls. SEO manages how traditional search engines like Google and Bing index your site. GEO controls how AI systems like ChatGPT, Claude, and Gemini discover and use your content.
Search Engine Robots
Standard directives from the Robots Exclusion Protocol, honored by all major search engines. These control the meta robots tag and the /robots.txt file.
| Setting | Standard | Default | Effect |
|---|---|---|---|
| Allow Indexing | Robots Exclusion Protocol | On | When off, search engines will not add your pages to their results (noindex) |
| Allow Link Following | Robots Exclusion Protocol | On | When off, search engines will not follow links on your pages (nofollow) |
AI & LLM Controls
These settings control whether AI systems can crawl and use your site content for training or generating responses. The platform uses a defense-in-depth approach with multiple enforcement layers.
Settings
| Setting | Standard | Default | Effect |
|---|---|---|---|
| Allow AI Bots | Robots Exclusion Protocol | On | Master toggle for AI bots reading your content. When off, all known AI bots are blocked via robots.txt and the page opts out with noai + noimageai — your store stops appearing in answers from ChatGPT, Claude, Perplexity, and Google AI Overviews |
| Allow AI Indexing | Experimental | On | When off, adds noai meta tag to opt out of AI training and retrieval. Not all providers honor this. |
| Allow AI Image Use | Experimental | On | When off, adds noimageai meta tag to opt out of AI image training. Not all providers honor this. |
| Blocked AI Bots | Robots Exclusion Protocol | Empty | Block specific bots by user-agent name, even when the master toggle is on |
| llms.txt Content | Emerging (llmstxt.org) | Empty (auto-generated) | Custom content for the /llms.txt endpoint. Leave empty to auto-generate from site config |
Enforcement Layers
AI bot controls are enforced at multiple levels for maximum coverage:
- robots.txt — Dynamic, database-driven rules. Primary read-side enforcement for well-behaved crawlers.
- Meta robots tags —
index/followplus thenoaiandnoimageaidirectives in the HTML<meta>tag. Database-driven, and the layer that reaches an AI bot which ignores robots.txt. - llms.txt — Structured site description endpoint that AI systems can consume.
Known AI Bots
The platform recognizes these AI bot user-agents for per-agent robots.txt blocking. Note that most vendors run more than one crawler: a training bot that collects content to build models, and a separate search or user-initiated bot that fetches pages at the moment someone asks a question. Blocking only the training bot leaves your store readable by the assistant itself, which is why both appear below.
| Bot | Organization | Type | What it does |
|---|---|---|---|
| GPTBot | OpenAI | Training | Trains OpenAI foundation models |
| OAI-SearchBot | OpenAI | Search index | Indexes for ChatGPT search — NOT covered by blocking GPTBot |
| ChatGPT-User | OpenAI | User-initiated fetch | Fetches a page when a ChatGPT user asks about it (OpenAI says robots.txt may not apply) |
| ClaudeBot | Anthropic | Training | Trains Anthropic models |
| Claude-SearchBot | Anthropic | Search index | Indexes for Claude's search — NOT covered by blocking ClaudeBot |
| Claude-User | Anthropic | User-initiated fetch | Fetches a page when a Claude user asks about it |
| anthropic-ai | Anthropic | Training | Legacy token, kept so older saved config keeps working |
| Google-Extended | Training | Gemini training and grounding only — NOT Google Search, NOT AI Overviews | |
| Google-CloudVertexBot | Training | Crawls sites for Vertex AI grounding | |
| PerplexityBot | Perplexity | Search index | Indexes for Perplexity answers |
| Perplexity-User | Perplexity | User-initiated fetch | Fetches a page when a Perplexity user asks about it — Perplexity says it generally IGNORES robots.txt |
| Meta-ExternalAgent | Meta | Training | Trains Meta AI models |
| Meta-WebIndexer | Meta | Search index | Builds Meta AI's search index |
| Meta-ExternalFetcher | Meta | User-initiated fetch | Fetches a page for a Meta AI request — Meta says it MAY bypass robots.txt |
| FacebookBot | Meta | Training | Legacy Meta crawler token |
| Applebot-Extended | Apple | Training | Apple Intelligence training — does NOT affect Siri or Spotlight search |
| CCBot | Common Crawl | Open dataset | Open dataset that many other models train on |
| Bytespider | ByteDance | Training | ByteDance / TikTok model training |
| Amazonbot | Amazon | Training | Alexa and Amazon AI services |
Google-Extended only controls Gemini AI training and grounding. Blocking it does not affect your store's Google Search rankings or indexing — and it does not remove you from Google AI Overviews, which are generated from Google's ordinary search index. The only setting that changes what AI Overviews can use is Allow Indexing, and turning that off removes your store from Google Search entirely.ChatGPT-User, Perplexity says Perplexity-User "generally ignores" them, and Meta says Meta-ExternalFetcher "may bypass" them. Anthropic documents Claude-User as honoring them. So blocking AI crawlers reliably stops your content being ingested and indexed; it does not guarantee nobody can pull up your store through an assistant.llms.txt
The /llms.txt endpoint provides a structured text description of your site that AI systems can read. This is an emerging standard (see llmstxt.org) that helps AI chatbots accurately answer questions about your business.
If you leave the llms.txt content field empty, the platform auto-generates content from your site name, description, location, and contact information. For best results, write custom content that includes:
- A clear description of what your business does
- Your location and service area
- Products and services offered
- Contact information
- FAQ section answering common customer questions