Integration

ChatGPT and Foundation CMS

People now ask ChatGPT what they used to type into a search box. Foundation CMS sites are set up to be found there: OpenAI's search and answer crawlers are allowed by default, its training crawler is refused, and every publish can be pushed to Bing, whose index ChatGPT search answers from.

Overview

This integration connects Foundation CMS with ChatGPT. Foundation CMS is Swarm Labs' multi-tenant content management system on Cloudflare Workers. It holds a site's content and hands it to the site as whole, typed records, and every site runs as its own Worker with its own database and media bucket. Foundation CMS keeps one register of crawlers and treats them by what they do. OpenAI runs three, each independent: OAI-SearchBot indexes pages so ChatGPT can cite them, ChatGPT-User fetches a page because someone asked about it, and GPTBot collects pages to train models.

Business Context and Core Use Case

A business wants to be cited when someone asks ChatGPT for a supplier, but may not want its copy used to train a model. Those used to be one decision; they are not. The Foundation CMS default is the one a client would choose if asked: answer crawlers in, training crawlers out, per site and changeable.

The Applications Involved

Foundation CMS (Foundation CMS) writes each site's robots rules from its crawler policy, adds a content signal line, and pushes publishes to Bing through IndexNow.

ChatGPT (ChatGPT) finds and cites pages through its search crawler and fetches them for people who ask.

How the Integration Works

Each site's robots.txt is generated from its policy: OAI-SearchBot and ChatGPT-User are allowed, GPTBot is refused, and a Content-Signal line states search=yes, ai-input=yes, ai-train=no. When an editor publishes, a site with IndexNow switched on pushes the address to Bing, and ChatGPT search answers from Bing's index. Verifying the site in Bing Webmaster Tools is a setting too. Structured data that matches what is visible, and an llms.txt built from the content, round it off.

Immediate Operational Value

New pages can reach AI answers in minutes rather than waiting for a crawl, and nobody has to maintain a robots file by hand as new crawlers appear. Each crawler is listed by name with what allowing it gets the client, so the choice is informed rather than guessed.

Security, Access, and Governance

Crawler policy is per site and set by administrators. Refusing a crawler is a named rule in robots.txt; the register lists bots explicitly rather than guessing from user-agent strings, so a renamed bot shows up as a change somebody can see.

Summary

Foundation CMS sites let ChatGPT find and cite them while keeping their content out of training by default, and tell Bing, and through it ChatGPT search, the moment something new is published.

Frequently asked questions

Can ChatGPT cite a Foundation CMS site?

Yes. OAI-SearchBot, which indexes pages for ChatGPT's answers, and ChatGPT-User, which fetches pages for people, are allowed by default.

Is the site's content used to train OpenAI's models?

Not by default. GPTBot, OpenAI's training crawler, is refused unless the site's administrator allows it.

How quickly does ChatGPT see a new page?

With IndexNow on, Bing is told at publish time, and ChatGPT search answers from Bing's index, often within minutes.

Can we change these defaults?

Yes, per site and per crawler, in the site's search settings.

Want ChatGPT and Foundation CMS
wired up for you?