01 · Follow the source
“Scraping social media” describes several different things
- Search results and snippets
- An assistant can receive a URL, title or excerpt from a search index. That does not establish that it opened the post, read every comment or watched the video. An indexed copy may be older than the current post.
- A direct page request
- A crawler or user-request tool can request a public URL. The response may contain content, a login page, a rate limit or an access challenge. Public visibility in your browser does not establish automated access.
- A licensed API or other integration
- A platform can supply structured content under a separate agreement. Reddit’s announced partnerships are examples. A robots.txt restriction on public web crawling does not describe the rights or coverage of that API.
- Earlier model training
- A model may answer from learned knowledge without retrieving the post now. A correct quotation or a familiar brand name alone does not identify the training source or prove current platform access.
The provider comparison explains documented search and training roles. For a particular answer, inspect its cited URL and the content actually available there. A citation to an article quoting Reddit is different from a citation to the original thread.
02 · Search suppliers
Does ChatGPT use Bing instead of Google?
Microsoft announced Bing as ChatGPT’s default search experience in May 2023. That is a dated integration announcement. It does not establish that ChatGPT previously relied on Google, switched away from it, or uses Bing exclusively today.
OpenAI also documents its own OAI-SearchBot for search, GPTBot for potential training use and ChatGPT-User for user-initiated requests. A search supplier is one part of this system; it does not identify every route an answer can use.
Access does not transfer automatically
If Google or Bing can index a platform page, that alone does not establish what an assistant receives from a search service, whether it can open the page itself, or whether it has permission to train on it. We have not established a complete current supplier list for every ChatGPT search mode.
04 · A dated observation
Published crawler rules by platform
This table reports the robots.txt responses we fetched, not a test of which assistants can read each platform. It covers selected bot names on the five www hosts. APIs, other subdomains and logged-in access are outside its scope.
- Disallow / — the selected group requests no crawling anywhere on this host.
- Path rules — individual URL rules apply. This is not blanket permission or proof of access.
- Named / wildcard — whether the bot has a named group or falls back to
User-agent: *.
Scroll the table horizontally on smaller screens.
Robots.txt snapshot · 12 Sept 2026 | Bot / documented role | Reddit | LinkedIn | Facebook | Instagram | TikTok |
GooglebotGoogle Search | Disallow /Wildcard group | Path rulesNamed group | Path rulesNamed group | Path rulesNamed group | Path rulesWildcard group |
BingbotBing Search | Disallow /Wildcard group | Path rulesNamed group | Path rulesNamed group | Path rulesNamed group | Path rulesNamed group |
OAI-SearchBotOpenAI search | Disallow /Wildcard group | Path rulesNamed group | Disallow /Wildcard group | Disallow /Wildcard group | Disallow /Named group |
GPTBotOpenAI training | Disallow /Wildcard group | Disallow /Named group | Path rulesNamed group | Disallow /Named group | Disallow /Named group |
Claude-SearchBotAnthropic search | Disallow /Wildcard group | Path rulesNamed group | Disallow /Wildcard group | Disallow /Wildcard group | Disallow /Named group |
ClaudeBotAnthropic training | Disallow /Wildcard group | Disallow /Named group | Path rulesNamed group | Disallow /Named group | Disallow /Named group |
PerplexityBotPerplexity search | Disallow /Wildcard group | Disallow /Named group | Path rulesNamed group | Disallow /Named group | Disallow /Named group |
ChatGPT-UserUser-initiated requests | Disallow /Wildcard group | Disallow /Named group | Disallow /Wildcard group | Disallow /Wildcard group | Disallow /Named group |
Bot purposes and compliance policies are documented in the crawler reference. OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User requests. Its row here records the platform’s published signal, not an assurance that the tool follows it or can get through.
05 · What this means for your website
Make the original information easy to verify
Choose platforms where customers ask questions you can answer. Keep service details, original evidence and contact information on accessible pages on your own site too.
- Keep business names, profile descriptions and website links accurate across platforms. Connect them with the business identity and trust signals on your website.
- Contribute useful first-hand answers to relevant discussions. Disclose your connection to a business and follow the community’s rules.
- For a useful video, consider publishing an accessible written explanation on your site. Do not assume every search tool can extract its audio or images.
- Follow the Quick guide’s measurement routine, noting whether each citation points to your website, your post, an excerpt or someone else’s discussion.
Forum mentions and business profiles can supply evidence for people to assess. Their existence alone does not establish a ranking benefit in every assistant.
06 · Sources and limits
What we fetched, and when
On 12 September 2026, we made an unauthenticated GET request to each linked robots.txt URL using DiscoveryFiles-DocumentationReview/1.0. All five returned HTTP 200. We reviewed the named groups for the selected tokens and used the wildcard group where no named group existed.
Responses can vary by time, host, request location or user-agent. We did not impersonate the named crawlers or test their access to posts. Authentication, platform terms, technical enforcement and separate agreements still matter. A robots.txt rule neither grants training rights nor reveals what a model was trained on.
For documentation and announcement date labels, see how we date sources.