Deep dive

Public posts.
Different paths to an answer.

Social content can reach an AI answer through search results, a page fetch or a platform agreement. Those routes expose different parts of a platform.

Platform rules fetched and reviewed . Open source details for individual fetch times.

01 · Follow the source

“Scraping social media” describes several different things

Search results and snippets
An assistant can receive a URL, title or excerpt from a search index. That does not establish that it opened the post, read every comment or watched the video. An indexed copy may be older than the current post.
A direct page request
A crawler or user-request tool can request a public URL. The response may contain content, a login page, a rate limit or an access challenge. Public visibility in your browser does not establish automated access.
A licensed API or other integration
A platform can supply structured content under a separate agreement. Reddit’s announced partnerships are examples. A robots.txt restriction on public web crawling does not describe the rights or coverage of that API.
Earlier model training
A model may answer from learned knowledge without retrieving the post now. A correct quotation or a familiar brand name alone does not identify the training source or prove current platform access.

The provider comparison explains documented search and training roles. For a particular answer, inspect its cited URL and the content actually available there. A citation to an article quoting Reddit is different from a citation to the original thread.

02 · Search suppliers

Does ChatGPT use Bing instead of Google?

Microsoft announced Bing as ChatGPT’s default search experience in May 2023. That is a dated integration announcement. It does not establish that ChatGPT previously relied on Google, switched away from it, or uses Bing exclusively today.

OpenAI also documents its own OAI-SearchBot for search, GPTBot for potential training use and ChatGPT-User for user-initiated requests. A search supplier is one part of this system; it does not identify every route an answer can use.

Access does not transfer automatically

If Google or Bing can index a platform page, that alone does not establish what an assistant receives from a search service, whether it can open the page itself, or whether it has permission to train on it. We have not established a complete current supplier list for every ChatGPT search mode.

03 · Platform by platform

What the available evidence tells us

Reddit and other forums

Reddit announced expanded Google Data API access in February 2024, including public posts and comments, and OpenAI Data API access in May 2024 to bring Reddit content into ChatGPT and other products. These announcements describe partnerships at those dates; they do not expose every current contract term or guarantee that a particular answer uses the feed.

Our September 2026 fetch of Reddit’s robots.txt returned a wildcard Disallow: /. Read that alongside the API announcements: public crawl rules and partner supply are different routes. Neither establishes access to private communities or deleted material.

Other forums need individual checks. A public discussion, a members-only thread and an excerpt republished elsewhere have different access conditions. Do not apply Reddit’s agreements to independent forums.

LinkedIn

LinkedIn says its public profile is a simplified version whose visible sections depend on the member’s settings. That public version can appear in Google and Bing. This does not establish access to a member’s full profile, feed, messages or private groups.

The robots snapshot gives OAI-SearchBot and Claude-SearchBot individual path rules, while GPTBot and ClaudeBot receive a whole-site disallow. LinkedIn’s published treatment of search and training bots therefore differs. A path rule still needs to be evaluated against the specific URL and actual response.

Facebook

The fetched rules give Googlebot, Bingbot, GPTBot, ClaudeBot and PerplexityBot individual path restrictions. OAI-SearchBot and Claude-SearchBot fall under the wildcard whole-site disallow in this snapshot. Facebook’s robots.txt also says automated collection requires express written permission.

These rules do not establish which public Pages or posts an assistant successfully reads. They provide no evidence of access to private groups, messages or a logged-in feed. Facebook’s own AI features do not, by themselves, establish third-party access.

Instagram

Instagram’s fetched rules differ from Facebook’s: Googlebot and Bingbot have path rules, while all six assistant-related tokens in our table receive a whole-site disallow, either by name or through the wildcard group. Its robots.txt also requires express written permission for automated collection.

A search result containing a profile name or caption is not evidence that the assistant retrieved the complete post, comments, image or video. We have not verified a general third-party API agreement granting these assistants access to Instagram content.

TikTok

TikTok’s fetched robots.txt explicitly groups OAI-SearchBot, GPTBot, Claude-SearchBot, ClaudeBot, PerplexityBot and ChatGPT-User under Disallow: /. Googlebot instead falls under general path rules; Bingbot has its own path restrictions.

That distinction leaves open a search-index route for some URLs, subject to the applicable rules and actual access. It does not prove a particular video was indexed or that an assistant can fetch its video, audio, transcript or comments. Those are separate content types to verify.

Unknown means we have not established access

For these five platforms, this review does not establish a complete provider-by-provider map of successful retrieval, private integrations or training datasets. For Grok’s separately documented public X search, see the provider comparison.

04 · A dated observation

Published crawler rules by platform

This table reports the robots.txt responses we fetched, not a test of which assistants can read each platform. It covers selected bot names on the five www hosts. APIs, other subdomains and logged-in access are outside its scope.

  • Disallow / — the selected group requests no crawling anywhere on this host.
  • Path rules — individual URL rules apply. This is not blanket permission or proof of access.
  • Named / wildcard — whether the bot has a named group or falls back to User-agent: *.

Scroll the table horizontally on smaller screens.

Robots.txt snapshot · 12 Sept 2026
Bot / documented roleRedditLinkedInFacebookInstagramTikTok
GooglebotGoogle Search Disallow /Wildcard groupPath rulesNamed groupPath rulesNamed groupPath rulesNamed groupPath rulesWildcard group
BingbotBing Search Disallow /Wildcard groupPath rulesNamed groupPath rulesNamed groupPath rulesNamed groupPath rulesNamed group
OAI-SearchBotOpenAI search Disallow /Wildcard groupPath rulesNamed groupDisallow /Wildcard groupDisallow /Wildcard groupDisallow /Named group
GPTBotOpenAI training Disallow /Wildcard groupDisallow /Named groupPath rulesNamed groupDisallow /Named groupDisallow /Named group
Claude-SearchBotAnthropic search Disallow /Wildcard groupPath rulesNamed groupDisallow /Wildcard groupDisallow /Wildcard groupDisallow /Named group
ClaudeBotAnthropic training Disallow /Wildcard groupDisallow /Named groupPath rulesNamed groupDisallow /Named groupDisallow /Named group
PerplexityBotPerplexity search Disallow /Wildcard groupDisallow /Named groupPath rulesNamed groupDisallow /Named groupDisallow /Named group
ChatGPT-UserUser-initiated requests Disallow /Wildcard groupDisallow /Named groupDisallow /Wildcard groupDisallow /Wildcard groupDisallow /Named group

Bot purposes and compliance policies are documented in the crawler reference. OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User requests. Its row here records the platform’s published signal, not an assurance that the tool follows it or can get through.

05 · What this means for your website

Make the original information easy to verify

Choose platforms where customers ask questions you can answer. Keep service details, original evidence and contact information on accessible pages on your own site too.

  • Keep business names, profile descriptions and website links accurate across platforms. Connect them with the business identity and trust signals on your website.
  • Contribute useful first-hand answers to relevant discussions. Disclose your connection to a business and follow the community’s rules.
  • For a useful video, consider publishing an accessible written explanation on your site. Do not assume every search tool can extract its audio or images.
  • Follow the Quick guide’s measurement routine, noting whether each citation points to your website, your post, an excerpt or someone else’s discussion.

Forum mentions and business profiles can supply evidence for people to assess. Their existence alone does not establish a ranking benefit in every assistant.

06 · Sources and limits

What we fetched, and when

On 12 September 2026, we made an unauthenticated GET request to each linked robots.txt URL using DiscoveryFiles-DocumentationReview/1.0. All five returned HTTP 200. We reviewed the named groups for the selected tokens and used the wildcard group where no named group existed.

Responses can vary by time, host, request location or user-agent. We did not impersonate the named crawlers or test their access to posts. Authentication, platform terms, technical enforcement and separate agreements still matter. A robots.txt rule neither grants training rights nor reveals what a model was trained on.

For documentation and announcement date labels, see how we date sources.