The four ways AI gets its information are training, retrieving live online information, leveraging licensing partnerships, and user-initiated actions to pull data:
1. Training
AI companies train their models on data to help those models learn to understand and generate content. This training data is the foundational knowledge layer that models draw on to answer prompts.
AI companies collect LLM training data from sources like:
Models have knowledge cutoff dates, after which they’re no longer trained on new data. To reduce the risk of outdated answers, AI systems often draw on additional sources of information that I’ll touch on later.
For example, when I asked ChatGPT how much rain Hawaii got from Hurricane Lala with web search turned off, ChatGPT replied that it was unaware of that hurricane.
| Tool | Free Requests | Additional Data |
| Semrush Keyword Checker | 5/day | Keyword difficulty, search intent, cost per click (CPC), related keywords, SERP overview |
| SE Ranking | 5/day | Keyword difficulty, CPC, SERP overview, related keywords, ad history |
| Searchvolume.io | Unlimited | None |
| SpyFu | Unlimited | Keyword difficulty, CPC, ad history, related keywords, backlinks, ranking history |
| Google Keyword Planner | Unlimited | CPC, ad competition, trends |
AI systems retrieve live online information to supplement training data and improve their answers’ accuracy.
For example, ChatGPT retrieved live information from various sites to provide me with the latest U.S. stock market news:
AI companies use proprietary crawlers to fetch live online data. They also get live web information through search engines, which crawl and index content independently.
This table lists some of AI systems’ known sources of live online data:
How to help AI systems retrieve your site content
Help AI systems retrieve your site content by allowing crawlers from AI companies and search engines to access your website by verifying that your robots.txt file doesn’t have disallow rules targeting these crawlers.
Next, use Google Search Console and Bing Webmaster Tools to check if Google and Bing have indexed your pages. If they haven’t, AI companies won’t be able to retrieve these pages’ content from the search engines’ indexes.
In Google Search Console, for example, click the drop-down menu under “Indexing” in the left sidebar and select “Pages” to view the number of pages Google has and hasn’t indexed.
You can also view specific URLs that are and aren’t indexed. To view your indexed pages, click “View data about indexed pages.”
Several AI companies have struck licensing partnerships allowing their AI systems to use partner organizations’ content for training and live retrieval.
For example, OpenAI and Google have partnered with Reddit and Stack Overflow. OpenAI also licenses content from organizations like Yelp and Time.
Content from partner organizations often appears prominently in AI systems’ answers. In this conversation, for instance, ChatGPT heavily cited Reddit when I asked for recommendations for “cheap eats in austin texas according to locals”:
How to improve your presence in partner companies’ content
Improve your presence in partner companies’ content by identifying your target AI system’s partner companies, then taking concrete steps to appear in their content.
Let’s say you want to grow your visibility on ChatGPT. Since OpenAI licenses content from Reddit, Yelp, and Time, you could improve your presence in these companies’ content by participating in Reddit discussions, creating a Yelp directory listing, and pitching journalists at Time for coverage.
As much as possible, ensure the content about your brand on third-party sites uses the language and framing you want AI systems to adopt when talking about you.