Meta allegedly building own search engine to avoid training Google's AI
Meta reportedly engaged in heavy web scraping using bots named "meta-externalagent" and "meta-webindexer" with significantly higher access frequency than competitors. According to a former Meta employee, the company is attempting to acquire large amounts of data for training new AI models with parameters in the range of 5–10 trillion. It remains unclear whether Meta is primarily preparing its own search engine or focusing on data collection for its AI systems.
If Meta has started heavily crawling the internet and their bots named "meta-externalagent" and "meta-webindexer" are making orders of magnitude more requests than bots from Google, X, and others, something is definitely happening. Rumors about Meta planning its own search engine have circulated before but never materialized. The question is what the target of this sudden activity is: either they need data for their AI, which they allegedly plan to make accessible to everyone, or it's truly preparation for a search engine.
Why would Meta want to build its own search engine instead of using Google?
Building its own search engine would allow Meta to control its web index for AI training without risking Google using that data for its own training. Meta would gain independence from other search engines and competitive advantage.
What is the typical parameter range for modern large language models?
Current state-of-the-art models have parameters in the tens to hundreds of billions range. Meta's ambition to train models with 5–10 trillion parameters would represent a new class of AI systems requiring massive amounts of high-quality training data.
How can web scraper bots be identified?
Web scraper bots are identified primarily by their user-agent identity in HTTP requests and access patterns. Bots typically identify themselves by name (e.g., "Googlebot", "meta-webindexer"), though they may also be masked. Anomalies in access patterns, such as unusual request volumes or targeting specific services, can indicate unknown or unauthorized scraping.
- Chrome fails to delete user data from Google sites — lapcatsoftware.com 57 % match
- ChatGPT Work can now mimic your writing style — x.com 55 % match
- MatrAIx: Harvard and MIT researchers simulate 8.3 billion personas to evaluate AI systems and digital products — matraix.ai 54 % match