The Short Version
The claim is accurate as a statement of technical capability and widespread industry practice. Publicly posted online content is routinely scraped to train AI models—confirmed by academic research, corporate disclosures (e.g., Google's privacy policy), and the existence of major datasets like Common Crawl. However, the claim omits critical legal context: copyright law, privacy regulations, terms of service, and the EU AI Act (fully enforced in 2026) all impose significant restrictions. "Can be done" is true; "can be done freely and lawfully in all cases" is not.