Friday, Jul 24, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

TECH

The Claim

Publicly posted online content can be scraped and used to train artificial intelligence models.

The Short Version

The claim is accurate as a statement of technical capability and widespread industry practice. Publicly posted online content is routinely scraped to train AI models—confirmed by academic research, corporate disclosures (e.g., Google's privacy policy), and the existence of major datasets like Common Crawl. However, the claim omits critical legal context: copyright law, privacy regulations, terms of service, and the EU AI Act (fully enforced in 2026) all impose significant restrictions. "Can be done" is true; "can be done freely and lawfully in all cases" is not.

Caveats

  • 'Publicly posted' does not mean 'free to use'—most online content is copyrighted, and scraping it for AI training is the subject of 70+ active lawsuits as of early 2026.
  • The EU AI Act now requires AI developers to disclose training data sources, respect copyright opt-outs, and comply with transparency obligations—scraping without compliance steps may be unlawful.
  • Website terms of service, technical access controls, and computer-access laws (e.g., CFAA in the U.S.) can make automated scraping illegal even when content is publicly viewable in a browser.

The Receipts

  1. Ethical AI Scraping in 2026: Navigating the Legal Landscape | Use Apify

    Use Apify

  2. EU AI Act 2026: New Rules for Training Data and Copyright - Scalevise

    Scalevise

  3. Training Data or Taking Data? How AI Copyright Lawsuits Are Reshaping Creative Rights

    Training Data or Taking Data? How AI Copyright Lawsuits Are Reshaping Creative Rights

  4. Unveiling the Legal Battle: OpenAI Faces Lawsuit Over Data Collection Practices

    Unveiling the Legal Battle: OpenAI Faces Lawsuit Over Data Collection Practices

  5. Mindfully Training AI Models Using Public Data - i2Coalition

    i2Coalition

  6. The EU AI Act and copyrights compliance - IAPP

    IAPP

  7. Generative AI Training and Copyright Law - arXiv

    arXiv

  8. Harvard's Library Innovation Lab Launches Institutional Data Initiative

    Harvard Law School

  9. AI's legal frontier: What Europe's privacy regulators say about scraping personal data - Zyte

    Zyte

  10. Google's New Privacy Policy Confirms AI Data Scraping - Thurrott.com

    Thurrott.com

+ 14 more sources — see the full list on Lenz

Filed Under

Artificial Intelligence Models

More Fact Checks