Close Menu
GeekBlog

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    The NSA Wants You to Reboot Your Router. A Reboot Only Fixes Half the Problem.

    August 25, 2026

    One in Three Web Pages Made Since ChatGPT Was Not Written by a Human

    August 25, 2026

    2026 Is the Worst Year for Phone Sales Since 2013. Your Spec Sheet Is Paying for It.

    August 25, 2026
    Facebook X (Twitter) Instagram Threads
    GeekBlog
    • Home
    • Mobile
    • Tech News
    • Blog
    • How-To Guides
    • AI & Software
    Facebook
    GeekBlog
    Home»Tech News»One in Three Web Pages Made Since ChatGPT Was Not Written by a Human
    Tech News

    One in Three Web Pages Made Since ChatGPT Was Not Written by a Human

    Olivia HartmanBy Olivia HartmanAugust 25, 20269 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    For about a decade, “dead internet theory” was the kind of thing you rolled your eyes at. It started on message boards like 4chan, it had no data behind it, and its central claim sounded paranoid: that most of what you read online was already written by machines, and that the humans were quietly becoming a minority in their own network.

    Pew Research Center just went and counted.

    The result is not a vindication of the conspiracy version, which involved shadowy coordination and a web that had already “died” years ago. It is something stranger and more useful. The measurable share of AI written pages on the open web is still small. The share among pages published since ChatGPT arrived is not small at all.

    The short version

    • What Pew did: sampled roughly 490,000 English language web pages from the Common Crawl archive, 10,000 pages from each of 49 crawls between January 2021 and July 2026
    • The headline number: 10% of pages in the July 2026 snapshot show significant signs of AI authorship
    • The number that actually matters: restrict the sample to pages published after ChatGPT launched in November 2022 and it jumps to 35%
    • The trend: that figure sat near 5% in mid 2025. It doubled in roughly a year.
    • Where it concentrates: .com domains run about 10 times the rate of .edu and .gov, and roughly double .org
    • The tool: an open weight detection model from Pangram, which is good but not perfect, and that caveat matters
    • The other half of the story: separate industry data puts malicious bot traffic at 40% of all web requests, so the reading side of the internet is automating too

    What Pew actually measured, and what it did not

    The methodology is worth understanding, because almost every headline about this study flattens it into one scary percentage.

    Pew pulled from Common Crawl, a nonprofit archive that periodically snapshots huge swaths of the public web. Researchers took a random sample of 10,000 English language pages from each of 49 separate crawls spanning January 2021 through July 2026. That gives you something most AI content studies lack: a clean before and after, with almost two years of pre ChatGPT baseline to compare against.

    Each page was run through a detection model built by Pangram, which looks for statistical fingerprints in word choice and sentence construction that separate machine writing from human writing. The model is open weight, which means outside researchers can inspect and reproduce the work rather than taking a vendor’s word for it.

    What Pew did not measure is just as important. This is a count of pages, not a count of readers. A single AI generated page sitting on a dead domain with no visitors counts exactly the same as a heavily trafficked news article. The study also flags pages with “significant signs” of AI authorship, which is not the same as pages written entirely by a bot with no human involvement.

    Recommended for you:

    The NSA Wants You to Reboot Your Router. A Reboot Only Fixes Half the Problem.
    Tech News·Aug 25, 2026

    The NSA Wants You to Reboot Your Router. A Reboot Only Fixes Half the Problem.

    Share of web pages showing signs of AI authorship Pew Research Center, Common Crawl sample of roughly 490,000 English language pages

    All pages, mid 2025 5%

    All pages, July 2026 10%

    Post ChatGPT pages only 35%

    0% 20% 40%

    The gap between 10% and 35% is old pages. Most of the web predates the tools that could have written it.

    The 10% versus 35% gap explains everything

    Here is the part people miss. Ten percent sounds almost reassuring. It is not, because the denominator is stuffed with pages that could not possibly be AI written.

    The internet is an accumulation, not a snapshot. Every random sample of live web pages is dominated by material published years or decades ago: archived forum threads, university course pages, product listings, government records, blogs abandoned in 2014. None of that could have come from a language model, because the models did not exist.

    Filter the sample down to pages published after November 2022 and you remove the ballast. What remains is the web as it is currently being made, and roughly a third of it carries machine writing signatures. As Pew put it, one in ten webpages may seem modest, but the internet includes a mix of new and old material, and a lot of pages in a random sample could not have been written by AI in the first place.

    That framing turns the study from a curiosity into a trend line. The rate is not a fixed property of the web. It is a property of what gets published each month, and that number keeps climbing.

    It is not spread evenly, and the pattern is predictable

    The domain breakdown is the most quietly damning part of the report. AI writing clusters exactly where you would expect it to cluster: on commercial pages, where publishing volume converts into money.

    Domain typeShare showing AI authorshipWhy it lands there
    .comAbout 1 in 10 pagesVolume is the business model. More pages, more search surface, more ad inventory.
    .org4.6%Mixed bag of nonprofits, communities and marketing sites wearing a nonprofit hat
    .eduAround 1%Slow publishing, named authors, institutional review, reputational risk
    .govAround 1%Approval chains and legal accountability make generated text expensive to risk

    A tenfold gap between commercial and institutional domains tells you the incentive, not the technology, is doing the work here. Nobody at a state agency is racing to publish 400 pages a month. Plenty of affiliate sites are.

    The detection caveat nobody wants to talk about

    Pangram’s model is strong, and Pew was upfront about using it, but AI detection has a reputation problem it earned honestly. Detectors have historically produced false positives on formal, plain, or non native English writing, and they degrade as models get better at sounding casual.

    There is a second complication that no detector currently handles well: the hybrid page. A human writes an outline, a model drafts the body, the human edits it back into shape. Is that page AI written? Statistically it may read as machine authored. In practice it is a person using a tool, which is what most working writers now do.

    Read the number as a floor and a ceiling at once

    Detection error cuts both ways. Lightly edited machine text can slip past a classifier and be counted as human, which pushes the real figure up. Heavily edited human writing can trip the classifier, which pushes it down. What survives both objections is the shape of the curve, and the curve is going one direction only.

    The other half of dead internet theory is doing fine

    Dead internet theory always had two claims. One was about who writes the content. The other was about who reads it. The second claim has been quietly winning for years, and it does not need a language model to be true.

    Imperva’s 2026 Bad Bot Report put malicious automated traffic at 40% of all web requests, up from 37% the year before, with total automated traffic including legitimate crawlers sitting above half. That means a majority of the connections hitting a typical website are not people. Scrapers, credential stuffers, price bots, scalper scripts, and now a fast growing category of AI agents fetching pages on someone’s behalf.

    Put the two findings side by side and you get the actual 2026 internet: a machine writing to a machine, with humans as an increasingly optional part of the loop. The theory got the vibe right and the mechanism wrong. Nobody planned this. It is just what happens when publishing costs collapse on one side and automated fetching gets cheap on the other.

    Can any of it be reversed

    Probably not through detection. The economics run the wrong way: a generated page costs a fraction of a cent, and a classifier that flags it has to be right often enough to survive being wrong in public.

    Three things could plausibly bend the curve, and only one of them is technical.

    Recommended for you:

    Fewer Than Ten Are Known to Exist. A French Cereal Promo Made These Nintendo Handhelds a $12,000 Set.
    Tech News·Aug 24, 2026

    Fewer Than Ten Are Known to Exist. A French Cereal Promo Made These Nintendo Handhelds a $12,000 Set.

    • Provenance at the source. Anthropic has begun embedding invisible watermarks in text produced by Claude, designed to persist through editing. Whether that survives contact with copy and paste at scale is an open question, and it only covers models that opt in.
    • Search engines demoting volume. Ranking systems already punish thin content. The difference now is that thin content is indistinguishable from thick content at a glance, so the signal has to shift toward things machines cannot fake cheaply: original reporting, named authors, verifiable expertise.
    • Readers migrating to places with friction. Some platform executives argue people will flock to spaces that verify humans. History suggests otherwise. Users are famously inert, and a decade of complaints about feed quality has moved very few of them anywhere.

    The platforms are already tangled in this. LinkedIn declared war on AI slop while running a business built on posting volume, which is roughly the problem in miniature. Meta spent a year flooding its own products with generated material and then went looking for a detection tool to clean up afterward. Even the estates of dead celebrities have been dragged in, which is why Robin Williams’ children switched his Instagram account back on after 12 years purely to have a place to push back against AI recreations of him.

    None of that is a solution. It is an immune response, and right now it is losing on volume.

    What to do with this if you read the web for a living

    The practical takeaway is not “trust nothing.” It is that the default assumption has flipped. For most of the internet’s life, you could assume a page existed because a person wanted it to exist. That is no longer a safe default for anything published in the last three years on a commercial domain.

    Which pushes the useful signals back to boring, old fashioned ones. Does the page name an author who exists elsewhere? Does it cite a primary source you can open? Does it contain a specific number, date, or quote that would be expensive to fabricate and easy to check? Does it say anything that would be embarrassing to be wrong about?

    Pew’s finding is not that the internet died. It is that a third of what got built recently was built by something that does not care whether you read it. The web still works. You just have to bring more of your own judgment to it than you used to, and that requirement is not going away.

    AI Content Artificial Intelligence Bots ChatGPT Dead Internet Theory Pew Research
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
    Previous Article2026 Is the Worst Year for Phone Sales Since 2013. Your Spec Sheet Is Paying for It.
    Next Article The NSA Wants You to Reboot Your Router. A Reboot Only Fixes Half the Problem.
    Olivia Hartman

      Olivia Hartman is GeekBlog's general technology reporter, covering the wider world of tech beyond smartphones — AI and software, laptops and PCs, gaming, streaming, space, science, consumer gadgets, deals and the policy stories shaping the industry. A versatile journalist with a nose for what actually matters, Olivia turns breaking news and product launches into accessible, no-hype reporting for everyday readers.

      Related Posts

      9 Mins Read

      The NSA Wants You to Reboot Your Router. A Reboot Only Fixes Half the Problem.

      7 Mins Read

      Mattel Is Paying Someone $50,000 for Being Good at UNO, and You Can Qualify From the Couch

      8 Mins Read

      Fewer Than Ten Are Known to Exist. A French Cereal Promo Made These Nintendo Handhelds a $12,000 Set.

      9 Mins Read

      A Laptop Caught Fire Mid Cabin on American Airlines 2398, and the FAA Wants to Know Why

      8 Mins Read

      AliExpress Was Playing Silent Sound Through Your Speakers to Work Out Who You Are

      9 Mins Read

      Nvidia’s AI Security Alliance Tripled to 120 Members. OpenAI and Anthropic Still Won’t Join.

      Top Posts

      How to Use Microsoft Teams: A Beginner’s Guide

      July 7, 20262 Views

      Zip to APK: Convert ZIP Archives Into Installable Android Packages Quickly

      January 16, 20262 Views

      Samsung Set a Date for the Galaxy S26 FE. The Only Real Upgrade Is the Chip.

      August 20, 20261 Views
      Stay In Touch
      • Facebook

      Subscribe to Updates

      Get the latest tech news from FooBar about tech, design and biz.

      Most Popular

      Best Stores for Buying MP3 and Digital Music You Can Keep Forever (2026)

      August 2, 2025932 Views

      Discord will require a face scan or ID for full access next month

      February 9, 2026770 Views

      Trade in your old phone and get up to $1,100 off a new iPhone 17 at AT&T – here’s how

      September 10, 2025383 Views
      Our Picks

      The NSA Wants You to Reboot Your Router. A Reboot Only Fixes Half the Problem.

      August 25, 2026

      One in Three Web Pages Made Since ChatGPT Was Not Written by a Human

      August 25, 2026

      2026 Is the Worst Year for Phone Sales Since 2013. Your Spec Sheet Is Paying for It.

      August 25, 2026

      Subscribe to Updates

      Get the latest creative news from FooBar about art, design and business.

      HEICJPG.online - Convert HEIC to JPG online
      Facebook
      • About Us
      • Contact us
      • Privacy Policy
      • Disclaimer
      • Terms and Conditions
      © 2026 GeekBlog

      Type above and press Enter to search. Press Esc to cancel.