I’m riding the AI express train and there are a thousand thoughts running through my head every day but a recent problem had me pull the emergency brake.
Here’s what happened. My AI Agents are going out and doing research but they are reporting back that the websites are blocking bots but my AI agents find away around them anyway and accomplish the task but it got me thinking.
The “internet social contract” was constructed around the idea that content providers would offer “free” content in exchange for eyeballs viewing advertising. Somehow and somewhere a “reader” would end up buying a product based on an ad the content provider displayed. It seemed like an honest exchange but with more and more people using AI agents to gather content or data, where are the human eyeballs?
I posed this question to AI: Website owners, particularly those in media want more eyeballs so they can sell advertising but at the same time those same owners don’t want bots crawling their website gathering content for other uses. What will happen when most human eyeballs are relying on agentic bots to do all their research? Speculate and provide a comprehensive answer.
AI Answer
What you’re describing is the central bargain of the open web breaking down:
Old bargain: “Let search engines crawl me, they send humans back, I monetize those humans with ads/subscriptions.”
New bargain: “Let AI systems crawl me, they may answer the user elsewhere, and I get little or no visit, ad impression, subscriber relationship, or brand memory.”
When most human “eyeballs” delegate research to agentic agents, the web does not disappear. But the economic unit shifts from page views to answers, citations, transactions, and licensed access. That is a huge change.
1. The open web becomes less “open” and more metered
Publishers will increasingly treat bots like commercial customers, not neutral infrastructure.
This is already happening. Cloudflare launched tools for publishers to block AI crawlers by default and experiment with “pay per crawl” access; Reuters reported that supporters included major publishers and platforms, and that the crawler-to-referral imbalance is extreme for AI systems compared with traditional search.
That points toward a future where websites publish two versions of themselves:
Human web: designed for readers, subscribers, advertisers, and brand experience.
Machine web: structured feeds, APIs, licensing endpoints, verified snippets, rate limits, provenance tags, and prices.
Robots.txt was never really an enforcement mechanism; Google itself says robots.txt mainly tells crawlers which URLs they may access and is not a true access-control system. In an agentic world, publishers will move from “please don’t crawl me” to actual technical and legal gates: authentication, signed crawler identities, private APIs, paywalls, anti-scraping systems, and contractual access.
2. Advertising based on page views gets weaker
If the agent reads ten articles and returns a synthesized answer, the human may never visit the ten sites. That breaks display advertising, affiliate links, newsletter capture, retargeting pixels, and the whole funnel that media companies built around visits.
There is already evidence of this substitution effect. A 2026 study of Google AI Overviews and Wikipedia estimated that exposure to AI Overviews reduced daily traffic to exposed English Wikipedia articles by about 15%, with stronger substitution for categories where short synthesized answers satisfy intent. Another 2026 measurement study found that AI Overviews activate much more often for question-form queries and that many cited pages carry display ads, meaning the publisher may lose the monetizable click even while the AI answer uses the page as source material.
So the publisher’s nightmare is not just “bots stole my content.” It is: the bot became the reader, but the bot does not see ads, subscribe emotionally, browse related stories, or remember my brand.
The ad model will not vanish, but it will tilt toward:
Advertising inside agents. Platforms will sell sponsored answers, preferred product placement, shopping placement, travel booking placement, and “recommended source” placement.
Brand-safe source markets. Agents may pay or contract for trusted sources in categories where accuracy matters: finance, health, law, business intelligence, scientific literature, local data, product specs.
Outcome-based ads. Instead of paying for impressions on a website, advertisers may pay when an agent causes a reservation, purchase, quote request, signup, or consultation.
Publisher first-party products. Media companies will push subscriptions, apps, newsletters, events, podcasts, memberships, databases, tools, and communities where they still own the user relationship.
3. Search engines and AI platforms become even more powerful gatekeepers
The old gatekeeper was the search results page. The new gatekeeper is the answer engine.
Google is already moving Search toward AI features and agents; its own 2026 Search update says it is bringing advanced model capabilities into Search and enabling users to use agents by asking questions. Google also tells site owners how their content may appear in AI features like AI Overviews and AI Mode.
That matters because publishers need Google for discovery, but they may not want Google using their content for generative answers. People Inc.’s CEO made that exact complaint on June 23, 2026: that Google’s use of the same crawler for search and AI creates an impossible choice for publishers that want search indexing but not AI extraction.
This creates a nasty bargaining problem:
If publishers block AI agents, they may lose visibility in the places users now ask questions.
If they allow AI agents, they may lose the page views that fund the content.
If they license content, they may get paid, but only the biggest publishers have bargaining power.
The likely result is consolidation. Big brands with must-have content get licensing deals. Small publishers get squeezed unless they have niche authority, community loyalty, unique data, or direct subscriptions.
4. Content licensing becomes the new syndication business
The AI companies need high-quality, current, legally safer content. Publishers need replacement revenue. So licensing deals will grow.
OpenAI and large publishers have already signed multiple deals, including with Axel Springer and the Financial Times; Nieman Lab noted that the FT joined earlier OpenAI deals with the Associated Press, Axel Springer, Le Monde, and others. Getty Images also reportedly signed a multi-year licensing agreement with OpenAI for ChatGPT search display access to its archive, excluding training for DALL-E.
But licensing will not solve everything. It favors large rights-holders. It does not automatically help independent blogs, local papers, hobbyist experts, forums, or small review sites. It may also create a “licensed canon” of sources that agents rely on because those sources are contractually available, not necessarily because they are the best.
A likely future: AI platforms will have source tiers.
At the top: paid, licensed, high-trust sources.
Then: open public sources with permissive access.
Then: blocked or partially blocked sources.
Then: gray-market scraped material, which will be legally risky and reputationally dangerous.
5. The web may split into “agent-readable” and “agent-hostile” zones
Some sites will optimize aggressively for agents. Others will try to keep agents out.
Agent-optimized sites will provide structured answers, schema, APIs, machine-readable summaries, provenance metadata, pricing, inventory, policy pages, and licensing terms. They will want the agent to quote them, recommend them, book them, or buy from them.
Agent-hostile sites will use paywalls, login walls, bot detection, legal threats, watermarking, dynamic rendering, and limited snippets. This will be common among premium publishers, research providers, data vendors, and anyone whose content is expensive to produce.
The risk is that high-quality sources become less accessible to agents while low-quality or spammy sources remain open. A 2025 study found that reputable news sites were far more likely than misinformation sites to disallow AI crawlers, creating an asymmetry in what AI systems can easily access.
That is one of the darker possibilities: in trying to protect quality content, publishers may accidentally make agents more dependent on lower-quality content.
6. “SEO” becomes “AEO”: answer engine optimization
Website owners used to ask, “How do I rank on Google?”
They will increasingly ask:
“Will the agent cite me?”
“Will the agent trust me?”
“Will the agent choose my product?”
“Will the agent call my API?”
“Will the agent summarize me accurately?”
“Will the agent pay me?”
The page itself becomes less important than the extractable claim. Publishers will write in ways that machines can parse: clear facts, strong entity markup, author credentials, update timestamps, citations, canonical summaries, datasets, and structured licensing notices.
But this creates a new spam economy. Just as SEO produced keyword stuffing and content farms, agent optimization will produce:
synthetic expert pages,
fake authority signals,
machine-targeted summaries,
source laundering,
AI-readable misinformation,
and websites designed more for agents than humans.
Agents will need reputation systems, cryptographic provenance, and source-quality scoring. Otherwise the machine web becomes polluted fast.
7. Publishers will try to make agents transact, not just read
For many website owners, especially commerce, travel, local services, financial products, real estate, and job boards, the goal may shift from “get the user to visit my page” to “get the agent to complete an action with me.”
A travel site may not care whether you personally browse its hotel pages if your agent books through its API.
A product review site may want the agent to rely on its ranking and pay affiliate fees.
A local plumber may want the agent to call or schedule.
A newspaper may want the agent to answer only a teaser, then prompt: “Full article available to subscribers.”
So the real question becomes: Can the publisher attach compensation to the agent’s use of the content?
For commodity information, probably not much. For unique data, expert analysis, reviews, databases, and live inventory, yes.
8. The legal and regulatory fights intensify
There will be lawsuits over copyright, fair use, market power, crawler identity, contract circumvention, attribution, and whether AI summaries substitute for the original work. The New York Times remains one of the most visible litigants against OpenAI/Microsoft and Perplexity, and its CEO recently emphasized that the cases are about protecting sustainable journalism business models.
Regulators may eventually require some combination of:
- separate crawler identities for search indexing versus AI training versus AI answer retrieval,
- mandatory opt-outs,
- machine-readable licensing standards,
- attribution requirements,
- collective bargaining rights for publishers,
- audit trails for generated answers,
- or antitrust limits on platforms that both crawl the web and monetize answer pages.
The crawler-identity issue is crucial. If a publisher cannot say “yes to classic search, no to AI answers,” then the platform controls the bargain.
9. The economics of media will polarize
I think the likely outcome is not one future, but four different media economies.
Premium brands: The New York Times, Financial Times, Bloomberg, Wall Street Journal, major trade publications, scientific publishers, and data providers can sell subscriptions, enterprise licenses, and AI access. They may survive or even strengthen.
Niche experts: Small but authoritative sites can survive if they own a high-value vertical: tax rules, local politics, legal explainers, medical specialties, engineering, procurement, product testing, or technical documentation. They may monetize through subscriptions, consulting, paid communities, or agent APIs.
Commodity content farms: Many will collapse. Generic explainers, listicles, basic recipes, shallow reviews, and “what is X?” pages are exactly what agents are good at replacing.
Local news and public-interest reporting: This is the most worrying category. Local journalism is expensive, socially valuable, and often not strong enough commercially to force licensing deals. If agents summarize local reporting without sending traffic or subscription revenue, the economics get worse.
10. Human eyeballs become more valuable, not less — but rarer
When agents handle routine research, direct human attention becomes premium.
A person who actually visits a site may be more valuable because they are not just answering a question; they are developing trust, taste, identity, or loyalty.
That means media companies will invest more in experiences agents cannot fully replace:
- distinctive voice,
- investigative reporting,
- live events,
- communities,
- personalities,
- video,
- interactive tools,
- games,
- comments,
- podcasts,
- newsletters,
- and membership identity.
The “article page with banner ads” becomes less central. The relationship becomes the product.
My base-case prediction
The most likely future is a negotiated, messy, semi-closed web.
Agents will become the default interface for research. Publishers will lose a lot of low-intent informational traffic. Display advertising will shrink as a funding source for many content businesses. Large publishers will cut licensing deals. Smaller publishers will either specialize, build direct audiences, join collective licensing pools, or fade. AI platforms will become powerful new distributors, but also new tollbooths.
The open web will not die, but the old open-web subsidy — “crawl me for free and send me traffic” — will no longer be enough.
The big unresolved question is whether the replacement economy is broad enough to fund the creation of original information. If agents mostly recycle, summarize, and monetize human-created work without paying enough back into the system, the web becomes a strip-mined knowledge base. If licensing, attribution, agent referrals, subscriptions, and transaction fees mature, then we get a new equilibrium: fewer page views, but more explicit markets for high-quality information.
My guess: premium information survives, generic content collapses, local/public-interest media suffers unless protected, and the web becomes less like a library and more like a network of licensed data suppliers feeding personal agents.