All insights
AI TrendsAI Opportunity Hunter

Data Is the New Oil, but AI Is the New Gatekeeper

Data Monopoly, Algorithmic Recommendation Power, and the Future Control of the Digital Economy

The most important economic decision of the next decade may happen before a consumer realizes that a decision is being made.

Someone asks an AI assistant to find a hotel. The system does not return ten blue links. It reads the request, infers the traveler’s priorities, compares availability, predicts which properties fit the budget and schedule, and presents three choices. A business owner asks how to reach new customers. The answer depends on whether the recommendation system can retrieve the company, trust its data, understand its offer, and rank it above thousands of alternatives. A worker asks an enterprise agent to prepare a report. The agent decides which documents count as relevant evidence and which colleagues or systems should be consulted.

In each case, the visible interface is a conversation. The hidden economic function is allocation. The system allocates attention, trust, traffic, opportunities, and sometimes money.

This is why the familiar phrase “data is the new oil” is no longer enough. Data is valuable, but raw accumulation does not automatically create power. The decisive advantage comes from converting repeated activity into better prediction, better ranking, lower friction, and more activity. The company that controls this feedback loop can become more than a software provider. It can become the gatekeeper between users and the economy.

The central question is simple:

Whoever controls data may control AI decision-making power. But when AI becomes the intermediary between humans and businesses, who controls the recommendation algorithm?

This article traces the answer from the early history of artificial intelligence to the emerging age of agents. It examines why data creates network effects, how recommendation systems shape demand rather than merely reflect it, why small businesses may face a discoverability shock, and how the European Union, the United States, and China are beginning to regulate the new bottlenecks. It also compares Google, Amazon, Meta, TikTok, Tesla, OpenAI, Microsoft, and NVIDIA, then develops three broad scenarios for 2030–2040.

The conclusion is deliberately qualified. Data is a strategic asset, but not a magic moat. The durable power of AI comes from a combination of data, computing, models, distribution, identity, trust, and regulatory legitimacy. If AI becomes the invisible decision-maker of society, the challenge will not only be building intelligent machines. It will be ensuring that the intelligence layer remains contestable and aligned with human interests.

A short executive summary

Five claims organize the analysis.

First, data creates advantage through feedback loops, not possession alone. A large dataset matters when it is relevant to a decision, updated by new behavior, connected to a model, and deployed at sufficient scale. Google’s 2009 research on large datasets, Meta’s production recommendation research, and recent company filings all point to the same mechanism: activity generates signals; signals improve products; improved products attract more activity.

Second, the AI gatekeeper is the integrated system around the model. The relevant control point includes the retrieval index, ranking objective, tool permissions, identity layer, payment relationship, cloud infrastructure, and interface. Model weights can be copied or licensed while distribution, context, and transaction authority remain concentrated.

Third, recommendation algorithms allocate opportunity. A feed decides which creator receives an impression. A shopping assistant decides which seller enters the shortlist. An enterprise copilot decides which document is considered relevant. The output may look like neutral assistance even when the objective includes conversion, margin, advertising yield, retention, or a platform’s own ecosystem priorities.

Fourth, AI agents increase the stakes because they can act after recommending. Search results became a route to a website. An agent can become a route to a purchase, booking, software deployment, or business decision. The market power of the interface therefore grows as the user delegates more judgment.

Fifth, the future is not predetermined. Concentrated AI distribution, interoperable agents, sovereign AI blocs, and competitive local intelligence can coexist. Regulation can improve contestability, but portability or transparency will not work if users cannot switch, rivals cannot obtain meaningful inputs, or businesses cannot understand why they are excluded.

Chapter 1: The birth of AI and the evolution of data as strategic power

From symbolic rules to statistical systems

Artificial intelligence began as an attempt to represent intelligence explicitly. In the 1950s, researchers explored symbolic reasoning, search, logic, and games. The machine was given rules, representations, and procedures for manipulating them. The approach was intellectually ambitious because it treated reasoning as something that could be formalized.

The expert-systems era extended that idea into narrow domains. A medical or industrial system could encode the knowledge of specialists in a rule base and apply it consistently. The limitation was not that rules were useless. It was that the world contained too many exceptions, ambiguous signals, and changing contexts for experts to write a complete program. The cost of maintaining the rules grew with every new edge case.

Machine learning changed the direction of travel. Instead of asking humans to describe every rule, researchers asked systems to infer patterns from examples. The data became part of the program. A classifier learned from labeled images; a fraud model learned from transaction histories; a search system learned from queries and clicks. The “software” was no longer only the code written by engineers. It included the dataset, the objective function, the model architecture, and the feedback process.

The deep-learning breakthroughs that became visible after 2012 were partly advances in neural architectures and training methods, and partly the arrival of larger datasets and cheaper parallel computing. Image recognition improved when models could be trained on millions of examples with powerful hardware. Speech recognition improved when systems could absorb diverse audio. Recommendation systems grew more sophisticated as platforms recorded granular user behavior at scale.

The rise of the web created a particularly valuable type of data: interaction data. A page view says that something was seen. A query, click, dwell time, return visit, purchase, skip, correction, or complaint says something about what happened next. Those signals are imperfect, but at scale they can become training material and business intelligence at the same time.

The transformer and the language-model economy

The transformer architecture, introduced in the 2017 paper “Attention Is All You Need,” changed the economics of sequence modeling. It enabled systems to process relationships among tokens more effectively and to scale pretraining across text and later images, audio, video, and code. The model was no longer built for a single task; it learned broad statistical structure and could be adapted to many tasks.

The generative-AI era made this capability visible to ordinary users. The ChatGPT release in 2022 did not invent machine learning, language models, or conversational interfaces. It made a general-purpose model into a mass-market product. The interface turned a difficult research achievement into an everyday habit: ask, receive, revise, and ask again.

That habit produces a new kind of data. A chat session can include a user’s intent, constraints, corrections, preferred format, domain vocabulary, and reaction to an answer. An enterprise assistant can observe which documents were accepted, which workflows were completed, and where a human intervened. The information is more contextual than a simple click, although it also raises stronger privacy and governance questions.

Agents move the loop from prediction to action

After 2024, the industry’s attention shifted toward AI agents. The distinction is not perfectly standardized, but the basic idea is that a system can plan a task, choose tools, inspect results, revise its plan, and continue until it completes the work or asks for help.

OpenAI’s ChatGPT agent announcement describes browsing, code execution, research, connected data, and browser actions. Google’s AI Mode announcement describes query fan-out, cited synthesis, and agentic help with tickets, reservations, and shopping. These are product disclosures, not independent proof of reliability or market dominance. They do reveal the direction of the interface.

The traditional loop was:

Data → Model → Prediction → User decision

The agentic loop is closer to:

Data → Model → Recommendation → Tool call → Outcome → New behavioral data

The outcome is now part of the learning signal. Did the hotel booking succeed? Did the code change pass review? Did the customer accept the recommended product? Did a human reverse the action? In theory, this makes systems more useful. In economic terms, it also makes the agent a more powerful intermediary.

The evolution of strategic assets

The dominant asset has changed with each digital era. The pattern is not deterministic, but it clarifies why data and recommendation power matter.

| Era | Primary scarce asset | Typical control point | Economic advantage | |---|---|---|---| | Industrial | Factories, machines, energy, logistics | Production capacity | Lower unit cost and reliable supply | | Early software | Code, distribution, operating systems | Software platform | Reusable functions and switching costs | | Internet | Attention, links, identity, marketplaces | Digital distribution | Network effects and transaction scale | | Mobile platforms | Location, social graph, app ecosystem | Device and app layer | Persistent context and targeted access | | Generative AI | Compute, training data, model capability | Foundation-model stack | General-purpose prediction and creation | | Agentic AI | Context, permissions, ranking, trust | Decision and action layer | Delegated choice and transaction control |

The point is not that data stopped mattering. It became embedded in a larger system. In the agentic era, controlling the route from question to action may be more important than simply owning a large archive.

Chapter 2: Why data creates monopoly power in AI

Data is not a commodity in the ordinary sense

Oil is a physical resource that can be extracted, refined, transported, and consumed. Data behaves differently. Multiple companies can copy a public fact. A dataset can be reused without being depleted. Its value depends on freshness, exclusivity, granularity, legal permission, labels, representativeness, and the model that uses it.

The phrase “data is the new oil” is therefore a metaphor, not an economic identity. A company can have more data and still lose if the data is noisy, irrelevant, inaccessible to engineers, or poorly connected to a valuable workflow. Conversely, a smaller dataset can be strategically important if it captures rare events, proprietary outcomes, or high-quality feedback that competitors cannot obtain.

Google’s first-party paper The Unreasonable Effectiveness of Data argued in 2009 that very large datasets can outperform more handcrafted approaches in difficult language tasks. The paper supports a technical point: scale can compensate for elegant rules when the system can learn from many examples. It does not prove that data automatically creates a legal monopoly. The economic question is whether the data is exclusive, whether it improves a product, and whether rivals can obtain a comparable signal. Meta’s production recommendation research shows why the data advantage is also an infrastructure problem: personalization models use dense and sparse behavioral features at a scale that can influence data-center design.

Data network effects

A network effect exists when a service becomes more valuable as participation grows. A telephone becomes useful when more people have a telephone. A marketplace becomes useful when more buyers and sellers are present. A data network effect adds a learning mechanism:

  1. More users generate more queries, clicks, purchases, watch time, ratings, corrections, and support requests.
  2. The platform turns those interactions into features, labels, forecasts, rankings, and fraud signals.
  3. Better prediction lowers search costs or improves conversion.
  4. A better experience attracts more users, sellers, advertisers, creators, developers, or employees.
  5. The enlarged ecosystem creates more interactions, restarting the loop.

The loop can be positive, but it is not guaranteed to be healthy. A model optimized for engagement can amplify sensational content. A marketplace optimized for conversion can privilege familiar brands. A search system trained on click data can learn from previous ranking decisions, making its own exposure look like proof of relevance. Feedback can improve the product or reproduce the platform’s existing biases.

Direct, indirect, and cross-market effects

Platforms often have several reinforcing sides. A search engine connects users, advertisers, publishers, and businesses. A social network connects viewers, creators, and advertisers. A marketplace connects consumers, sellers, logistics providers, and payment services. A cloud AI platform connects developers, model providers, chip suppliers, and enterprise buyers.

More activity on one side can make the service better for another side. Advertisers pay for access to attention; their revenue funds infrastructure and distribution. Sellers provide inventory; more inventory improves the shopping experience; more shoppers attract sellers. Developers build tools for a compute platform; their tools make the platform more attractive to other developers.

This is why the data advantage is often not a single database. It is a connected system of identity, APIs, ranking, payments, business relationships, and infrastructure. A startup may reproduce a model architecture while lacking the distribution channel that generates the next million high-quality feedback signals.

Barriers to entry are real, but not always permanent

Large incumbents have several advantages:

  • They can collect data from a high-frequency consumer or enterprise product.
  • They can subsidize experimentation with advertising, commerce, cloud, or hardware revenue.
  • They can integrate data across services, subject to privacy law and user choice.
  • They can train models and run inference at lower average cost because infrastructure is already deployed.
  • They can place a new AI feature in front of an existing user base.
  • They can use trust, identity, payments, and support relationships to complete actions.

None of these advantages is necessarily permanent. Data can become stale. Models can make old differences less important. A new interface can collect a more useful signal. Open-source systems can lower the cost of experimentation. Regulation can impose portability or interoperability. A focused company can win with a better workflow even without a general-purpose dataset.

The more precise claim is that data creates a dynamic barrier. It raises the cost and time required to catch up, particularly when the incumbent uses feedback to improve faster than entrants can acquire users. Whether the barrier becomes a monopoly depends on switching costs, multi-homing, interoperability, infrastructure access, and the ability of rivals to reach customers without passing through the incumbent.

The feedback-loop equation

For analytical purposes, imagine a platform’s recommendation quality as a function:

Q = f(D, M, R, I, O, T)

where D is data quality and coverage, M is model capability, R is retrieval and ranking design, I is infrastructure, O is the objective function, and T is trust and distribution. Improving one input may have little effect if another is weak. A company with extraordinary data but poor product distribution may not learn quickly. A company with an excellent model but no permission to access relevant context may produce generic results.

The strategic danger is that a firm can control several variables at once. It owns the data source, the training system, the interface, the identity layer, and the transaction channel. At that point, competition is not between models in a laboratory. It is between ecosystems with different access to the learning loop.

Chapter 3: The rise of AI agents and the death of traditional search

Search was already a recommendation system

Search engines were never neutral libraries. They crawled, indexed, ranked, and presented a limited set of results. A user who typed “best laptop for video editing” did not receive the whole web. The user received a ranked interpretation shaped by relevance, authority, freshness, location, language, commercial signals, and the search engine’s own design.

The important change in AI search is not that ranking suddenly appeared. It is that the ranking output becomes an explanation and, potentially, an action plan.

From ten thousand results to one shortlist

Consider two interfaces.

Search-era request: “Find hotels in Kyoto for three nights.”

The platform returns pages of options. The user compares prices, reviews, neighborhoods, cancellation terms, and availability. The platform influences the order but leaves substantial work to the user.

Agent-era request: “Book the best quiet hotel in Kyoto near a station, under my budget, with a flexible cancellation policy.”

The agent may infer what “best” means, retrieve a subset of properties, compare structured facts, ask one clarifying question, and complete a reservation. The user experiences convenience. The hotel experiences a new distribution channel whose rules are partly invisible.

The agent must decide which sources to trust, how to reconcile conflicting information, what alternatives to omit, whether a partner has a commercial relationship, and when a user’s preference overrides a predicted objective. These are marketplace decisions, even if they are implemented as model calls.

Google’s answer-and-action layer

Google’s 2025 AI Mode announcement describes “query fan-out,” where a question is decomposed into multiple searches. It also describes Deep Search and agentic workflows involving shopping, tickets, and reservations. The company presents the features as a way to reduce the user’s research burden. The strategic implication is that Google’s index, shopping graph, maps, identity, and commercial infrastructure can be combined into a single decision surface.

The publisher’s problem changes. It is no longer sufficient to appear on a results page. A business may need to be crawled, understood, cited, represented accurately, and deemed suitable for the action the user wants to take. The platform can own the customer relationship even when the transaction is completed on another website.

The U.S. Department of Justice’s 2025 Google remedies make the data issue concrete. The DOJ announcement describes court-ordered access to certain search-index and user-interaction data and search or text-ad syndication for qualified competitors. The remedy is tied to a specific litigated case; it is not a universal rule that every platform must open every dataset. Its importance is conceptual: regulators treated data and distribution access as relevant inputs to search competition.

Amazon’s shopping agent

Amazon’s 2025 Form 10-K describes competition not only from retailers, but also from search engines, comparison-shopping sites, social networks, virtual assistants, and AI-enabled ways of discovering or acquiring goods. That wording recognizes that the shopping journey can begin outside a traditional catalog.

An AI shopping agent could compare price, delivery, product compatibility, return policy, seller reliability, and user preference. Amazon has an unusual position because it combines a first-party retailer, third-party sellers, fulfillment, advertising, customer-service history, and transaction data. A model that controls the shortlist can influence not only what is purchased, but which sellers receive the opportunity to compete.

The commercial question is not whether recommendation is useful. It is who sets the objective. Does “best” mean lowest price, highest quality, fastest delivery, greatest margin, highest advertising bid, most reliable return outcome, or some weighted combination? A user can delegate a decision without realizing that the platform’s economics are embedded in the weighting.

OpenAI and the personal context layer

OpenAI’s ChatGPT agent disclosure shows another route to gatekeeper power. A general assistant can browse sites, execute code, work with connected information, and create deliverables. Its advantage may come less from owning a marketplace and more from becoming the user’s proxy across marketplaces.

If the same agent knows a person’s schedule, budget, preferred brands, past purchases, workplace systems, and tolerance for risk, it can produce highly personalized recommendations. The data is valuable because it is contextual. The privacy risk is correspondingly high: a system that can make a good recommendation may also know more about the user’s priorities than any individual merchant.

The agent can become a gatekeeper even when it does not own the underlying index. It chooses which sources to inspect, which APIs to call, which products to shortlist, and which evidence to summarize. A company may be absent from the final answer because of retrieval, ranking, data quality, a safety policy, a commercial rule, or an interface limit. The user sees one answer, not the counterfactual market.

The end of browsing is not inevitable

The phrase “death of search” is rhetorically powerful but analytically too strong. Users will continue to browse when they want exploration, verification, entertainment, social proof, or control. Agents will not always have sufficient confidence to act. Regulators may require disclosure and alternatives. Businesses may insist on direct relationships.

The more plausible shift is a layered market. Search remains a discovery tool; agents become a research and action layer; specialized marketplaces retain transaction authority; and users move between them depending on trust and complexity. The risk is greatest in high-volume, low-attention decisions where a default recommendation can quietly determine demand.

Chapter 4: Algorithmic recommendation power

Algorithms do not only predict demand

A recommendation system begins with prediction. It estimates what a user may watch, click, buy, read, or complete. But the act of recommending changes the environment from which future data is collected. A product that receives more impressions has more opportunities to receive clicks. A creator who is shown to more viewers has more chances to build a following. A news story placed near the top becomes more likely to be discussed.

The recommendation therefore performs two functions:

  1. It predicts what might happen.
  2. It allocates the opportunity for that outcome to happen.

This is the difference between measuring demand and shaping demand. The platform’s ranking policy becomes an allocation rule for attention and revenue.

TikTok: discovery without a conventional social graph

TikTok’s official documentation says recommendations can use user interactions such as watch behavior, likes, shares, follows, comments, searches, and skips, along with content information and user information such as language or device settings. It also says signals from users with similar interests may influence recommendations.

The For You feed can introduce a creator a user does not follow. This can lower one barrier to discovery: a new account does not need an established friend network to receive a test audience. But the trade-off is centralized ranking power. Eligibility, safety, diversity, freshness, predicted interest, and commercial rules determine who receives distribution.

TikTok’s public help page is not a full algorithmic audit. It lists categories of signals, not every feature, weight, model, or enforcement rule. That distinction matters. Transparency about inputs can help users understand the system without revealing enough for outsiders to reconstruct or challenge its real allocation policy.

YouTube and the economics of watch time

Video recommendation systems must balance immediate engagement with long-term satisfaction, safety, creator diversity, and advertiser suitability. A system that optimizes only watch time may favor content that provokes strong reactions. A system that optimizes satisfaction may need signals that arrive later, such as whether viewers return, complain, or abandon the platform.

The feedback loop can generate a winner-take-more dynamic. Popular videos create more training data and more social proof. More data improves personalization, which can increase distribution. The loop may also produce discovery opportunities for niche creators if the system values relevance rather than only historical popularity. The outcome depends on the objective and the constraints, not on the existence of recommendation itself.

Netflix and Spotify: recommendation as culture

Entertainment recommendation is often treated as a convenience problem, but it also shapes cultural exposure. A Netflix interface determines which stories appear as salient options. Spotify affects which artists are discovered, repeated, and included in playlists. The platforms use signals such as prior behavior, similarity, context, and popularity, then learn from what users do after a recommendation.

The economic issue is not that one platform controls all taste. It is that the cost of ignoring the platform rises when it controls a large share of discovery. A creator, studio, or musician can maintain a product and still become commercially invisible if the recommendation system does not surface it.

Amazon: recommendation becomes shopping infrastructure

In commerce, recommendation affects prices, demand, inventory, advertising, and seller survival. A “frequently bought together” module is visible merchandising. A conversational agent’s decision to include or exclude a product is more opaque merchandising. A seller may not know whether a rival won because of price, delivery, review quality, advertising, margin, data completeness, or model error.

A future requirement for AI-mediated commerce may be a clear separation between relevance and paid influence. The user does not need the entire model specification, but should be able to know when a choice is sponsored, when a platform benefits from the transaction, and which user preference or constraint caused the recommendation.

Hotels, restaurants, and local services

Local businesses face an especially direct version of the problem. A restaurant can be highly rated by its customers and still be missing from an AI answer if the model cannot verify its opening hours, menu, location, dietary options, or reservation availability. A small hotel can offer excellent service but lack structured, consistent information across public sources.

In a link-based world, the user might discover the business through a map, a review, a search result, a recommendation from a friend, or an advertisement. In an agent-mediated world, the user may see only a handful of options. A business that does not enter the retrieval layer may effectively disappear from demand.

Helping versus influencing

The boundary between assistance and influence is not binary. Every interface organizes information. The important distinctions are:

  • Is the objective disclosed?
  • Can the user change the main preference settings?
  • Are commercial relationships visible?
  • Can the user ask for alternatives and see why they differ?
  • Is there a route to appeal an incorrect representation?
  • Can businesses obtain machine-readable facts about eligibility and errors?

The more consequential the decision, the stronger these controls should be. Recommending a song is not the same as recommending a loan, medical provider, job candidate, or supplier.

Chapter 5: The small-business problem

From SEO to AI visibility

For two decades, small businesses learned to compete for digital attention through search-engine optimization, local listings, advertising, branding, reviews, links, and direct relationships. The strategy was imperfect, but the surface was legible enough to study. A business could see impressions, rankings, clicks, and conversions.

AI recommendations compress that surface. A user may receive a paragraph, a table of three options, or an agent’s final action. The business may not know which alternatives were considered or why it was omitted. The result is a potential discoverability shock.

Google Search Central’s guidance on AI features says there is no special AI file or secret markup required for AI Overviews or AI Mode. It emphasizes familiar fundamentals: crawlability, helpful content, internal links, structured data that matches visible information, and up-to-date business or merchant profiles. The guidance is reassuring in one sense: businesses do not need a completely new technical language. It is also revealing in another: the platform decides whether a business can be crawled, indexed, interpreted, and served.

Machine-readable business identity

An AI-ready business needs more than persuasive copy. It needs a reliable public identity:

  • exact name, address, region, and service area;
  • current hours, holiday exceptions, and availability;
  • products, prices, variants, inventory, and constraints;
  • policies for delivery, cancellation, returns, accessibility, and privacy;
  • verifiable reviews and evidence of service quality;
  • structured data consistent with the visible page;
  • stable URLs, APIs, feeds, and credentials for authorized agents;
  • a process for correcting errors and outdated representations.

This is not a guarantee of recommendation. It is the minimum data layer that helps a model represent a business accurately. The commercial question is whether future agents will privilege official feeds, trusted aggregators, paid partnerships, or a mixture of all three.

Agent optimization could repeat the problems of SEO

“AI Search Optimization” or “Agent Optimization” will likely become a market category. Some of the work will be useful: improving data quality, provenance, page structure, service descriptions, and customer experience. Some of it may repeat the worst practices of SEO: gaming citations, generating synthetic reviews, stuffing documents with model-directed language, and creating content designed to manipulate retrieval rather than help customers.

The platform response will be another learning loop. It may reward freshness, first-party evidence, independent corroboration, user satisfaction, and reliable transaction outcomes. Businesses will attempt to influence those signals. The more important AI recommendations become, the more valuable and contested the signals will be.

The case for neutral access

A fair agent economy should not require every small business to negotiate separately with every AI provider. Open schemas, portable catalogs, transparent eligibility rules, and secure APIs can reduce the distribution burden. Regulators may also require platforms to provide meaningful explanations or access to data about how business information is used.

Yet access alone does not solve the ranking problem. If every business can submit data, the agent still needs to rank them. The policy challenge is to preserve differentiation based on quality while preventing secret commercial influence, discriminatory exclusion, and unreviewable errors.

Chapter 6: Can AI recommendations be objective?

Recommendation is a choice of objective

The phrase “objective recommendation” hides a design decision. A platform must decide what it is optimizing. Possible objectives include relevance, satisfaction, price, quality, novelty, diversity, safety, speed, margin, advertising revenue, retention, or some weighted function.

Suppose an agent recommends hotels using the following fictional weighting:

| Factor | Weight | |---|---:| | Verified customer satisfaction | 40% | | Price and total cost | 30% | | Location and travel time | 20% | | Commercial partnership | 10% |

This table would not make the recommendation neutral, but it would make the trade-off visible. The user could decide that commercial partnership should be zero, or that accessibility should receive a higher weight. In practice, the model may use hundreds of features and dynamically adjust them. The principle remains: ranking is an allocation of probability and attention, not a natural fact.

Commercial bias

Commercial influence can enter through explicit advertising, commissions, preferred-partner programs, default tools, data partnerships, or an objective that includes platform margin. Paid placement is not automatically illegitimate. Advertising supports much of the internet, and users may prefer a cheaper service funded by commercial relationships.

The risk is opacity. A conventional advertisement can be labeled. A synthesized answer that quietly favors a partner may appear to be independent research. The user may not know whether a product was recommended because it was the best match, because its data was easier to retrieve, because it paid for access, or because the platform earns more from it.

Data bias and historical inequality

Training and feedback data reflect who was visible, who had access, whose behavior was recorded, and which outcomes were previously rewarded. A small business in a low-connectivity region may have fewer reviews and less structured information. A local language may have fewer examples. A product category used by a minority may appear less frequently even when its customers value it strongly.

The model does not need an explicit discriminatory rule to reproduce the inequality. It can use proxies such as geography, language, device quality, price, popularity, or prior ranking. A quality filter can become a barrier when the historical data was shaped by unequal distribution.

Popularity bias and self-fulfilling rankings

Popularity is a useful signal, but it becomes dangerous when it is both the input and the output. A business receives more exposure because it was already popular; it receives more transactions because it received more exposure; the model interprets the transactions as evidence of superior relevance.

The loop can suppress new entrants even when they offer a better product. Platforms sometimes counter this with exploration, diversity constraints, randomized exposure, or fairness objectives. Those interventions create their own trade-offs: a user may see a less likely option, a seller may receive low-quality traffic, and the platform must decide how much experimentation is acceptable.

Platform self-preferencing

A platform that owns a marketplace, payment system, cloud service, or content library can favor its own products without publishing an explicit rule. It may use superior internal data, default placement, integration advantages, or a ranking objective that values ecosystem retention.

The proper analysis is not “every first-party recommendation is illegal.” Consumers may benefit from integration. The question is whether the platform can use control over the interface to disadvantage rivals that depend on access to the same demand, and whether the user has a realistic alternative.

Cultural and regional bias

A global model must interpret local preferences, social norms, languages, and regulations. A recommendation trained mainly on one region may treat its assumptions as universal. An agent can also produce different answers across jurisdictions because the retrieval index, legal obligations, data residency, and available services differ.

Regional adaptation is not automatically bias. It can be necessary for relevance and safety. The concern arises when a user cannot understand why the system changed the choice, when businesses cannot correct an inaccurate representation, or when cultural assumptions affect access to jobs, credit, healthcare, education, or public information.

The black-box problem is partly economic

Model opacity is often discussed as a technical problem: engineers cannot fully explain internal representations. For recommendation, opacity is also a market problem. A business excluded from a shortlist may not know whether it failed a quality threshold, lacked data, lost an auction, was outside the retrieval index, or was filtered by a safety policy.

Publishing a model card will not answer every case. A practical system might provide an explanation of the dominant factors, a category-level reason for exclusion, a way to request correction, and an independent audit route. Users need enough information to exercise choice without receiving a playbook for gaming the system or exposing personal data.

Chapter 7: How should AI recommendation systems be regulated?

No single law covers the new gatekeeper. Different instruments address different layers: safety, privacy, competition, platform responsibility, data access, and national security.

The European AI Act: risk and accountability

The EU AI Act, Regulation 2024/1689, establishes a risk-based framework intended to support human-centric and trustworthy AI while protecting health, safety, fundamental rights, democracy, and the rule of law. It creates obligations that vary by system and risk category, including transparency, documentation, risk management, and responsibilities across providers and deployers.

For the data-gatekeeper thesis, the AI Act matters because it makes accountability part of the product lifecycle. A recommendation or agent system used in a consequential context cannot be treated as a black-box feature with no owner. The Act does not, by itself, solve market concentration. A compliant model can still operate inside a dominant distribution channel. Safety compliance and competition are related but distinct policy goals.

The Digital Markets Act: contestability and data access

The Digital Markets Act targets large “gatekeepers” providing core platform services. Its architecture addresses fairness, contestability, interoperability, user choice, data combination, and portability. The European Commission’s gatekeeper framework includes Alphabet, Amazon, Apple, ByteDance, Meta, and Microsoft, with other services designated over time.

The DMA is directly relevant to data feedback loops. It can limit certain cross-service combinations of personal data without consent and create rights around access to data generated through platform activity. If a user can move relevant data or a business can obtain meaningful performance information, switching becomes more possible.

But portability is not the same as copying a recommendation engine. A user may export posts, transactions, or profile information while leaving behind embeddings, inferred preferences, fraud signals, aggregate learning, ranking features, and the social context that makes the service valuable. A business may receive a dashboard without receiving a route to the customers who generated the data.

Future enforcement will face a hard question: which derived data is necessary for meaningful competition, and which derived signal is a protected trade secret or a privacy risk? The answer cannot be “share everything.” It must be scoped to the bottleneck and safeguarded against abuse.

The Digital Services Act: transparency and alternatives

The Digital Services Act requires online platforms to explain the main parameters of recommender systems and the options recipients have to modify or influence them. Very large platforms and search engines must offer at least one option that is not based on profiling for their recommender systems. The DSA also addresses advertising transparency, systemic-risk assessment, mitigation, and researcher access under safeguards.

This is a meaningful baseline. A user should not have to accept personalization as the only route to information. An alternative feed or search mode can reveal how much a ranking changes when the platform cannot use an individual profile.

The DSA does not require the publication of model weights or every feature. Its emphasis is intelligibility, choice, accountability, and systemic-risk research. Agentic search will test the interpretation. If an agent retrieves, summarizes, and acts, is the entire workflow a recommender system? The answer will determine whether a legal obligation designed for feeds extends to conversational transactions.

GDPR: data rights without automatic model portability

The General Data Protection Regulation gives individuals rights around lawful processing, access, erasure, transparency, and data portability. Article 20 provides a right to receive certain personal data in a structured, commonly used, machine-readable format and to transmit it to another controller in defined circumstances.

The gatekeeper problem is that the most valuable signal may be inferred rather than directly provided. A user may provide a click, but the platform stores a predicted preference, user embedding, fraud score, or relevance feature. The legal right to portability of personal data does not automatically require the platform to export its model weights or ranking logic.

That boundary is not necessarily a flaw. Model internals can expose trade secrets and security vulnerabilities. It does mean that privacy rights and competition policy must work together. Meaningful switching may require portable preferences and histories while keeping proprietary model architecture protected.

China: algorithm governance and data sovereignty

China’s Provisions on the Administration of Algorithmic Recommendations apply to recommendation services including generation and synthesis, personalized pushing, ranking, search and filtering, and dispatch or decision algorithms. The provisions require safety responsibility systems, regular review of mechanisms and outcomes, disclosure of basic principles and main operating mechanisms, options that are non-personalized or easy to disable, and functions for selecting or deleting user labels.

This model combines information governance, platform responsibility, user control, anti-manipulation provisions, and national-security concerns. It is not simply a privacy regime. It shows that governments may view recommendation systems as infrastructure for social order as well as commercial technology.

United States: antitrust remedies and infrastructure controls

U.S. competition policy has traditionally relied more heavily on case-specific antitrust enforcement than on a single ex ante gatekeeper law. The Google search remedies demonstrate a structural approach: the Department of Justice describes access to certain search-index and user-interaction data, and search or text-ad syndication, as part of the court-ordered response.

The remedy is narrower than a general data-sharing mandate. Its design recognizes that rivals may need certain inputs to achieve scale, while privacy, security, and commercial confidentiality limit what can be disclosed. The test for future cases will be whether the data is an essential competitive input or merely a convenient asset.

The infrastructure layer adds another dimension. The Bureau of Industry and Security’s May 31, 2026 guidance links certain advanced-computing licensing requirements to entities headquartered in Country Group D:5, including China, or Macau, even when the entity is located elsewhere. The policy demonstrates that AI control is no longer only about personal data. Chips, cloud geography, corporate ownership, model access, and end use are becoming connected.

Five principles for regulation

The following framework is more practical than demanding complete algorithmic transparency.

  1. Explain the recommendation at the right level. Users need dominant factors, alternatives, and meaningful control, not a dump of proprietary code.
  2. Disclose commercial influence. A commission, paid placement, platform-owned product, or default tool should be identifiable when it affects a recommendation.
  3. Audit high-impact systems. Independent evaluators should test bias, manipulation, ranking stability, privacy, and failure under adversarial conditions.
  4. Make switching usable. Portability requires formats, authentication, permissions, real-time updates, and a destination that can actually use the data.
  5. Target the bottleneck. Remedies should address the point that prevents competition—data access, defaults, APIs, identity, payments, or distribution—rather than imposing a symbolic obligation no rival can use.

Chapter 8: Case studies of AI data power

Google: from index to answer engine

Google’s original advantage came from organizing the web and selling access to attention. Its search index, ranking system, advertising market, browser, mobile operating system, maps, shopping graph, and user-interaction signals now intersect with generative AI.

Alphabet’s 2025 Form 10-K describes AI as reshaping advertising and records the final judgment in the U.S. search case. Google’s AI Mode describes a search experience that can break questions into multiple searches, synthesize sources, and help evaluate tickets, reservations, and shopping options.

The gatekeeper risk has two layers. First, Google can decide which sources become evidence in an answer. Second, it can decide which action options are presented. A publisher may receive fewer direct clicks even if its information helped train or supply the answer. A small business may be represented accurately or inaccurately without a clear route to appeal.

The opportunity is also substantial. AI search can reduce the cost of research, surface obscure sources, translate information, and help users compare complex options. The public-interest question is whether the system preserves source visibility, attribution, choice, and competition as the interface becomes more concise.

Amazon: commerce as a prediction loop

Amazon has first-party retail, third-party sellers, logistics, payments, advertising, cloud infrastructure, and customer-service signals. The company’s filing identifies AI-enabled discovery as part of its competitive environment. The same integration that creates convenience can create dependence.

For sellers, ranking is access to demand. Inventory, price, delivery, reviews, advertising, returns, and conversion influence visibility. An AI agent could make the shortlist even smaller by completing a purchase after evaluating only a few options.

The policy question is not whether Amazon should recommend products. It is whether sellers can understand the ranking factors, whether first-party products receive advantages unavailable to rivals, and whether paid influence is clear. A marketplace can be efficient and still require contestability.

Meta: social graph plus discovery engine

Meta’s 2025 filing says its AI efforts include content recommendation, an AI-powered discovery engine, advertising delivery, targeting, and measurement, while also describing Llama and a mixture of open and closed models.

The company’s data advantage is not simply the number of posts. It is the relationship among people, content, interactions, devices, identities, advertisers, and feedback. Recommendation controls attention; attention creates advertising inventory; advertising funds infrastructure and new products.

Meta’s open-model strategy complicates the monopoly narrative. Releasing capable models can broaden access and create downstream competition. It can also expand the platform’s influence over tools, developer practices, and distribution. The analytical question is where control remains: weights, data, compute, APIs, user interface, and recommendation policies can be open or closed separately.

TikTok and ByteDance: the ranking asset

TikTok demonstrates that a platform can compete through recommendation quality even when a user’s social graph is less central to discovery. In its official explanation of content recommendations, the company describes a mix of interaction, content, and user-information signals. A new creator can receive distribution based on predicted relevance, but the ranking system decides the test audience, safety eligibility, diversity, and continuation.

This creates a distinctive power structure. The platform does not merely host a social graph; it allocates the path by which a creator might build one. The recommendation engine becomes a market for attention and a mechanism for cultural selection.

As ByteDance operates under regulatory scrutiny and geopolitical pressure, recommendation algorithms also become strategic assets of national policy. The question extends beyond ownership of code. It includes who can audit, govern, transfer, or control the systems that allocate information.

Tesla: physical data and the update channel

Tesla’s 2025 Form 10-K describes AI in Full Self-Driving, Robotaxi, and robots, alongside more than $20 billion of expected 2026 capital expenditure driven partly by AI initiatives, data centers, and AI-enabled assets. The company’s potential advantage is a loop connecting real-world observations, edge cases, fleet operations, simulation, training, validation, and software updates.

Physical-world data is expensive and uneven. A large fleet can observe rare events, but the observations may be noisy, geographically narrow, or difficult to label. Safety and privacy constraints limit what can be collected and reused. The data advantage becomes durable only if the company can turn observations into validated improvements and deploy them safely.

The gatekeeper risk is the installed base. A company that controls the vehicle or robot, the sensors, the update channel, and the service relationship may decide which model improvements reach users. This can create a powerful product loop and a difficult switching environment.

OpenAI: the context broker

OpenAI’s agent product illustrates an interface-centered strategy. The system can browse, use tools, run code, and work with connected context. If users rely on it to research, compare, plan, and act, the company can become a broker of intent even without owning every underlying service.

The most valuable data may be the user’s context: goals, constraints, history, preferences, files, and corrections. The model can learn which sources and tools lead to satisfactory outcomes. This is a high-value feedback loop, but it also creates a concentration risk around personal and enterprise context.

An agent should not become a silent commercial proxy. Users need to know which sources were considered, how external relationships affect the choice, what information the agent used, and how to revoke permissions.

Microsoft: enterprise context as distribution

Microsoft’s 2025 Form 10-K connects Microsoft 365 Copilot, LinkedIn, Dynamics, Power Platform, Azure AI, and agents across the intelligent cloud and edge. Its potential advantage is not a single consumer feed. It is the enterprise context layer: email, documents, meetings, identity, code, business applications, security policies, and permissions.

An enterprise agent that can retrieve and act within those boundaries may be more valuable than a standalone model. It can suggest a document, identify a colleague, draft a customer response, or initiate a workflow. The risks are equally direct. A ranking error can hide a critical document, expose a confidential source, or create a misleading summary that carries institutional authority.

NVIDIA: upstream gatekeeper power

NVIDIA does not decide which hotel a consumer sees, but it may influence which AI systems can be trained and served. Its fiscal 2026 filing describes a full-stack platform spanning GPUs, networking, systems, software, libraries, models, training datasets, and services. It reports more than 7.5 million developers using CUDA and related tools.

The ecosystem advantage is a form of upstream gatekeeping. If access to the dominant compute and software stack affects model cost, latency, and deployment, then infrastructure shapes downstream competition. This does not mean alternatives cannot emerge. It means that model openness alone may not remove dependence if compute, networking, cloud access, and developer tooling remain concentrated.

A comparative view

| Company or ecosystem | Data advantage | AI decision advantage | Market impact | |---|---|---|---| | Google | Queries, clicks, web index, maps, shopping, identity | Retrieval, synthesis, answer ranking, agentic search | Controls access to information and discovery | | Amazon | Transactions, inventory, reviews, fulfillment, seller data | Product ranking, purchase matching, delivery prediction | Allocates commerce demand and seller visibility | | Meta | Social interactions, content, identity, ad outcomes | Feed recommendation, discovery, targeting | Allocates attention and advertising value | | TikTok/ByteDance | High-frequency watch and interaction signals | For You ranking and content discovery | Allocates creator reach and cultural exposure | | Tesla | Fleet sensor data, interventions, edge cases | Driving and robot model updates | Connects physical data to installed hardware | | OpenAI | Prompts, context, tool outcomes, user corrections | Agentic research and action selection | Can mediate intent across services | | Microsoft | Enterprise documents, identity, workflows, code | Copilot and agent recommendations | Shapes workplace information and actions | | NVIDIA | Compute ecosystem, developer tools, model infrastructure | Enables training and inference at scale | Sets upstream conditions for AI competition |

The table does not rank companies by “most data.” It shows that gatekeeper power can appear at several layers. A consumer interface, an enterprise context system, a commerce market, and a compute platform can each control a different route to value.

Chapter 9: AI, data, and geopolitical competition

Data sovereignty has two meanings

Data sovereignty is often reduced to where servers are located. In AI, it has at least two layers.

Jurisdictional sovereignty concerns which government can access, regulate, compel, or restrict data based on where it is collected, stored, processed, or transferred.

Capability sovereignty concerns whether a country has the chips, data centers, models, cloud services, talent, and industrial applications needed to use data at frontier scale.

A nation can possess large domestic datasets but depend on foreign compute. It can build advanced chips but lack high-quality local interaction data. It can train a strong model but depend on foreign operating systems, cloud software, or distribution. AI policy therefore combines data policy with industrial policy.

The U.S.–China competition is a stack competition

The United States and China compete across chips, cloud infrastructure, model capability, data governance, talent, and applications. The BIS 2026 guidance shows how export control can link advanced computing access to corporate headquarters, parent ownership, destination, and end use. China’s algorithmic recommendation provisions show how domestic control over ranking systems is treated as a matter of information governance, user rights, and national security.

Neither system is a simple mirror of the other. The strategic similarity is that both treat AI infrastructure as consequential beyond a private software market. The difference is how authority, market access, information control, and innovation are balanced.

Fragmentation costs

Global firms may need region-specific stacks:

  • separate data storage and processing regions;
  • different model weights, filters, and moderation policies;
  • local cloud and chip suppliers;
  • contracts governing whether enterprise data can be used for training;
  • controls on telemetry and cross-border model improvement;
  • region-specific identity, payments, and audit systems;
  • fallback models when a frontier API or hardware supply is restricted.

Fragmentation can improve resilience and local accountability. It can also reduce economies of scale, duplicate infrastructure, and produce different answers to the same question because data-access rules and ranking objectives vary by jurisdiction. A global recommendation market may become a set of partially compatible regional markets.

National infrastructure versus private gatekeepers

Governments may respond by treating compute, public data, and AI services as national infrastructure. That can support research, security, and public-interest applications. It can also create new gatekeepers if access is allocated by a small number of state-backed or regulated providers.

The policy challenge is to build capability without replacing one opaque intermediary with another. Public infrastructure should publish access rules, protect privacy, support independent evaluation, and enable multiple downstream applications.

China: sovereignty plus managed circulation

China adds a model that is different from both the American competition story and the European rights-and-contestability story. The Data Security Law frames data security as connected to national sovereignty, security, and development interests. Its scope can reach certain processing outside China when national security, public interests, or the rights of people and organizations in China are affected.

The implication is not simply that data should be locked inside national borders. China’s 2024 guidelines on public-data resources describe an effort to make public data available in an orderly way while protecting national security, personal information, and business secrets. The stated target is a more systematic development and utilization of public data resources by 2030.

This is best understood as a sovereignty-and-circulation model. Sensitive data is classified and governed; public and enterprise data is expected to become an economic input through controlled channels. The approach may reduce dependence on foreign cloud and data gatekeepers, but it can create domestic administrative bottlenecks. National control is not the same as competitive access.

China’s existing algorithm-recommendation rules, discussed above, add user controls, disclosure, safety responsibility, and anti-manipulation duties. Together, the policies show that a recommendation system can be treated as commercial infrastructure, information infrastructure, and a national-security concern at the same time.

India: public rails designed to unbundle the intermediary

India offers a different counter-design: use shared protocols and consent rails to let multiple firms participate in the same transaction. The Open Network for Digital Commerce describes a network in which buyers and sellers do not have to use the same application. Its design unbundles buyer experience, seller onboarding, fulfillment, payments, and other functions across participants.

ONDC is not proof that ranking is neutral or that every seller receives equal discovery. The network still needs search, reputation, logistics, identity, and dispute-resolution rules. Its significance is structural: the application that controls the buyer interface does not have to control the whole marketplace. A protocol can become a contestable layer rather than a private company’s exclusive route between supply and demand.

India’s Account Aggregator framework applies a similar idea to financial data. The Ministry of Finance describes consent-based transfer from one financial institution to another on a customer’s instruction, with no information shared without explicit consent. Its March 2026 update reported 179 Financial Information Providers, 989 Financial Information Users, more than 2.88 billion enabled financial accounts, and 284.6 million accounts linked by users. These are official program statistics, not proof that switching costs have already fallen across the entire financial sector.

The value of the framework is that portability becomes an active permission and interoperability service rather than a one-time download. Multiple authorized providers can access usable context without every provider building a private copy of the country’s financial data.

India’s Ayushman Bharat Digital Mission extends the pattern into healthcare. Its documented architecture is federated: records remain in the systems where they are created; there is no single central repository of every health record; and independent systems interoperate through common registries, identity, and consent arrangements. That separates interoperability from centralization. The identity and consent layers remain important infrastructure, but one application does not have to own every record to enable a useful health-data ecosystem.

Africa: sovereignty, local ownership, and regional scale

Africa faces a different gatekeeper problem. The issue is not only a private platform becoming too powerful. Governments, researchers, and local companies may lack sufficient access to datasets, compute, language resources, and bargaining power to build alternatives at all.

The African Union Data Policy Framework seeks to harmonize national data-governance systems, create a shared data space, support intra-African digital trade, protect rights, and enable sustainable development. The AU Continental AI Strategy identifies shortages of high-quality, inclusive, locally produced datasets and warns that much data about African populations is available to a small number of companies. It calls for African-owned, people-centered, and inclusive AI capabilities.

The tension is important. A national silo may protect sovereignty while making data too fragmented to support competitive AI. A common regional space can create scale, but only if it includes local ownership, fair access, security, and rights. The African policy challenge is therefore to combine sovereignty with interoperability rather than choose one against the other.

There are also practical counterexamples. DHIS2 describes an open-source, locally owned health-information system used in more than 80 countries. Each organization can run its own instance, retain local data, modify the software, and choose storage arrangements consistent with local law. Open code does not remove the need for skills, infrastructure, or implementers, but it changes dependency from one proprietary vendor to a local ecosystem that can inspect and adapt the system.

Masakhane’s participatory machine-translation work provides another route around a global data deficit. Researchers and language communities released datasets, benchmarks, models, code, and evaluations for African languages. A local-language data network cannot match the compute of a frontier laboratory, but it can prevent a small set of globally abundant languages from becoming the only gateway to useful AI representation and search.

These cases broaden the definition of a gatekeeper. It can be a private recommender, a national data-residency regime, a public identity service, a cloud provider, or a language model that lacks local coverage. The remedy depends on which layer prevents participation.

Chapter 10: Future scenarios for 2030–2040

The following scenarios are forecasts, not facts. They describe different combinations of technology, policy, and market structure.

Scenario 1: The AI oligopoly world

By 2030, a small group of companies controls most discovery and routine digital action. Search engines, operating systems, commerce platforms, social networks, workplace software, and cloud AI are connected. Model capability becomes more interchangeable, but distribution remains concentrated.

Agents choose sources, vendors, and tools through private ranking systems. Businesses optimize for machine-readable evidence, platform eligibility, and preferred APIs. A consumer may still reach the open web, but the default path from intention to transaction runs through a few interfaces.

By 2040, the largest platforms resemble economic operating systems. They manage identity, context, agent permissions, payments, reputation, and service discovery. Competition authorities restrict self-preferencing and require selected data access, but the combined value of trust, infrastructure, and distribution remains difficult to reproduce.

The benefit is convenience and integration. The cost is reduced pluralism, greater exposure to ranking errors, and a risk that commercial priorities become invisible public infrastructure.

Scenario 2: The open and interoperable AI ecosystem

Standardized agent protocols, portable user context, open identity, interoperable catalogs, and cloud switching reduce dependence on one interface. Users move preferences, histories, permissions, and transaction records among agents. Businesses publish verified machine-readable offerings once and reach multiple systems.

Ranking becomes more competitive because agents can compare sources, explain objectives, and switch providers. Specialized agents compete on privacy, local expertise, public-interest ranking, price, or workflow reliability.

This scenario requires more than open model weights. Portability needs schemas, authentication, permission management, audit trails, real-time updates, and a destination that can actually use the data. Dominant platforms may comply formally while keeping the most important derived signals and transaction capabilities proprietary.

The benefit is user choice and innovation. The cost is fragmentation, inconsistent quality, and the difficulty of coordinating safety across many independent agents.

Scenario 3: A regulated AI economy

Governments create a layered framework. High-impact recommendation and agent systems must disclose commercial influence, provide meaningful explanations, support user controls, maintain audit records, and allow independent evaluation. Competition law focuses on defaults, data access, self-preferencing, and interoperability. Critical infrastructure providers face resilience and security obligations.

Regulation does not eliminate concentration. It changes the terms under which concentration operates. Large firms can still invest more, but they cannot silently turn a ranking system into an unreviewable allocation mechanism.

The benefit is greater trust and rights. The cost is compliance complexity, slower deployment, and the possibility that large firms handle regulation more easily than startups. Regulators will need proportional rules, sandboxes, and shared testing infrastructure.

Scenario 4: Sovereign AI blocs and local intelligence

Export controls, data-residency rules, national-security reviews, and regional cloud markets divide the world into partially separate AI spheres. Companies maintain multiple versions of models and recommendation systems. Local data and domestic distribution become strategic priorities.

At the same time, smaller models, efficient inference, edge devices, and sector-specific data lower the cost of building useful local agents. Enterprises keep sensitive information in private environments and use specialized models rather than one universal platform.

This scenario can produce both concentration and pluralism. A country may depend on one domestic cloud while allowing many local applications. International interoperability declines, but domain-specific competition increases.

When gatekeepers can be bypassed

The concentration thesis becomes too pessimistic if it assumes that every useful AI system must train on one platform’s raw data. Several vertical systems suggest a more constructive path: separate data custody from model training, separate marketplace applications from network protocols, or use synthetic data to let developers build before they obtain sensitive records.

The MELLODDY federated drug-discovery project brought ten pharmaceutical companies into a collaborative learning arrangement without requiring them to centralize their confidential compound datasets. The model aggregated learning updates while the partners retained local data. The project reported improvements for participating classification and regression models. This is a research and industry collaboration, not proof that federated learning solves data governance. The federation still needs secure aggregation, compatible schemas, incentives, defenses against poisoning, and a trusted coordination process.

Healthcare provides a second class of examples. MITRE Synthea generates synthetic patient populations and electronic health records, including an accessible synthetic dataset and APIs for experimentation. Synthetic data lets a startup test schemas, retrieval, workflow design, and evaluation without waiting for a hospital to release raw records. It can lower the entry cost of health-data software.

Synthetic data is not a substitute for representative clinical evidence. It can miss rare conditions, reproduce the assumptions of its generator, or create a false sense of accuracy. A 2024 review in Nature Reviews Bioengineering notes both its potential for privacy, data scarcity, and underrepresentation and the need for rigorous quality and trust measurement. The positive lesson is narrower: synthetic data can be an access-enabling layer that reduces early dependence on a dominant data owner.

Open tooling can create a similar effect. OpenFL provides a reusable framework for collaborative training and evaluation without directly sharing sensitive data. The availability of open code lowers the engineering cost of building a federation. It does not guarantee privacy, secure updates, or economic viability; those must be tested in each deployment.

India’s ONDC, ABDM, and Account Aggregator systems illustrate the protocol version of the same idea. They unbundle the application from the network, let data remain with its origin, and use common identity, consent, and interoperability rules. DHIS2 demonstrates the local-ownership version. Masakhane demonstrates the community-data version.

The common pattern is not “make all data public.” It is to split a gatekeeper bundle into contestable functions:

| Function | Gatekeeper bundle | Bypass or mitigation pattern | |---|---|---| | Data custody | One central repository | Federated local repositories or synthetic sandbox | | Discovery | One application controls ranking | Protocol-level discovery with multiple applications | | Model training | One owner receives raw data | Local training and secure aggregation | | Software | Proprietary workflow and format | Open code, APIs, and common schemas | | Identity and consent | Platform-controlled permissions | Regulated or public consent and identity rails |

The warning is just as important as the optimism. Control can reappear in the coordinator, identity provider, cloud, standard-setting body, gateway, or model host. Federated learning can reduce raw-data concentration while increasing dependence on one orchestration platform. Open source can reduce software lock-in while leaving compute and distribution concentrated. A successful bypass must be measured at the whole stack.

Which scenario is most likely?

These scenarios are not mutually exclusive. Concentrated interfaces can exist inside sovereign blocs. Open agents can operate through dominant app stores. Regulation can improve portability in one region while infrastructure remains concentrated globally.

The most plausible path is a hybrid: a few global infrastructure and interface firms, regulated regional services, and a long tail of specialized agents that compete on proprietary context, workflow integration, and trust. Model intelligence will matter, but the winners will also control data quality, permissions, distribution, and the ability to complete an action.

Chapter 11: The new digital power structure

Factories, platforms, and decision systems

The industrial age rewarded control of production. Factories, machines, logistics, and energy determined who could make goods at scale.

The internet age rewarded control of distribution. Search engines, social networks, marketplaces, app stores, and payment systems determined which producers could reach demand.

The AI age may reward control of decision systems. The key question becomes which information, product, person, source, or action enters the user’s field of attention—and what the software does next.

This is not a claim that humans stop deciding. It is a claim that delegation changes the structure of choice. When users ask for a shortlist, accept a default, or authorize an agent to act, the interface’s objective becomes part of the market. The system can influence outcomes without issuing an order.

The gatekeeper stack

The future control point is best represented as a stack:

  1. Data layer: user activity, public information, transactions, sensors, business records.
  2. Compute layer: chips, data centers, networking, cloud access, energy.
  3. Model layer: foundation models, embeddings, classifiers, ranking policies.
  4. Retrieval layer: indexes, catalogs, APIs, provenance, source selection.
  5. Identity and permission layer: who is the user, what can the agent access, and what may it do?
  6. Interface layer: search, chat, feed, workplace software, marketplace, device.
  7. Transaction layer: payment, booking, purchase, publication, workflow execution.
  8. Trust and governance layer: explanations, audits, appeals, safety, regulation.

A firm does not need to own all eight layers to exercise power. It may control the interface and identity while relying on another model. It may control compute and software while serving many interfaces. It may own a marketplace and use several models. The most strategic position is where layers reinforce one another and make switching difficult.

What should remain human?

The answer is not to prevent delegation. Delegation is one of the reasons to build useful AI. The answer is to identify decisions where people need visibility, alternatives, and appeal.

Low-stakes personalization can be largely automated. High-stakes decisions about employment, credit, healthcare, education, housing, and access to public information require stronger controls. Users should be able to ask what was considered, why an option was selected, what commercial relationships exist, and how to correct an error.

Businesses also need rights. A recommendation system that determines market access should provide machine-readable eligibility information, correction channels, and meaningful data about performance. Otherwise, “fair competition” becomes a promise made by a platform that competitors cannot test.

The difference between data power and legitimacy

An AI system can be accurate and still lack legitimacy if its objective is hidden, its errors cannot be appealed, or its benefits are distributed unfairly. It can be transparent and still be harmful if users have no practical alternative. It can be open and still be concentrated if compute and distribution remain controlled.

Legitimacy requires more than technical quality. It requires a credible relationship among users, businesses, platforms, and regulators. That relationship depends on privacy, contestability, understandable rules, secure portability, and accountability when the system’s recommendation causes harm.

Chapter 12: How to measure gatekeeper intensity

There is no accepted universal “AI Gatekeeper Intensity Index.” That is not a reason to avoid measurement. It is a reason to publish a dashboard whose units, denominator, time window, and source data are explicit. A composite score can be useful for diagnosis, but it should never disguise judgment as a scientific constant.

Data Accessibility Coverage

For a defined use case, score each required data field for usable access:

Data Accessibility Coverage (DAC) = Σ wⱼ × aⱼ / Σ wⱼ

Here, wⱼ is the importance of field j and aⱼ is a 0–1 score combining lawful permission, machine readability, freshness, provenance, quality, granularity, geographic coverage, and API or export reliability. A dataset should not receive a high score merely because it can be downloaded. Stale, aggregated, legally unusable, or unauditable data is not meaningful access.

For example, a local restaurant discovery agent might require identity, location, opening hours, menu, price range, dietary options, reservations, and cancellation policy. A dataset with only a name and address has broad coverage in a formal sense but low coverage for the actual task.

DAC should be reported by dimension. A platform can have excellent data quality but low portability, or high portability but poor freshness. Combining those into one number too early hides the bottleneck.

Portability and switching friction

Policy language often promises portability without testing whether a new provider can use the data. A practical measure is a migration experiment:

Portability Success Rate = priority tasks completed after import / priority tasks tested

Report median migration time, engineering hours, direct fees, downtime, re-consent rate, schema-conversion loss, and quality loss. For an AI agent, priority tasks might include importing preferences, retrieving a prior conversation, restoring a business catalog, calling a payment provider, or recreating a workflow in a rival system.

Switching Friction Cost = one-time migration cost + recurring interoperability cost + expected failure cost

The OECD’s work on data portability, interoperability, and competition explains why usable portability can lower switching costs. The formulas above are this article’s analytical proposals, not established legal metrics. The test should include both an individual and a small business. Portability that works only for an enterprise with a dedicated data team is not meaningful contestability.

Recommendation concentration and alternative exposure

For a defined query set, calculate the share of impressions, shortlist slots, or completed transactions attributed to each supplier or source. Apply a Herfindahl–Hirschman-style concentration measure:

Recommendation Concentration Index (RCI) = Σ sᵢ²

where sᵢ is supplier i’s percentage share of exposure or action. The U.S. Department of Justice explains HHI as the sum of squared market shares. Its familiar merger-screening thresholds should be treated only as a reference; an AI recommender is not automatically an antitrust market.

RCI should be paired with four operational measures:

  • Top-K alternative rate: the percentage of eligible suppliers that appear in the top K across matched queries.
  • New-entrant exposure rate: the share of impressions received by qualified suppliers below a defined age or prior-volume threshold.
  • Geographic and language coverage: the share of exposure going to local regions or languages eligible under the request.
  • Commercial disclosure rate: the percentage of recommendations where paid placement, affiliate relationships, or platform-owned inventory is clearly disclosed.

These measures distinguish concentration caused by legitimate quality differences from concentration created by a closed ranking loop. The DSA’s recommender-transparency and non-profiled-option rules provide a legal reason to ask for this kind of evidence, even though the law does not prescribe these formulas.

Compute-provider concentration and failover

Data access is only one gate. A vertical may have open, federated data and still depend on one cloud, accelerator stack, or model API.

Compute Provider Concentration = Σ cᵢ²

where cᵢ is provider i’s share of inference spend, reserved capacity, or production requests. Report training, inference, and networking separately. Add a failover test: what percentage of production traffic can move to a second provider within a stated time and quality or cost tolerance?

A company with 80% compute concentration but a tested two-week failover is exposed differently from a company with 50% concentration and no viable alternative. The metric should therefore combine concentration with recovery time and performance loss.

Compliance Burden Ratio

Rules can protect users and still impose uneven operational costs. Measure the burden rather than assuming it:

Compliance Burden Ratio (CBR) = annual privacy, security, audit, localization, consent, reporting, and legal-review cost / relevant revenue

Report cash expense, staff hours, and one-time implementation separately from recurring cost. For a public-interest project or early-stage company, use operating budget and show the denominator rather than comparing CBR with commercial revenue.

CBR is not an argument that privacy or sovereignty safeguards are wasteful. It identifies whether a rule is operationally contestable and whether shared compliance tooling, standards, or regulatory sandboxes are needed. A large incumbent may absorb the same absolute compliance work at a much lower ratio than a startup or local provider.

Action Authority Ratio

An agent that summarizes information has less direct allocation power than one that can purchase, publish, hire, approve, or route money.

Action Authority Ratio (AAR) = value-weighted automated actions without human confirmation / value-weighted total actions

Report transaction type, reversibility, error rate, appeal rate, and the maximum value or risk of an action. A high AAR is not automatically bad; it is a measure of how much decision authority sits behind the interface. It should trigger stronger testing in high-impact domains.

A cautious composite score

If a reader needs one dashboard number, a proposed Gatekeeper Intensity Score (GIS) could combine normalized values:

GIS = w₁·RCI + w₂·(1 − DAC) + w₃·SFI + w₄·PCI + w₅·CBR + w₆·AAR

Here, SFI is normalized switching friction and PCI is compute-provider concentration; the weights sum to one. These weights are not validated. A credible publication should publish the full vector, run sensitivity analysis across weights, show confidence intervals, and avoid ranking countries or companies from a single composite.

The better use is diagnostic. If GIS is high, ask why: is the problem recommendation concentration, inaccessible data, infrastructure dependence, compliance asymmetry, or excessive action authority? Different causes require different remedies. The index is useful only when it sends the reader back to the underlying evidence.

Chapter 13: What should be portable, auditable, or shared?

The policy choice is not total openness versus total secrecy. Different layers require different remedies.

| Layer | Typical control point | Possible remedy | Main safeguard | |---|---|---|---| | User-provided data | Posts, files, profiles, transactions | Export, portability, interoperable APIs | Consent, privacy, security | | Interaction data | Queries, clicks, watch time, purchases | Narrow access for qualified researchers or rivals | Anonymization, aggregation, purpose limits | | Derived signals | Embeddings, inferred interests, rankings | Explanations, controls, audit access | Trade secrets, privacy, anti-gaming controls | | Public index | Web pages, product and business facts | Provenance and citation rules | Copyright, attribution, freshness | | Model and retrieval | Weights, ranking logic, tool policy | Testing, audit, non-discrimination | Safety, abuse prevention, intellectual property | | Infrastructure | Chips, cloud, data centers, identity | Switching standards and resilience rules | National security, reliability, supply chains |

The remedy should match the bottleneck. Requiring a platform to publish model weights may not help a small hotel if the problem is opaque retrieval or lack of transaction access. Requiring raw data sharing without privacy controls may harm users. A non-profiled feed can improve user choice without creating a competitor. A data-access rule is useful only if a qualified rival can turn the data into a viable product.

Independent audits should test more than bias in a frozen model. They should examine how the system changes ranking when a commercial relationship is introduced, whether new entrants receive meaningful exploration, how errors are corrected, what data is retained, how the agent behaves under conflicting objectives, and whether users can actually switch.

Chapter 14: Recommendations for executives, founders, and policymakers

For technology executives

Treat data governance as product infrastructure. Document which data is collected, why it is needed, how long it is retained, whether it is used for training, and how users can revoke access. Separate relevance from commercial influence. Create internal evaluation for ranking stability, recommendation bias, source attribution, and agent action.

Do not assume that a larger model solves the gatekeeper problem. A more capable system can make a better recommendation or a more persuasive error. Invest in provenance, permissioning, fallback modes, and user-visible controls.

For founders and small businesses

Build machine-readable identity before buying “AI optimization.” Keep facts current across first-party pages, profiles, feeds, and authorized APIs. Collect independent evidence of outcomes rather than generating generic content. Monitor whether agents describe the business accurately and create a correction process.

Look for opportunities around the gatekeeper stack: provenance, catalog interoperability, agent permissions, evaluation, audit logs, local data, vertical workflow, and trustworthy transaction infrastructure. A narrow product with reliable outcome data can be more defensible than a generic chatbot layered over a public API.

For investors

Separate data volume from data advantage. Ask:

  • Is the data exclusive or merely abundant?
  • Does it contain outcome feedback or only content?
  • Can the company legally and technically use it?
  • Does the model improve a measurable workflow?
  • Does the product control distribution or depend on another gatekeeper?
  • Can users and businesses switch?
  • Is the moat in data, compute, model quality, trust, or integration?

Valuation should reflect the cost of maintaining the feedback loop, not just the size of the dataset. Data centers, labeling, privacy, security, and customer acquisition can overwhelm an apparently attractive advantage.

For policymakers

Define the market at the decision surface. A search engine, feed, marketplace, workplace copilot, and agent may belong to different markets, but each can become an intermediary that controls access to opportunity. Competition analysis should follow the user journey from discovery to action.

Require useful explanations and commercial disclosure without demanding that companies reveal security-sensitive code. Support secure data access for qualified rivals and researchers. Fund public-interest testing. Coordinate privacy, competition, consumer-protection, and AI-safety agencies so that an organization cannot shift responsibility from one legal silo to another.

Finally, preserve room for pluralism. Public procurement, open standards, interoperable identity, and portable business catalogs can help prevent a few private agents from becoming the only practical route to participation in the digital economy.

Conclusion: Is data the most valuable asset of the AI era?

Data is one of the most valuable assets of the AI era, but the answer is not simply yes. Data without relevant outcomes is noise. Data without compute cannot be turned into a capable model. Data without distribution does not generate a feedback loop. Data without trust may be unusable. Data without legal permission creates liability rather than advantage.

The real strategic asset is a system:

Data + Computing + Models + Distribution + Trust + Regulatory Legitimacy

Whoever controls several of these layers can influence not only what people know, but what they are shown, which businesses they discover, which products they buy, which sources they trust, and which actions their software is authorized to take.

That power can create enormous public value. Better recommendations can lower search costs, help small businesses find customers, make complex information more accessible, and let people delegate tedious work. It can also create a new form of invisible market control. A company may not need to own every product or publisher if its AI decides which ones enter the user’s consideration set.

The future battle may therefore not be “Who has the biggest AI model?” It may be “Who controls the intelligence layer between humans and the economy?”

The answer should not be left entirely to the systems that benefit from becoming gatekeepers. Users need meaningful choice. Businesses need fair access. Researchers need evidence. Regulators need remedies that match the actual bottleneck. And the designers of AI systems need to remember that a recommendation is never only a prediction when it changes what happens next.

If AI becomes the invisible decision-maker of society, the greatest challenge will not only be building intelligent machines, but ensuring that intelligence remains aligned with human interests.

References

  1. Google Research, “The Unreasonable Effectiveness of Data” (2009) — supports the technical argument that large, relevant datasets can improve statistical language systems; it does not by itself prove monopoly power.
  2. Meta AI Research, “The Architectural Implications of Facebook’s DNN-based Personalized Recommendation” (2020) — documents the production-scale architecture and infrastructure demands of recommendation models.
  3. Alphabet 2025 Form 10-K — provides company disclosure on AI, advertising, search distribution, and the final judgment in the U.S. search case.
  4. Amazon 2025 Form 10-K — supports analysis of marketplace sellers, advertising, fulfillment, AI-enabled discovery, and competition.
  5. Meta Platforms 2025 Form 10-K — describes AI-powered content discovery, advertising, Llama, infrastructure, and the company’s data-driven business.
  6. Microsoft 2025 Form 10-K — supports the enterprise-context analysis of Copilot, LinkedIn, Azure AI, Microsoft 365, and agents.
  7. Tesla 2025 Form 10-K — provides company disclosure on Full Self-Driving, Robotaxi, robots, AI-enabled assets, and AI-related capital expenditure.
  8. NVIDIA fiscal 2026 Form 10-K — supports the upstream infrastructure analysis of GPUs, networking, CUDA, models, training data, and the developer ecosystem.
  9. TikTok, “How TikTok recommends content” — provides TikTok’s own description of interaction, content, and user-information signals; it is not a complete ranking audit.
  10. OpenAI, “Introducing ChatGPT agent” (2025) — documents the product’s stated browser interaction, research, code execution, and tool-use workflows.
  11. Google, “AI Mode in Google Search: Updates from Google I/O 2025” — supports the discussion of query fan-out, Deep Search, shopping, reservations, and agentic search.
  12. Google Search Central, “AI features and your website” — provides current official guidance on crawlability, structured data, business information, and AI-search eligibility.
  13. U.S. Department of Justice, “Department of Justice Wins Significant Remedies Against Google” (2025) — establishes the government’s description of search-data access, syndication, and distribution remedies.
  14. Regulation (EU) 2024/1689, Artificial Intelligence Act — provides the enacted EU risk-based framework for AI accountability, transparency, safety, and fundamental rights.
  15. Regulation (EU) 2022/1925, Digital Markets Act — establishes the EU gatekeeper framework covering contestability, fairness, interoperability, data use, and portability.
  16. Regulation (EU) 2022/2065, Digital Services Act — supports recommender transparency, non-profiled alternatives, advertising transparency, systemic-risk duties, and researcher access.
  17. Regulation (EU) 2016/679, General Data Protection Regulation — establishes data-subject rights, including Article 20 portability, and helps distinguish personal data from derived model signals.
  18. China CAC, Provisions on the Administration of Algorithmic Recommendations — supports the discussion of algorithm responsibility, disclosure, user controls, label deletion, and anti-manipulation obligations.
  19. U.S. Bureau of Industry and Security, Advanced Computing Guidance (May 31, 2026) — documents how advanced-compute licensing connects AI infrastructure, corporate headquarters, cloud geography, and export controls.
  20. China Data Security Law — establishes the sovereignty, security, and development framing used to analyze China’s controlled data circulation model.
  21. China Government, Public-Data Resource Guidelines (2024) — supports the distinction between data sovereignty and regulated public-data utilization.
  22. Open Network for Digital Commerce (ONDC) — documents India’s protocol-oriented attempt to unbundle buyer, seller, logistics, and payment applications.
  23. Government of India, Account Aggregator Framework — provides official evidence for consent-based, interoperable financial data sharing and program adoption figures.
  24. Ayushman Bharat Digital Mission, Federated Architecture FAQ — supports the Indian healthcare example in which records remain distributed while systems interoperate.
  25. African Union Data Policy Framework — supports the analysis of shared African data space, cross-border data governance, rights, and digital trade.
  26. African Union Continental AI Strategy — documents Africa’s data-access, local-language, local-ownership, and inclusive-AI challenges.
  27. DHIS2, Local Ownership and Data Sovereignty — provides an open-source, locally operated health-information counterexample to proprietary data centralization.
  28. Masakhane, Participatory Research for African-Language Machine Translation — supports the community-owned language-data network example.
  29. MELLODDY Federated Drug Discovery — provides a research and industry example of collaborative model training without centralizing confidential pharmaceutical datasets.
  30. MITRE Synthea — documents an open synthetic-health-data resource that lowers early experimentation barriers without exposing patient records.
  31. OpenFL Documentation — supports the availability of reusable open infrastructure for federated training and evaluation.
  32. Nature Reviews Bioengineering, Synthetic Data in Biomedicine (2024) — supports the balanced assessment of synthetic data’s privacy and coverage benefits alongside quality and trust limitations.
  33. OECD, Data Portability, Interoperability and Competition — supports the economic rationale for measuring usable portability rather than merely formal export rights.
  34. U.S. Department of Justice, Herfindahl–Hirschman Index — provides the official definition of HHI used as a screening reference for recommendation concentration.