Back to research
5 min readActive thesis

The Decentralization Trade: Why GPU and Memory Demand May Disappoint

Models are commoditizing faster than expected, and companies will increasingly self-host instead of renting from hyperscalers. The opportunity is in software.

Educational content only. Nothing on this page is financial or investment advice. Full disclaimer

The Decentralization Trade: Why GPU and Memory Demand May Disappoint

The consensus AI trade has been simple for three years: a handful of companies will spend whatever it takes to build the biggest datacenters, train the biggest models, and rent access to them to everyone else. Every capex forecast, every memory and GPU demand curve, every hyperscaler multiple is built on that assumption. I think the assumption is wrong, and the market has not started pricing the alternative.

Models were always going to end up open and cheap to run

Frontier AI labs spent years positioning proprietary, closed models as the only credible option for serious use cases. That moat is eroding faster than most forecasts assume. Open-weight models are closing the capability gap with the frontier on a shorter timeline every cycle, and the cost of running a good-enough model keeps falling. Once a capability becomes replicable outside the handful of labs that first built it, it stops behaving like a scarce asset and starts behaving like a commodity input. That is the direction this is heading, whether or not it is complete yet.

A commodity input does not get rented from a small number of gatekeepers forever. It gets owned.

Why companies will host their own models

The market's mental model is still "small and medium companies will call an API hosted by a hyperscaler." I think that undersells how quickly the calculus flips once open-weight models are good enough. Once a company can run a model that meets its needs on infrastructure it controls, sending proprietary data to a third party stops being a convenience and starts being a liability: a privacy exposure, an IP exposure, and a dependency on a vendor whose incentives are not the same as the client's.

That tension between platform companies and the AI providers they rely on is already visible. Apple's public friction with OpenAI over the terms of their arrangement is just one example, not the whole argument, but it illustrates the point: even the largest, best-capitalized companies in the world are uncomfortable being structurally dependent on an external model provider for something as sensitive as how their platform and their users' data are handled. Smaller companies with real IP to protect will reach the same conclusion faster, because they have less leverage to negotiate around it.

The direction this points is a slow rotation away from one giant centralized model serving everyone, and toward each company running its own model, on its own stack, sized for its own workload.

This means GPU and memory demand is not what the market is pricing

This is the part of the thesis that actually moves numbers. The current capex and memory demand forecasts are built on the idea that a small number of buyers will need effectively unlimited compute and memory, scaling toward infinity as usage grows. Decentralization breaks that assumption in two ways.

First, no individual company hosting its own model needs infinite compute or infinite memory. It needs enough compute and memory to serve its own workload, which is a bounded, specific number, not an open-ended one. A thousand companies each provisioning "enough" is a very different demand curve than a handful of hyperscalers provisioning "as much as physically possible," even if the sum of the parts is large.

Second, individual users are heading the same direction. Local models are improving quickly, and most individual usage is narrow: a few repeated tasks, not arbitrary open-ended reasoning across every domain at once. At some point, running a smaller model locally, tuned to what a person actually does, starts to beat routing every query through a hosted frontier model. That is another chunk of demand that never shows up as GPU-hours purchased from a centralized provider.

None of this means GPU and memory demand goes to zero. It means the shape of the curve is different from what is currently priced in, and "different from priced in" is where trades are made.

Where the real opportunity was always sitting

If compute decentralizes, the monopoly premium currently priced into the hosting layer, the businesses that assumed they would own the chokepoint between every company and its AI capability, is at risk of being wrong. That premium gets built on the idea that access to frontier compute is the scarce resource. Once companies can provision their own, the scarcity moves elsewhere.

It moves to the businesses that actually have the customers: the software companies with real distribution, real workflows, and real client relationships, the ones who can now run their own model behind their own product instead of paying a toll to a hosting layer for the privilege. The value was never really in owning the biggest datacenter. It was in owning the demand, the customer, and the use case that the datacenter was rented out to serve. Decentralization does not remove that value, it just returns it to where it was always going to end up.

What would change my mind

  1. The capability gap stays wide. If frontier proprietary models keep pulling meaningfully ahead of open-weight alternatives instead of the gap closing, the incentive to self-host weakens and the centralized rental model holds up longer than this thesis assumes.
  2. Regulation favors certified large providers. If compliance regimes end up requiring AI workloads to run through a small number of accredited hosting providers rather than allowing companies to self-host, that would reverse the direction of this argument entirely.
  3. Self-hosting costs more than it saves. If the total cost of running, securing, and maintaining a private model stack stays higher than an API bill for most companies, the economics never flip and centralized hosting remains the default.

Absent those, the direction of travel favors fragmentation over concentration.

The bottom line

The obvious observation, that inference needs more compute as AI usage grows, is not the interesting part of this story. The interesting part is who ends up owning that compute and who captures the value sitting on top of it. I think the answer is more companies, running smaller, private, purpose-built model stacks, not a handful of hyperscalers renting out an ever-expanding fortress of GPUs. That is a real risk to the demand curve currently priced into GPUs and memory, and it is a reason to look past the hosting layer toward the software companies that actually own the customer.

We track setups like this on higher timeframes across the market. If you want to see how this thesis evolves in real positions, follow the Midas Index.

This article is for educational purposes only and is not financial advice.

The conversation continues inside the Cabal Network

Public research is only the surface. Members get every thesis and position update live, for $19.99 a month.

  • Private Telegram group with deeper discussion of every investment thesis
  • Early notifications on new investments, thesis updates and trade swings
  • Capital rotation alerts before they hit the public research feed
Get Cabal Network Access

Read next

Never miss a new article

Get an email whenever we publish new research. No spam, unsubscribe anytime.

Masari Cabal

A private sovereign syndicate built for higher-timeframe swing trading, strict capital preservation, and systemic macro wealth compounding. Powered securely by Whop.

Risk Disclosure

Trading financial assets involves substantial risk of capital loss. Our custom TradingView mathematical indicators and private Telegram channels are strictly for research and educational functions. We are not registered brokers or fiduciary financial advisors.

© 2026 Masari Cabal. All rights reserved.

Handled securely via encrypted Whop processing hubs.