AI Portfolio 002: Compute
The first layer of the AI stack is also the most physical one: chips, memory, networking, power and cooling. What is fact, what is thesis, and what I am not claiming.
The first note in this series argued that artificial intelligence is crossing a boundary. It is moving from software that helps a person work to software that can operate parts of an economic system. I sketched a mental portfolio with six layers where that shift could show up economically: compute, models, data and context, agentic systems, economic rails, and physical AI. If you have not read that note, the short version is enough to follow this one. If you have, this is where we open the first layer in depth.
Compute.
It is the layer everyone already understands is expensive. It is also the layer most people describe imprecisely, as if "GPUs" were the whole story. It is not. Compute is chips, but it is also memory, networking, buildings, power and cooling, bundled together and sold, in effect, as a single unit of capacity.
What compute actually is
Start with the chip. A GPU or a purpose-built accelerator does the arithmetic that trains and runs a model. But a chip by itself does almost nothing. It needs memory sitting next to it, fast enough to feed it data without starving it. It needs networking, so thousands of chips can act as one machine during training and so a fleet of chips can serve many requests at once during inference. It needs a building with enough power delivered to the rack and enough cooling removed from the rack, because a modern AI server draws far more power per square foot than the data centers built for the previous computing era.
Training and inference are different economic problems inside the same layer. Training is a capital-intensive, front-loaded exercise: build a very large cluster, run it for months, produce a model. Inference is an operating cost that scales with usage: every query, every agent action, every token generated draws power and occupies capacity, continuously, for as long as the product exists. The industry spent the last few years mostly talking about training. I think the more interesting economics now sit in inference, because that is the cost line that scales with adoption rather than with a single build cycle.
I already see this as an operator, not only as someone reading earnings calls. Running models in production means watching a line item that moves with usage, not one that gets negotiated once a year. That single fact changes how I read every AI company's cost structure. A company whose usage costs scale roughly with revenue behaves very differently from one whose usage costs scale faster than revenue.
Networking is the part of compute that gets the least attention and probably deserves more. Training a large model is not one chip doing one job. It is thousands of chips that need to exchange intermediate results constantly, fast enough that the network does not become the slowest part of the system. Inference has a related but different networking problem: routing many simultaneous requests to available capacity without adding latency a user can feel. My thesis is that interconnect bandwidth, not raw chip count, increasingly decides how much useful training and serving capacity a given cluster actually delivers. That is a thesis, not a fact. I have not sourced a dated, cited figure for it here, and I am marking it as such rather than dressing it up as settled.
Cooling follows directly from power density. A modern AI server rack draws far more power, and produces far more heat, than the racks that filled data centers a decade ago. Air cooling struggles at that density. That is why liquid cooling, direct to chip in many new builds, has gone from a specialty technique to something close to a default assumption in new AI data center design. I state that as an industry thesis, grounded in what is publicly visible about how new capacity is being built, not as a sourced statistic.
The demand is real and it is documented
I do not need to speculate about whether AI compute demand exists. It shows up in earnings.
NVIDIA's Data Center segment reported $89.0 billion in revenue for the quarter ended July 26, 2026, up 117 percent from a year earlier (NVIDIA Corporation, quarterly results, as of August 27, 2026). That is one company, one quarter, and it is still worth sitting with for a moment. It means demand for training and inference capacity grew faster than the chips could apparently be produced and deployed.
Microsoft told investors it expects roughly $190 billion in capital expenditure through calendar year 2026, a 61 percent increase over the prior year, and that it expects Azure to remain capacity constrained through the year regardless (Microsoft fiscal 2026 Q3 earnings call, as reported by The Motley Fool, as of July 27, 2026). Read that sentence again. A company spending at that scale is telling its own shareholders it still cannot build fast enough to meet demand. That is not a company hedging. That is a company describing a bottleneck it has not yet solved with money alone.
The bottleneck is not only chips. Memory is its own constraint. SK Hynix told investors that its 2026 DRAM, NAND and high bandwidth memory production is essentially sold out, much of it committed to Nvidia (SK Hynix Q3 2025 earnings call, as reported by Bloomberg, as of October 28, 2025). High bandwidth memory takes meaningfully more wafer capacity per gigabyte than ordinary memory. When an accelerator maker cannot get enough of it, the constraint on how much AI compute exists in the world shifts from chip design to a much narrower set of memory fabrication lines.
And then there is power. The International Energy Agency projects that global data center electricity consumption will nearly double to around 945 terawatt-hours by 2030 in its base case, growing about 15 percent a year from 2024 onward, with demand from AI-optimized data centers more than quadrupling over that period (IEA, "Energy and AI," as of April 10, 2025). Chips can, in principle, be manufactured faster than new grid capacity and cooling infrastructure can be permitted and built. That asymmetry is why several hyperscalers are now negotiating directly with power generators and, in some cases, financing generation themselves. The bottleneck moved. It used to be the chip. Increasingly, it is the substation.
Industry thesis, company thesis, actual position
I want to be precise about what kind of claim I am making, because the word "Portfolio" in this series title is not evidence that I hold anything. It is a mental model I am building in public, one layer at a time. Three different things can be true at once, and I want to keep them separate every time I write about a company.
An industry thesis is a view on the sector: I believe compute will remain a real economic bottleneck for AI for years, not quarters, because power and advanced memory cannot scale as fast as demand.
A company thesis is a view on a specific company or technology, without any claim of ownership: I think a company that controls scarce power capacity, advanced packaging, or high bandwidth memory supply has more durable leverage than a company that only assembles servers from parts anyone can buy.
An actual position is a disclosed holding. I am not asserting one here. None has been independently verified elsewhere in this corpus as of this note, and I am not going to imply one by naming a company as if I owned it.
Readers should hold me to that distinction every time this series names a company.
What I would look for in this layer
Using the same lens from Note 001: bottleneck, distribution, data advantage, switching cost, economic participation, operating leverage, survival. In compute specifically, the bottleneck question matters most right now. Does a company control something scarce, whether that is advanced packaging, high bandwidth memory supply, or contracted power, or is it simply buying the same commodity everyone else can buy? Survival matters almost as much. Building or operating at this scale requires balance sheets that can absorb multi-year capital commitments before revenue catches up, and not every entrant will make it through that gap.
None of that is a stock pick. It is a checklist for reading a company's disclosures, not a recommendation.
What would change my mind
I would take this thesis less seriously if capital expenditure kept climbing while data center utilization stayed flat, which would suggest hyperscalers were building ahead of real demand rather than behind it, the opposite of what current earnings and capacity commentary describe today. I would also revise it if new memory or interconnect technology relieved the HBM bottleneck faster than currently sold-out capacity implies, or if a genuinely more efficient training or inference architecture reduced compute intensity per unit of useful output enough to blunt the power constraint. None of that has happened yet. All of it is worth watching, and I will note here when it does.
What this note does not claim
It does not claim a ticker. It does not claim a holding. It does not claim that compute is the only layer worth attention, only that it is the layer with the clearest, most dated evidence behind it right now. The next note in this series opens Models: what happens to pricing power and margin as the models running on top of all this compute keep getting better, cheaper and more interchangeable.
Subscribe to ONCHAIN FX to follow the thesis as it evolves.

Read in Portuguese: Portfólio de IA 002: Computação
Subscribe to ONCHAIN FX to follow the thesis as it evolves.