← Nancy's Stack
Deep Dive · Semiconductor & AI Infrastructure

How AI Is
Actually Made

The complete 20-stage production chain — from sand and rare earth minerals to the AI assistants, cloud platforms, and enterprise software that runs the modern world. Every key company, every ticker, every role in the chain.

20
Production Stages
$975B
2026 Global Chip Market
$500B
AI Chips Alone
$7.6T
AI Capex 2026–2031
Company Tier:
Dominant / Monopoly position
Major player
Notable / Rising
01
Raw Materials
The Earth's Crust — Where Everything Starts
What Happens
Everything begins underground. Silicon — the core of every chip — comes from quartz sand, one of the most abundant minerals on Earth. But abundance doesn't mean simplicity: the processing chain transforms raw sand through multiple refinement steps before it becomes usable. Beyond silicon, AI infrastructure requires rare earth elements (neodymium for motors and magnets, lanthanum for precision optics, dysprosium for magnets in cooling systems), cobalt and lithium for backup power batteries, tantalum for capacitors inside chips, copper for all wiring, and increasingly uranium as data centres shift to nuclear power. China controls approximately 85% of global rare earth processing — making export restrictions a potential systemic risk to every stage downstream.
Critical chokepoint: A single Chinese restriction on rare earth exports could constrain global chip production within months. This geopolitical risk sits underneath the entire AI supply chain.
Key Companies
MPMP Materials — only US rare earth miner/processorDominant
LynasLynas Rare Earths — largest ex-China rare earth companyMajor
RIORio Tinto — iron ore, copper, lithiumMajor
FCXFreeport-McMoRan — world's largest copper minerMajor
UUUUEnergy Fuels — US uranium + rare earthsNotable
02
Silicon Purification & Wafer Production
Sand → Crystal → Wafer: The Blank Canvas
What Happens
Raw quartz sand is chemically processed into polysilicon — silicon purified to 99.9999999% (nine nines). This polysilicon is then melted at 1,414°C and a seed crystal is slowly pulled from the melt while rotating, growing a single-crystal cylindrical boule via the Czochralski process. The boule is sliced into discs called wafers — typically 300mm in diameter for advanced chips — then polished to atomic-level flatness (variations of less than 1 nanometre across the entire surface). These wafers become the blank canvas that every chip is printed on. The purity requirements are extraordinary: a single atom of the wrong element can destroy a chip's electrical properties.
Japan and South Korea dominate this stage, with Shin-Etsu and Sumco controlling over 50% of global silicon wafer supply between them.
Key Companies
Shin-EtsuShin-Etsu Chemical (4063.T) — world's largest silicon wafer makerDominant
SumcoSumco Corp (3436.T) — #2 silicon wafer producer globallyDominant
WackerWacker Chemie (WCH.DE) — polysilicon leaderMajor
GWC.TWGlobalWafers — #3 silicon wafer producerMajor
SK SiltronSK Siltron — Samsung subsidiary, growing rapidlyNotable
— Design Layer —
03
Electronic Design Automation (EDA)
The Software That Designs Chips — The Invisible Gatekeeper
What Happens
Modern AI chips contain 80–100+ billion transistors. No human team can manually design the layout of 80 billion components — it requires Electronic Design Automation (EDA) software. Engineers specify what a chip should do in code (like a programming language for hardware), and EDA tools automatically synthesize, place, and route billions of transistors across multiple layers of silicon to implement that design. EDA tools also simulate the chip's behaviour across millions of conditions before a single wafer is processed. Without EDA software, there are no new chips. The US weaponised this in 2022–2023 by restricting Synopsys and Cadence tools from Chinese chip designers — one of the most devastating export controls possible, because it struck at the design stage before any manufacturing occurs.
Synopsys + Cadence command over 70% of the EDA market globally — a genuine software duopoly at the foundation of the entire semiconductor industry.
Key Companies
SNPSSynopsys — EDA market leader, also acquiring ANSYS (simulation)Dominant
CDNSCadence Design Systems — EDA #2, strong in custom/analog chipsDominant
SIEGYSiemens EDA (Mentor Graphics) — #3, acquired by Siemens in 2017Major
ANSYSANSYS — simulation software (being acquired by Synopsys)Notable
04
Chip Architecture & IP Licensing
The Grammar of Computing — ARM
What Happens
Before designing a chip, engineers need to choose an Instruction Set Architecture (ISA) — the fundamental "language" the chip will speak, defining what instructions it can execute and how it handles data. ARM dominates: every iPhone, every Android phone, every Apple Mac, every AWS Graviton server, and NVIDIA's Grace CPU uses ARM architecture. ARM does not manufacture chips — it designs the architecture and licenses it. Chip designers pay ARM an upfront licence fee (millions of dollars) plus a royalty of roughly 1–2% of every chip sold that uses ARM's architecture. This creates a royalty stream that compounds as global chip volumes grow. The challenger is RISC-V — an open-source ISA that companies like Google and SiFive are pushing as an alternative to avoid ARM licence fees.
ARM architecture runs in 99% of smartphones globally and is rapidly capturing AI servers, data centres, and embedded AI devices. It is the architectural standard of modern computing.
Key Companies
ARMArm Holdings — dominant ISA licensor, 99% smartphone marketDominant
RISC-VRISC-V International — open-source ISA alternative, growing in AI/embeddedNotable
INTCIntel — x86 ISA (dominant in servers/PCs but losing AI to ARM)Major
SiFiveSiFive — RISC-V chip designer (private)Notable
05
Chip Design — Fabless Semiconductor Companies
Designing the AI Engines Without Making Them
What Happens
Fabless companies design chips entirely in software using EDA tools, send the blueprint to a foundry, and never touch manufacturing equipment. NVIDIA designs the H100/H200/Blackwell GPUs that dominate AI training — the actual blueprint goes to TSMC in Taiwan to manufacture. The critical distinction in 2026 is between general-purpose GPUs (NVIDIA — can run any AI workload) and custom ASICs (Application-Specific Integrated Circuits — designed for one specific task). Amazon's Trainium/Inferentia, Google's TPU, and Marvell's custom XPUs for Amazon and Google are custom chips that run 3–5x more efficiently for specific AI inference tasks than a general NVIDIA GPU. As AI workloads scale to billions of inferences per day, the economics of custom silicon become compelling — each chip costs more to design but far less to run.
NVIDIA holds 70–80% of the AI training GPU market. The race is now for the inference market — where custom ASICs could displace general-purpose GPUs for specific applications.
Key Companies
NVDANVIDIA — H100/H200/Blackwell GPUs, 70-80% AI GPU market shareDominant
MRVLMarvell — custom XPUs for Amazon + Google, networking ASICsMajor
AVGOBroadcom — custom AI chips for Google TPU and MetaMajor
AMDAMD — MI300X GPU challenger, CPU for AI servers (EPYC)Major
QCOMQualcomm — edge AI chips for smartphones and PCsNotable
INTCIntel — Gaudi AI accelerators, seeking AI relevance after GPU missNotable
— Manufacturing Layer —
06
Lithography Equipment
The Machine That Makes the Machine — ASML's EUV Monopoly
What Happens
Chip blueprints are printed onto silicon wafers using lithography — projecting circuit patterns onto photosensitive material using light. For the most advanced chips (3nm, 2nm, 1.8nm), only Extreme Ultraviolet (EUV) light is precise enough. EUV light is produced by firing a 50,000-pulse-per-second laser at tin droplets, creating plasma that emits light with a 13.5 nanometre wavelength. This light is reflected through a series of mirrors with surface tolerances measured in fractions of atoms. Each ASML EUV machine contains 100,000+ components, 2km of cabling, weighs 180 tonnes, costs $380 million, and is shipped in 40 cargo containers. ASML is the only company on Earth that can manufacture EUV machines — a true technological monopoly. Without ASML, TSMC cannot make NVIDIA's next GPU. The entire advanced semiconductor industry rests on one Dutch company in Eindhoven.
ASML's "high-NA EUV" machines (next generation, €380M+ each) are already on order for the sub-2nm nodes needed for AI chips in 2027–2030. Supply is already constrained years in advance.
Key Companies
ASMLASML Holding — sole global maker of EUV lithography machines. Absolute monopoly.Monopoly
NikonNikon (7731.T) — older DUV lithography only (mature nodes)
CanonCanon (CAJ) — older DUV + nano-imprint for mature nodes
07
Semiconductor Capital Equipment
The Toolbox of Chip Manufacturing — 1,000+ Process Steps
What Happens
Beyond ASML's lithography, manufacturing a chip requires dozens of other specialised machines. Deposition equipment (Applied Materials, TEL) deposits ultra-thin films of materials atom by atom onto the wafer. Etching equipment (Lam Research) removes unwanted material with extreme precision — like carving with atomic-scale chisels. Ion implantation fires atoms into the silicon to change its electrical properties. Inspection and metrology machines (KLA) examine every wafer at microscopic scale to detect defects — because a single contamination particle can destroy an entire chip. A single 300mm wafer passes through 1,000+ individual process steps across 50+ different machines, taking 3–4 months from start to finished die.
Applied Materials, Lam Research, and KLA together capture roughly 45% of the global semiconductor equipment market. Their revenue leads foundry revenue by 12–18 months, making them excellent leading indicators of the semiconductor cycle.
Key Companies
AMATApplied Materials — deposition, materials engineering, #1 equipment makerDominant
LRCXLam Research — etching and deposition, critical for 3D NANDDominant
KLACKLA Corporation — process control, inspection, defect detectionDominant
8035.TTokyo Electron (TEL) — cleaning, thermal, deposition equipmentMajor
ACLSAxcelis Technologies — ion implantation specialistNotable
08
Wafer Fabrication — The Foundry
Where Blueprints Become Silicon — TSMC's Global Chokepoint
What Happens
A foundry takes a chip designer's blueprint and physically manufactures it using all the equipment described above. The foundry doesn't own the chip design — it's a pure manufacturing service. TSMC takes NVIDIA's H100 blueprint and uses ASML's EUV machines and AMAT's deposition tools to print and build billions of transistors on a wafer. A modern chip fab (fabrication plant) costs $15–25 billion to build, requires several years to construct, and must operate in cleanrooms 1,000x cleaner than a hospital operating theatre — a single particle of dust can destroy a chip. TSMC manufactures 90%+ of the world's most advanced chips (below 5nm), including every NVIDIA AI GPU, every Apple chip, and AMD's most advanced processors. This geographic concentration in Taiwan creates enormous geopolitical risk for the global technology industry.
TSMC's 2nm process (N2) entering production in 2025 packs 292 million transistors per square millimetre. For context, a human hair is 75,000 nanometres wide; transistors are now 2 nanometres.
Key Companies
TSMTSMC — 90%+ of advanced chips. Makes NVDA, Apple, AMD, ARM-based designsDominant
SamsungSamsung Foundry — #2, strong in mobile chips, struggles at leading edgeMajor
INTCIntel Foundry — 18A process, targeting external customers from 2025–2026Major
GFSGlobalFoundries — mature nodes (12nm+), strategic supplier for defence/autoNotable
SMICSMIC (981.HK) — China's largest foundry, constrained by US export controls
09
Memory Manufacturing & HBM
The AI Brain's Bandwidth — High Bandwidth Memory Changes Everything
What Happens
AI models are enormous — GPT-4 has ~1.8 trillion parameters, each a number needing to be rapidly accessed during inference. Standard DRAM chips sit on a circuit board connected to the GPU via copper traces — limited bandwidth. High Bandwidth Memory (HBM) stacks 8–12 layers of DRAM vertically using Through-Silicon Vias (microscopic wires passing through each chip), then places this tower directly on the same package as the GPU, connected by thousands of tiny bumps. An NVIDIA Blackwell GPU uses 192GB of HBM3e providing 8 terabytes per second of bandwidth — enough to transfer the entire Library of Congress in about 2 seconds. Without sufficient HBM supply, the most powerful GPU becomes a bottleneck waiting for data. HBM supply has been the binding constraint on AI chip shipments since 2023, with demand consistently exceeding production capacity.
SK Hynix controls approximately 61% of the global HBM market, Samsung around 17%, and Micron the remainder and growing rapidly. Demand exceeds supply through at least 2027.
Key Companies
SKHYSK Hynix — HBM market leader (61%), NVIDIA's primary HBM supplierDominant
SamsungSamsung Electronics — DRAM + NAND + HBM recovering after yield issuesMajor
MUMicron Technology — HBM3e gaining share, DRAM + NAND + $22B pre-soldMajor
KioxiaKioxia — NAND flash leader (formerly Toshiba Memory)Notable
WDCWestern Digital (SanDisk brand) — NAND flash, storage solutionsNotable
10
Advanced Packaging
Bonding Multiple Chips Into One System — CoWoS & 3D Stacking
What Happens
Modern AI chips are too large and defect-prone to manufacture as a single piece — yields would be uneconomically low. Instead, chips are manufactured as separate smaller dies and then bonded together into a unified system. TSMC's CoWoS (Chip on Wafer on Substrate) places multiple GPU compute dies and HBM memory stacks on a silicon interposer — a thin piece of silicon with thousands of micro-bumps — creating a package that behaves as one integrated system. The NVIDIA Blackwell GB200 uses two GPU compute dies plus six HBM3e memory stacks, all bonded via CoWoS. CoWoS packaging capacity has been the #1 bottleneck limiting NVIDIA GPU shipments since Blackwell launch — not GPU die production, but the packaging. TSMC is spending billions expanding CoWoS capacity specifically for this constraint.
Advanced packaging has become as strategically critical as chip fabrication itself — it's where the final system is assembled, and CoWoS demand exceeds supply by 2:1 in 2025–2026.
Key Companies
TSMTSMC — CoWoS, SoIC, InFO packaging. AI bottleneck critical path.Dominant
ASXASE Technology — world's largest OSAT (Outsourced Semiconductor Assembly & Test)Major
AMKRAmkor Technology — advanced packaging, TSMC partnerMajor
JCETJCET Group — China's leading OSAT, growing in advanced packagingNotable
— Infrastructure Layer —
11
Optical Interconnects & Transceivers
Moving Data at the Speed of Light — The AI Bandwidth Backbone
What Happens
Inside an AI data centre, thousands of GPUs must share data constantly — passing gradient updates during training, activations during inference, and model parameters between layers. Copper cables work over centimetres but degrade rapidly over metres. Optical transceivers convert electrical signals into laser light pulses, transmit them through fibre optic cables at the speed of light, then convert back to electrical at the destination. A single 1.6T (terabit per second) optical transceiver can carry the data equivalent of 200 Netflix 4K streams simultaneously. As AI clusters scale to 100,000+ GPUs, the bandwidth between them becomes the limiting factor. The transition from 400G to 800G to 1.6T transceivers is happening now, driven by AI cluster scale. NVIDIA invested $2 billion directly into Coherent because GPU compute is limited by optical bandwidth.
Optical transceiver revenue is expected to double from $8B to $16B+ between 2024 and 2026 driven almost entirely by AI data centre demand.
Key Companies
COHRCoherent Corp — world's largest optical transceiver maker, NVDA invested $2BDominant
LITELumentum — laser components, transceivers, LIDARMajor
MTSIMACOM Technology — photonic integrated circuits, analog ICsMajor
FNFabrinet — contract manufacturer of optical transceiversNotable
CIENCiena — optical networking systems, long-haul fibreNotable
12
Networking Silicon
Traffic Control for 100,000 GPUs — The Invisible Routing Layer
What Happens
Once GPUs are connected with optical fibres, dedicated networking ASICs manage how data flows between them — routing packets, managing congestion, and ensuring that 10,000+ GPUs can communicate simultaneously without bottlenecks. These are not commodity chips: routing AI training traffic requires extremely low latency (microseconds) and zero packet loss, because a single dropped packet during a training step forces the entire cluster to retry. Marvell's Ethernet networking chips and Broadcom's Tomahawk switch ASICs compete for the scale-out AI networking market. NVIDIA's InfiniBand offers an alternative networking fabric that bundles compute and networking more tightly. Separately, Data Processing Units (DPUs) offload networking and security tasks from the GPU, handled by NVIDIA's BlueField and Marvell's LiquidIO.
The choice between InfiniBand (NVIDIA) and Ethernet (MRVL/AVGO) for AI cluster networking is a multi-billion dollar battleground with significant implications for each company's AI data centre revenue.
Key Companies
MRVLMarvell — Ethernet networking ASICs, custom XPUs, networking for AIDominant
AVGOBroadcom — Tomahawk switch ASICs, #1 in data centre Ethernet switchingDominant
NVDANVIDIA — InfiniBand networking, ConnectX NICs, BlueField DPUMajor
ANETArista Networks — data centre network switches and routers, AI-scaleMajor
CSCOCisco Systems — enterprise networking, AI-scale switchingNotable
13
Physical Connectors & Components
The Physical Plumbing — Every Plug, Socket and Cable
What Happens
Every cable plugs into a connector. Every circuit board connects to another via physical connectors. Every power supply, drive, network card, and fibre module requires a precisely engineered physical interface. A single AI server chassis can contain hundreds of individual connectors — power connectors, PCIe connectors, network connectors, storage connectors, and fibre optic connectors. As AI servers push to 1,000W+ per GPU and 120kW per rack, the connectors must handle enormous current without generating excessive heat or resistance. Connector density inside AI servers is rising sharply: more GPUs per server means more connectors per square centimetre of circuit board. The engineering challenge is significant — connectors that worked at 400W don't simply scale to 1,000W.
Key Companies
APHAmphenol — world's 2nd largest connector maker, 1.24 book-to-bill in AIDominant
TELTE Connectivity — world's largest connector maker by revenueDominant
MolexMolex (Koch Industries, private) — major connector and component supplierMajor
Bel FuseBel Fuse (BELFA) — power connectors, magneticsNotable
14
Power Infrastructure
The Electricity Constraint — AI's Biggest Physical Bottleneck
What Happens
A single NVIDIA Blackwell GB200 NVL72 rack draws 120 kilowatts — enough electricity to power 40 homes. A large AI data centre requires 500–1,000MW of electricity — comparable to a small city. Power is now the primary constraint on AI expansion: not chips, not buildings, but electricity. The challenge is threefold: generation (building new gas, nuclear, or renewable plants), transmission (upgrading substations and power lines to handle the load), and distribution inside the data centre (converting high-voltage AC to the precise DC voltages each chip needs). GE Vernova makes the gas turbines powering new data centres. Eaton and Schneider make the power distribution units inside. Vertiv makes the thermal management systems preventing chips from melting. The $1.4 trillion in AI data centre electrification needed by 2030 is the most certain near-term spending commitment in the entire energy sector.
Data centre electricity demand is forecast to more than double by 2030. GE Vernova has $163 billion in backlog driven substantially by AI power demand — nearly 4x its annual revenue.
Key Companies
GEVGE Vernova — gas turbines + grid equipment, $163B backlog, AI power leaderDominant
ETNEaton — power distribution, UPS systems, data centre power managementMajor
SU.PASchneider Electric — energy management software + hardware, AI data centreMajor
VRTVertiv Holdings — data centre thermal management, cooling, UPS systemsMajor
PWRQuanta Services — constructs and upgrades power grid infrastructureNotable
15
Server Assembly & Integration
Putting It All Together — The AI Server Stack
What Happens
Someone must take NVIDIA GPUs, Micron HBM, Marvell networking silicon, Coherent optical transceivers, and Amphenol connectors and integrate them into functioning server systems deployable in data centres. AI server integrators manage supply chains, qualify components, design chassis and cooling systems, handle logistics, and provide enterprise support. The NVIDIA DGX H100 reference design contains 8 H100 GPUs, 80 network ports, thousands of components, and requires liquid cooling — all integrated and tested. Dell's PowerEdge servers are NVIDIA's primary enterprise distribution partner. Super Micro specialises specifically in AI servers with more aggressive thermal designs and faster time-to-market on new GPU platforms. As AI servers hit 1MW+ per rack, the thermal and power engineering at this integration stage becomes as complex as the chip engineering itself.
Dell's AI server revenue grew 757% year-over-year in Q1 FY2027, with $60B in AI server revenue guided for FY2027 — making Dell one of the largest beneficiaries of the AI infrastructure buildout.
Key Companies
DELLDell Technologies — NVIDIA's primary enterprise server partner, $60B AI guidedDominant
SMCISuper Micro Computer — AI server specialist, aggressive GPU integrationMajor
HPEHewlett Packard Enterprise — enterprise AI servers, Cray supercomputersMajor
2317.TWFoxconn (Hon Hai) — server manufacturing + AI server expansionNotable
0992.HKLenovo — global server manufacturer, AI server pushNotable
— Cloud & Network Layer —
16
AI Cloud Infrastructure — Neocloud
Renting Compute by the Hour — The GPU Leasing Market
What Happens
Neocloud companies purchase AI servers from Dell/Supermicro, install them in data centres with sufficient power, connect them with high-speed optical networks, and rent the GPU compute capacity to AI companies, research labs, and enterprises by the hour or under multi-year contracts. Most AI startups cannot afford to spend $500M building their own data centres — they rent. The business model: buy an NVIDIA H100 server for ~$300K, rent it at $2.50–$4.00/GPU-hour, recover the purchase cost within 6–12 months, then profit for the remaining 3–4 year hardware life. The critical differentiator between neoclouds is power and location — securing gigawatts of power access years in advance is the strategic moat. Companies that locked in power agreements in 2022–2023 are now delivering capacity that new entrants cannot replicate for 3–5 years.
The neocloud market is dominated by CoreWeave, Nebius, and IREN, with $2B+ in contract announcements from these companies in 2026 alone as hyperscalers outsource overflow demand.
Key Companies
CRWVCoreWeave — #1 neocloud, $7.5B IPO 2025, NVIDIA-backed, Microsoft customerDominant
NBISNebius Group — $46B committed backlog, Meta + Microsoft customerMajor
IRENIREN — $4B+ ARR target, 85% contracted, Microsoft + NVIDIA customerMajor
LambdaLambda Labs (private) — AI cloud for ML researchersNotable
CrusoeCrusoe Energy (private) — AI cloud from stranded natural gasNotable
17
Hyperscale Cloud Platforms
The Operating System of AI — AWS, Azure, Google Cloud
What Happens
Hyperscalers don't just provide compute — they provide the complete development environment, managed services, databases, APIs, security, and global distribution that enterprises build AI on. Most companies don't interact with GPUs directly — they provision a cloud instance and pay per hour. The hyperscalers abstract all hardware complexity, managing millions of servers, thousands of network switches, and complex cooling systems invisibly. Beyond providing infrastructure, all three major hyperscalers are also building and deploying their own AI models — AWS Titan, Azure OpenAI (GPT-4), and Google Gemini. The $700B+ annual hyperscaler AI capex (2026 estimate) represents the single largest capital deployment in technology history, driven by competitive pressure: no hyperscaler can afford to fall behind on AI infrastructure without ceding cloud market share to rivals.
AWS, Azure, and Google Cloud together represent over $300B in annual cloud revenue growing at 20–60% year-over-year, all increasingly driven by AI workloads.
Key Companies
AMZNAmazon / AWS — #1 cloud (33% share), Trainium/Inferentia custom chipsDominant
MSFTMicrosoft / Azure — #2 cloud, OpenAI investment, Copilot AI integrationDominant
GOOGGoogle / GCP — #3 cloud (+63% Q1 2026), Gemini AI, TPU chipsDominant
ORCLOracle Cloud — fastest growing hyperscaler, NVIDIA preferred partnerMajor
METAMeta — internal AI infrastructure only, 1M+ GPU cluster for LLaMANotable
18
Telecommunications & AI-RAN
Connecting Data Centres to the World — The 5G & AI-RAN Transition
What Happens
Data centres are islands — powerful but isolated. Connecting them to each other globally and to billions of end users requires telecommunications infrastructure: fibre networks, submarine cables, 5G radio systems, and core networking equipment. The critical evolution in AI is AI-RAN (AI Radio Access Network) — replacing traditional dedicated hardware in wireless base stations with general-purpose servers running AI software. Traditional base stations use specialised hardware that can only run network software. AI-RAN runs the same network software on standard servers with GPUs, enabling the network to dynamically optimise signal routing, interference cancellation, and capacity allocation using AI in real-time. NVIDIA invested $1 billion into Nokia specifically for AI-RAN, because AI-RAN uses NVIDIA GPUs as the core processing engine for every wireless base station globally.
AI-RAN could replace 5 million+ base stations globally with GPU-powered alternatives — a multi-hundred billion dollar transition over the next decade that Nokia and Ericsson are positioned to capture.
Key Companies
NOKNokia — #2 telecom equipment globally, AI-RAN pioneer, NVIDIA invested $1BDominant
ERICEricsson — #1 telecom equipment globally, 5G + AI-RAN developmentDominant
ANETArista Networks — data centre networking connecting AI clusters globallyMajor
CIENCiena — optical long-haul networking between data centresMajor
COMMCommScope — wireless infrastructure, small cell networksNotable
— AI Training, Inference & Applications —
19
AI Model Training & Inference
Training Once, Running Billions of Times — Where All Infrastructure Converges
What Happens
Training and inference are fundamentally different workloads that share the same infrastructure but have different economic profiles. Training means showing an AI model trillions of data examples until it learns patterns — GPT-4 training used ~25,000 NVIDIA A100 GPUs running for 90+ days at an estimated $50-100M cost. It happens once per model version. Inference is running that trained model to answer queries in real-time — every ChatGPT response, every AI image, every fraud detection. Inference happens billions of times daily and is where revenue is generated. The hardware requirements differ: training needs maximum throughput (running everything in parallel, latency less critical). Inference needs minimum latency per query (single user waiting for a response). This distinction is why hyperscalers are investing in custom inference chips through MRVL and AVGO — a bespoke chip runs inference for a specific model 3–5x more cheaply than a general NVIDIA GPU, critical at billions of queries per day.
Inference is growing faster than training as AI models deploy at scale — JPMorgan estimates inference will exceed training compute demand by 2027, shifting the competitive dynamic toward custom silicon over general-purpose GPUs.
Key Model Developers
OpenAIOpenAI — GPT-4, o3, ChatGPT. Backed by Microsoft. ~$3B revenue 2024.Dominant
GOOGGoogle DeepMind — Gemini 2.0, TPU-native training, integrated in all productsDominant
METAMeta AI — Llama 3 open source, largest open weight model familyMajor
AnthropicAnthropic — Claude (private, backed by Google $2B + Amazon $4B)Major
xAIxAI — Grok, 100K H100 cluster (Elon Musk, private)Notable
MistralMistral AI — European open-weight models (private, Paris-based)Notable
20
AI Applications & Software
Where AI Creates Real-World Value — The ROI Layer
What Happens
The top of the chain is where AI creates measurable business value — where the billions spent on chips, power, and cloud translate into products and services. Enterprise software companies embed AI into workflows: ServiceNow deploys AI agents that handle IT tickets, HR requests, and financial approvals autonomously. Salesforce's Einstein AI generates sales insights. Vertical AI companies apply AI to specific industries: Harvey for legal, Abridge for healthcare documentation, Suno for music creation. AI-native consumer companies — Perplexity (AI search), Character.ai (conversation), Midjourney (image generation) — build directly on the foundation model layer. This is also where the great ROI question is resolved: are enterprises genuinely more productive and profitable because of AI? The answer propagates back through every stage in the chain. Strong ROI validates continued capex; weak ROI triggers a reconsideration of the $700B+ annual spend.
The enterprise AI software market is projected to grow from $35B in 2024 to $280B by 2030 — a faster growth rate than the infrastructure layers below it, once adoption accelerates.
Key Companies
MSFTMicrosoft Copilot — AI across Office 365, GitHub, Azure. Widest enterprise reach.Dominant
NOWServiceNow — AI agents for enterprise IT/HR/finance workflows, $27.7B RPODominant
PLTRPalantir — AI for government and enterprise decision intelligenceMajor
CRMSalesforce — Einstein AI, Agentforce, AI in CRMMajor
ZETAZeta Global — AI-powered marketing intelligence, Palantir partnershipMajor
IONQIonQ — trapped-ion quantum computing systems for advanced AI and optimisation researchNotable

Join the stack

The next deep dive straight to your inbox.