The Cheap-Model Winners: Who Cashes In When AI Gets Commoditized
Special report, July 18, 2026. Brought to you by SignalDeck.live.
The thesis, in plain English
Something important is happening to the economics of artificial intelligence, and this week the market finally noticed, violently. The price of running an AI model (what the industry calls “inference”, the act of actually using a trained model, as opposed to “training,” which is building it) has been collapsing. A query that cost $30 per million tokens on a GPT-4-class model in March 2023 costs roughly $0.40 today, a ~95% price decline in two years. At the same time, “open-weight” models, models whose underlying parameters are published for anyone to download and run on their own hardware, the way Linux gave away the operating system, have closed most of the quality gap with the closed, proprietary frontier.
When the thing in the middle of a stack becomes cheap and abundant, value doesn’t disappear. It migrates. It moves down the stack, to the infrastructure that gets paid on usage, chips, memory, networking, cloud capacity, databases, monitoring. And it moves up the stack, to distribution, the companies with enormous installed customer bases who can now ship AI features to hundreds of millions of users at a fraction of last year’s cost. The layer that gets squeezed is the middle: the model labs themselves, whose product is being commoditized in real time. Think of what happened when bandwidth got cheap: the telecoms who sold raw connectivity suffered, while the companies built on top of cheap bandwidth (streaming, cloud, e-commerce) and the ones selling the equipment underneath both prospered.
The third group of winners is the FCF-rich incumbents, companies generating so much free cash flow (the cash left over after all expenses and capital spending; the money that funds buybacks, dividends, and war chests) that they can outspend everyone embedding AI into products customers already pay for.
This week put the thesis on trial. On Friday, July 17, 2026, the semiconductor index fell into a bear market, down 20% from its June high, triggered by the release of Kimi K3, a massive Chinese open-weight model that reignited the “cheap models will kill AI capex” fear. The Nasdaq fell 1.4%, the S&P 500 1.02%, and Nvidia briefly lost its most-valuable-company crown to Apple. Here’s the tell that this was sentiment, not fundamentals: TSMC, the company that physically manufactures the world’s AI chips, fell 7% on the same day it reported profit up 77%. When a stock falls 7% on a +77% profit print, the market is repricing fear, not facts.
The crux: Jevons paradox
The bear case says: if models get 50x cheaper, AI spending must collapse. The 19th-century economist William Stanley Jevons would disagree. He observed that when coal-powered engines became more efficient, coal consumption rose, because cheaper power made steam economical for thousands of new uses. That’s the Jevons paradox: efficiency expands demand rather than shrinking it.
The AI data says Jevons is winning, decisively:
Token volumes are exploding. Inference token volume is projected to grow roughly 24x by 2030, to around 120 quadrillion tokens per month. The inference market itself is projected to grow from $106B in 2025 to $255B by 2030 (a 19% annual growth rate).
Inference has taken over the workload. Inference is now roughly two-thirds of all AI compute in 2026, up from one-third in 2023 and half in 2025, and it represents 80-90% of an AI system’s lifetime cost. One “agent” task (an AI performing a multi-step job autonomously) fires off 10-20 separate model calls. Agents are token multipliers.
Prices collapsed and spending rose anyway. Token prices have deflated at a median of ~50x per year (with a range of 9-900x, expected to decelerate to 3-5x/yr through 2027). Yet total AI spend went up.
The key evidence, this fear already failed once. In January 2025, the DeepSeek shock wiped ~$600B off Nvidia in a single day on this exact “cheap model” thesis. What happened to hyperscaler capex afterward? It grew: combined 2025 capex came in around $410-450B, and 2026 is tracking toward ~$725B, up 64-77%, with Amazon at ~$200B, Google $180-190B, Meta $125-145B, and Microsoft $120B+. Goldman pegs cumulative 2025-27 spend at $1.15T.
Cheaper inference doesn’t shrink the pie. It makes a thousand previously-uneconomic AI use cases suddenly profitable, and every one of them consumes compute, storage, networking, and monitoring.
“Everyone runs their own”: open weights and self-hosting
The second leg of the shift: companies increasingly don’t need to rent intelligence from OpenAI or Anthropic at all. Open-weight model quality now sits within a few points of the closed frontier. Kimi K3, released July 17, 2026, at 2.8 trillion parameters, the largest open-weight model ever, with weights publicly available July 27, benchmarks behind only the very top US frontier models. Meta’s Llama 4 Maverick scores 85.5% on MMLU (a broad knowledge benchmark); Qwen3, DeepSeek-V3.2, and Kimi hit ~70% on SWE-bench (a real-world coding benchmark); MiniMax and GLM lead open leaderboards.
Meanwhile, self-hosting, running the model on your own rented or owned GPUs instead of paying per-token API fees, becomes cheaper than the API at roughly 10-30 million tokens per day of usage. The tooling to do it (vLLM for serving, Ollama for local deployment) has standardized. The drivers go beyond cost: privacy, data residency, regulatory sovereignty.
Read this correctly: it’s deflationary for the model labs, their product is the thing being commoditized, but it’s a tailwind for everyone building on top of models and everyone selling the hardware and plumbing underneath them. When intelligence is a free ingredient, the money is made by whoever owns the kitchen, the supply chain, or the restaurant with a line out the door.
The layer map, who wins
Hyperscale & AI clouds, Clear (ORCL debated): rent the compute everyone (including self-hosters) runs on; model-neutral toll collectors. Names: MSFT, AMZN, GOOGL, ORCL.
App-software franchises, Debated (seat-model risk): huge installed bases; cheap inference cuts AI COGS and creates new SKUs overnight. Names: MSFT (365), CRM, ADBE, NOW, INTU.
AI-native / model-agnostic, Clear: built to swap models freely; cheaper input = fatter margin. Names: PLTR, APP, DUOL, SHOP.
Data / dev / observability, Clear (GTLB debated): every AI app needs a database, a pipeline, and monitoring; consumption-priced. Names: SNOW, MDB, DDOG, GTLB.
Edge, on-device & self-host, Clear: small cheap models are exactly what runs on phones, edge networks, and SMB clouds. Names: NET, AAPL, QCOM, DOCN.
Inference silicon / memory / networking, Clear direction, debated shares: usage explosion needs chips, HBM memory, and Ethernet regardless of which model wins. Names: NVDA, AVGO, AMD, MRVL, MU, ANET.
Special case: Meta spans two layers at once, the largest distribution footprint on earth (3.56B daily users) and the open-weights arsenal (Llama) actively commoditizing the model layer for everyone else.
The scoreboard
All figures from company reports as compiled in our research digest; quarters as labeled.
MSFT, Hyperscale + apps · rev +18% (Q3 FY26, $82.9B) · FCF $15.8B/qtr (-22% YoY, capex-squeezed) · 1.6B Windows devices, 400M+ M365 seats · Strong tailwind.
GOOGL, Hyperscale + distribution · rev +22% (Q1 26, $109.9B) · FCF $10.1B/qtr, $64.4B TTM (margin fell 21%→9.2%) · ~$460B cloud backlog · Strong.
AMZN, Hyperscale, most model-neutral · rev +17% (Q1 26, $181.5B), AWS +28% · FCF $1.2B TTM, collapsed ~95% · AWS #1 at ~29-31% share · Strong.
META, Distribution + open weights · rev +33% (Q1 26, $56.31B) · FCF $12.4B/qtr · 3.56B daily users, Llama 1B+ downloads · Strong.
AAPL, Edge / on-device · rev +17% (Q2 FY26, $111.2B) · OCF $28.7B/qtr, net cash $62B · 2.5B+ active devices · Strong.
NVDA, Inference/training GPU · rev +85% (Q1 FY27, $81.6B) · FCF $48.6B/qtr · ~80% accelerator share · Two-sided.
AVGO, Custom ASIC + networking · rev +48% (Q2 FY26, $22.2B), AI semi +143% · FCF $10.26B/qtr (46% margin) · ~$73B AI backlog, 6 XPU customers · Strong.
AMD, 2nd-source GPU · rev +38% (Q1 26, $10.3B) · FCF $2.6B/qtr (record 25%) · Meta 6GW MI450 deal · Two-sided.
MRVL, Custom ASIC #2 · rev +28% (Q1 FY27, $2.418B) · FY28 revenue target $16.5B · #2 custom-ASIC house · Two-sided.
MU, HBM memory · Q3 FY26 rev $41.46B (supercycle) · adj FCF $18.3B (record) · HBM sold out for 2026 · Strong.
ANET, AI networking · rev +35% (Q1 26, $2.709B) · 47.8% non-GAAP op margin · #1 in >10GbE switching · Strong.
QCOM, On-device silicon · rev -2% (Q2 FY26, $10.6B) · auto >$5B run-rate, 1M+ vehicles · Neutral.
CRM, App franchise · rev +13% (Q1 FY27, $11.1B) · FY26 OCF $15B · 150K+ customers, Agentforce ARR $1.2B +205% · Two-sided.
ADBE, App franchise · rev +12.7% (Q2 FY26, $6.62B) · FCF ~$2.11B/qtr, ~$10.3B TTM · 32-33M Creative Cloud subs · Two-sided.
NOW, App franchise · subscription +22% (Q1 26, $3.77B) · FCF $1.67B/qtr (44% margin) · 8,800 customers, 85% of Fortune 500 · Two-sided.
INTU, App franchise · rev +10.4% (Q3 FY26, $8.56B) · FY26 FCF ~$7.4B · ~100M customers · Strong.
SAP, System of record · cloud +27% cc (Q1 26) · FY26 FCF ~€10B · cloud backlog €21.9B +25% · Two-sided.
WDAY, App franchise · FY26 sub rev guide +14% ($8.815B) · FCF ~$2.65B · 11,000+ orgs, 60%+ Fortune 500 · Two-sided.
PLTR, Ontology layer · rev +85% (Q1 26, $1.6B) · adj FCF $925M (57% margin) · 1,007 customers, US Army up-to-$10B deal · Strong.
APP, AI ad-tech · rev +59% (Q1 26, $1.84B) · FCF $1.29B/qtr · 5,306 e-commerce advertisers · Strong.
DUOL, AI-native consumer · rev +27% (Q1 26, $292M) · FCF $147.8M (50.6% margin) · 137.8M MAU, 12.5M paid · Strong.
SHOP, Commerce rails · rev +34% (Q1 26, $3.17B) · FCF $476M (15% margin) · GMV >$100B, ~5.5-6.9M merchants · Strong.
SNOW, Data cloud · rev +33% (Q1 FY27, $1.39B) · adj FCF $265.5M (19% margin) · 13,912 customers, ~50% use AI weekly · Strong.
MDB, AI-app database · rev +25% (Q1 FY27, $688M) · FCF $197.5M/qtr · 67,700 customers, 19 of top-20 US banks · Strong.
DDOG, Observability · rev +32% (Q1 26, $1.006B) · FCF $289M/qtr (29% margin) · 32,700 customers, ARR >$4B · Strong.
GTLB, AI dev platform · rev +23% (Q1 FY27, $264M) · adj FCF $147M (timing-boosted) · 10,682 base customers, >50% Fortune 100 · Two-sided.
ORCL, Neutral AI cloud · rev +21% (Q4 FY26, $19.2B), OCI +93% · FCF NEGATIVE $23.7B FY26 · RPO $638B (+363%) · Two-sided.
DOCN, SMB self-host cloud · rev +22% (Q1 26, $258M) · adj FCF margin 9-12% · 650K+ customers, AI ARR $170M +221% · Strong.
NET, Edge inference · rev +34% (Q1 26, $639.8M) · FCF $84.1M/qtr (13% margin) · 330+ cities, within 50ms of 95% of humanity · Strong.
The companies
Hyperscale & AI clouds
MSFT, Microsoft
What it does: The world’s largest enterprise software company plus the #2 cloud (Azure, ~24-25% share). Footprint: The single biggest distribution machine in enterprise computing, 1.6 billion Windows devices, 400M+ Microsoft 365 seats, over 90% of the Fortune 500, and 20M paid Copilot seats already. Earnings power: Q3 FY26 revenue of $82.9B grew 18% with a 46.3% operating margin; Azure grew 40% and the AI business hit a $37B annualized run-rate, up 123%. The blemish: quarterly FCF fell 22% to $15.8B because capex nearly doubled to $30.9B. How it wins when models cheapen: every 50x drop in inference cost drops the cost of serving Copilot to the largest paying enterprise base on earth, cheap models are a direct COGS cut on a product Microsoft has already sold. And Azure is model-agnostic: it hosts OpenAI and rivals, so it gets paid whichever model wins. Innovation angle: distribution-first AI, Copilot embedded in every surface a knowledge worker already touches. Verdict: the best-positioned “cheap AI + wide distribution” name in the market. Key risk: FCF keeps shrinking under capex intensity; the market’s patience for that is not unlimited.
GOOGL, Alphabet
What it does: Search, Android, YouTube, and the #3 cloud, plus the most vertically integrated AI stack of any hyperscaler. Footprint: billions of daily users across Search/Android/YouTube, and a cloud backlog of roughly $460B that nearly doubled quarter-over-quarter. Earnings power: Q1 2026 revenue $109.9B, up 22%; Google Cloud revenue of $20.03B grew 63% with operating income tripling to $6.6B from $2.2B. Net income of $62.6B looks spectacular but was inflated by investment gains, treat it with a discount. FCF was $10.1B for the quarter ($64.4B trailing), with FCF margin cratering from 21% to 9.2% on $35.7B quarterly capex ($180-190B planned for the year). Cash on hand: $126.8B. How it wins when models cheapen: this is the crucial part, Google designs its own AI chips (the TPU line, now “Ironwood”), so its cost of running AI falls faster than rivals renting GPUs. That lets it serve Gemini answers inside Search at near-zero marginal cost. Innovation angle: the TPU/Ironwood vertical stack is the most differentiated infrastructure asset among the hyperscalers. Verdict: the best combination of distribution and infrastructure control in the group. Key risk: the FCF margin collapse is real, and AI answers in Search may cannibalize the ad clicks that fund everything.
AMZN, Amazon
What it does: The #1 cloud (AWS, ~29-31% share) plus the retail empire. Footprint: AWS is the default enterprise cloud, and it just printed its fastest growth in 15 quarters, $37.6B, up 28%, at a 37.7% operating margin. Consolidated revenue was $181.5B (+17%) with a record 13.1% operating margin. Earnings power, read the footnotes: net income of $30.3B was inflated by a $16.8B gain on its Anthropic stake, and trailing FCF has collapsed roughly 95% to $1.2B under $44.2B of quarterly capex. That is the most honest number in this report: Amazon is spending essentially every dollar it generates. How it wins when models cheapen: AWS is the most deliberately model-neutral of the big clouds, its Bedrock service serves whichever model the customer wants, and its in-house Trainium chips (now a >$20B run-rate, carrying over 50% of Bedrock’s tokens, with OpenAI committing 2GW and Anthropic ~1GW) make Amazon cheaper to run any of them on. It wins regardless of which model wins, the clearest Jevons validation in the group. Innovation angle: neutrality plus proprietary silicon. Verdict: the best structural AI-infrastructure position. Key risk: FCF has effectively vanished, and the headline net income isn’t clean. You’re trusting management that the capex converts.
ORCL, Oracle
What it does: The database giant reborn as a wholesale GPU landlord for the AI era. Footprint & growth: Q4 FY26 revenue $19.2B (+21%); OCI, its cloud infrastructure arm, grew 93% to $5.8B; total cloud $9.9B (+47%). The staggering number: remaining performance obligations (contracted future revenue) of $638B, up 363%, with OpenAI representing roughly half. FY27 revenue is guided to $90B. Earnings power, the hard truth: FY26 FCF was negative $23.7B ($32B operating cash flow minus $55.7B capex), FY27 net capex is guided to ~$70B so FCF stays negative, funded by ~$43B of new debt and counting. S&P downgraded Oracle to BBB- on July 9, 2026, explicitly citing OpenAI concentration. How it wins when models cheapen: Oracle rents raw multi-vendor GPU capacity (Superclusters with >100K Nvidia Blackwells plus 50K AMD MI450s), the literal landlord for organizations that want to run their own models. Innovation angle: vendor-agnostic superclusters at scale. Verdict: the highest-torque hyperscale bet, and the most leveraged to one customer’s solvency. Key risk: half the backlog is OpenAI, and the whole build is debt-financed with negative FCF. This is a credit story wearing a cloud costume.
Distribution & open weights
META, Meta Platforms
What it does: The largest social distribution network in history, now weaponizing open-source AI. Footprint: 3.56 billion daily users. Llama has 1B+ downloads, the most-distributed open-weight model family on earth. Earnings power: Q1 2026 revenue of $56.31B grew 33%, the fastest of the four mega-caps, at a 41% operating margin (48.1% in the core apps). Net income of $26.8B included an $8B one-time tax benefit, so haircut it. FCF was $12.4B despite $19.8B of quarterly capex, with FY26 capex guided to $125-145B, the steepest hike of any hyperscaler. How it wins when models cheapen: Meta wants models to be free. By giving Llama’s weights away, it commoditizes the layer where rivals (OpenAI, Google) make money, while protecting the layer where it makes money, advertising. Cheaper AI directly feeds the ad engine: impressions +19%, price per ad +12%. Every dollar the industry shaves off inference is a dollar off Meta’s cost of running AI across 3.56B daily users. Innovation angle: open-weights-as-strategic-weapon is the single most elegant competitive move in this entire report. Verdict: the cleanest expression of “cheap and open AI as a distribution weapon”, and unlike most AI stories, it’s monetized in the ad P&L today. Key risk: hyperscaler-scale capex with no cloud revenue line to offset it, and that flattered net income.
Edge & on-device
AAPL, Apple
What it does: The premium device franchise, and, quietly, the purest edge-inference play in mega-cap tech. Footprint: 2.5B+ active devices, every one shipping with a Neural Engine (Apple’s on-device AI chip, standard since 2017). Earnings power: Q2 FY26 revenue $111.2B (+17%), Services at a record $31B (+16%), net income $29.6B (+22%), 49.3% gross margin, $28.7B operating cash flow, $62B net cash. No hedges needed, these are clean numbers. How it wins when models cheapen: the entire direction of cheap AI, distillation (compressing big models into small ones) and quantization (shrinking them to run on less memory), produces exactly the kind of model a phone’s NPU runs locally. Apple spends nearly nothing on training and captures the benefit of everyone else’s efficiency race. It is the only mega-cap with essentially zero training-capex exposure. Innovation angle: a 3B-parameter on-device model plus Private Cloud Compute, a privacy-preserving tier for heavier queries, a genuinely differentiated privacy architecture. Verdict: the purest “distribution + FCF” edge bet, with near-risk-free AI optionality, Friday’s flight to Apple (it briefly reclaimed most-valuable-company from Nvidia) shows the market pricing exactly this. Key risk: it’s still a hardware-cycle and China story; if cheap on-device AI never triggers an upgrade supercycle, the AI angle stays small.
NET, Cloudflare
What it does: A global edge network, servers in 330+ cities, within 50 milliseconds of 95% of the world’s population, that is becoming the toll road for AI traffic. Footprint & growth: Q1 2026 revenue $639.8M (+34%); 4,416 customers paying $100K+ (up 25%); $1M+ deals up 73%. Earnings power: still early, 11.4% operating margin, $84.1M FCF (13% margin). This is a growth asset, not a cash cow yet. How it wins when models cheapen: Workers AI runs inference at the edge and charges by consumption (“neurons”), so revenue scales with usage volume, not with which model is fashionable. AI Gateway sits in front of any model as routing/caching plumbing. And the boldest move: Pay Per Crawl and the Monetization Gateway let websites charge AI crawlers via HTTP 402, with AI crawlers default-blocked on ad-supported pages from September 15, 2026. Cloudflare is trying to become the payments layer of the agent economy. It also cut ~20% of its workforce to reorganize “agentic-AI-first.” Innovation angle: AI Gateway plus the Monetization Gateway is real agent-economy plumbing nobody else has shipped. Verdict: the neutral volume-toll on inference, wins as usage multiplies, indifferent to the model wars. Key risk: the richest valuation in this report at ~28-35x sales. The thesis can be right and the entry price still punishing.
Inference silicon, memory, networking
NVDA, Nvidia
What it does: The ~80% share leader in AI accelerators. Footprint & earnings power: Q1 FY27 revenue of $81.6B grew 85%, Data Center $75.2B (+92%), gross margin 75%, net income $58.3B, and FCF of $48.6B in a single quarter, plus an $80B buyback authorization and a dividend raised from $0.01 to $0.25. Hyperscalers are ~50% of data-center revenue. How the cheap-model era cuts both ways: the bear case is that model efficiency reduces training capex and that custom ASICs (Broadcom/Marvell chips built for specific customers) take inference share. The bull case is Jevons, aggregate inference demand explodes, and Nvidia is racing itself: the Rubin generation targets 10x lower inference cost per token, which is Nvidia choosing to be the deflation rather than its victim. Innovation angle: no company has ever self-deflated its own product this aggressively while growing 85%. Verdict: the largest, most liquid way to own inference-demand growth, with a quarterly FCF engine that funds the entire arms race. But it is genuinely two-sided in a way the toll-takers are not. Key risk: priced for 80%+ growth; one hyperscaler digestion quarter or visible ASIC substitution hits the multiple disproportionately hard. Friday’s bear-market plunge in semis is this risk expressing itself.
AVGO, Broadcom
What it does: The design partner behind custom AI chips (“XPUs”) for the biggest AI spenders, plus the networking silicon connecting them. Footprint: six custom-XPU customers, including Google, Meta, Anthropic, and reportedly a ~$10B OpenAI program; an AI backlog around $73B. Earnings power: Q2 FY26 revenue $22.2B (+48%), AI semiconductors $10.8B (+143%), FCF $10.26B, a 46% FCF margin that rivals or beats Nvidia’s, and non-GAAP net income of $12.1B. Guidance is staggering: Q3 AI semis of $16B (+200%), ~$56B for FY26, and a $100B+ AI target for 2027. How it wins when models cheapen: custom ASICs are the mechanism by which cheap models get served cheaply, when a hyperscaler wants inference at lowest cost per token, it commissions a Broadcom-designed chip. The commoditization of models literally routes through Broadcom’s order book. Innovation angle: the merchant custom-silicon model itself. Verdict: the second-most-direct inference-silicon play, with best-in-class cash economics. Key risk: extreme customer concentration, a handful of design wins are the whole story, and the VMware software business is soft.
Tight profiles, silicon and networking
AMD, The credible second source. Q1 2026 revenue $10.3B (+38%), Data Center $5.8B (+57%), record FCF of $2.6B (25% margin, 3x YoY), and a landmark 6GW MI450 deal with Meta. At only ~5-7% accelerator share, AMD wins as buyers deliberately diversify off Nvidia for inference, but it remains a distant #2 whose story depends on MI450 execution.
MRVL, The #2 custom-ASIC house behind Broadcom. Q1 FY27 revenue $2.418B (+28%), data center $1.83B, custom silicon set to more than double by FY28 toward a $16.5B revenue target. Same thesis as Broadcom, earlier stage, with heavier program-concentration risk, a single lost socket matters enormously.
MU, The quiet monopoly-adjacent winner: memory. Q3 FY26 revenue of $41.46B in a full supercycle, ~81% gross margin, adjusted EPS $25.11, record adjusted FCF of $18.3B. HBM (high-bandwidth memory, the specialized RAM every AI accelerator needs) is sold out for 2026, with HBM4 for Nvidia’s Rubin ramping at ~2x the prior generation. The bottleneck has moved to memory and packaging, Micron gets paid regardless of which chip or model wins.
ANET, Networking is the other agnostic toll. Q1 2026 revenue $2.709B (+35%) at a 47.8% non-GAAP operating margin; #1 in high-speed (>10GbE) switching; FY26 AI revenue target of $3.5B within ~$11.5B total. As hyperscalers convert from InfiniBand to Ethernet and inference nodes multiply, every new rack needs Arista switching, model-indifferent by construction.
QCOM, The on-device idea without Apple’s franchise. Q2 FY26 revenue of $10.6B actually fell 2% on handset softness, though automotive passed a $5B run-rate (1M+ vehicles on Snapdragon Ride) and IoT grew 9%. Its NPUs benefit from small-model proliferation across phones, cars, and PCs, but it’s diversifying off a declining core, a structurally weaker seat at the edge table than Apple’s.
App-software franchises
CRM, Salesforce
What it does: The system of record for the world’s sales and service organizations, 150K+ customers, 90% of the Fortune 500. Earnings power: Q1 FY27 revenue $11.1B (+13%) at a 34.8% non-GAAP operating margin; current remaining performance obligation $33.6B (+14%); FY26 operating cash flow of $15B; a $50B buyback with $27.5B returned in Q1 alone. The AI proof point: Agentforce, its AI-agent product, hit $1.2B ARR, up 205%; total AI ARR is $2.9B, up 200%; the platform processed 28.6 trillion AI tokens, up 152% quarter-over-quarter. How it wins when models cheapen: Agentforce is consumption-priced, so cheaper inference is direct gross-margin expansion on a product it cross-sells into an installed base (60%+ cross-sell already). Verdict: the strongest evidence that a seat-based incumbent can build a real consumption AI business. Key risk, and it’s the big one: Salesforce is the poster child of the seat-compression bear case, down 38% YTD 2026. If agents shrink headcount at its customers, the core seat business deflates faster than Agentforce grows. Genuinely two-sided.
NOW, ServiceNow
What it does: The enterprise workflow platform, the plumbing that routes IT, HR, and customer tickets across 8,800 customers, including 85% of the Fortune 500 and 630 customers paying over $5M a year. Earnings power: Q1 2026 revenue $3.77B with subscriptions up 22%, a 32% non-GAAP operating margin, and $1.67B of quarterly FCF at an exceptional 44% margin. The AI proof point, the best in enterprise software: Now Assist entered 2026 at $750M in contract value, with the target raised from $1B to $1.5B; customers who renew with Now Assist expand their contracts roughly 3x; the >$1M AI cohort grew 130% YoY. How it wins: cheap inference makes Now Assist cheaper to serve, and ServiceNow has the clearest demonstrated evidence that AI is expansion revenue, not cannibalization. It was a named beneficiary in this week’s software-over-semis rotation. Verdict: the strongest AI-monetization proof among the franchises. Key risk: the same agent that resolves a ticket can bypass the ticketing seat entirely; the stock was hit in the SaaSpocalypse, and AI contract value is still small against a ~$14.7B run-rate.
ADBE, INTU, SAP, WDAY, tight profiles
ADBE, The creative-software monopoly at a crossroads. Q2 FY26 revenue $6.62B (+12.7%, 97% subscription), TTM FCF ~$10.3B, a $25B buyback, and 32-33M Creative Cloud subscribers. AI-First ARR passed $500M (3x YoY) with Firefly at ~$300M, and its indemnified, IP-safe model is a genuine enterprise moat. The two-sided problem: model-native tools (Midjourney-class) attack the core from below, and AI ARR is still under 2% of revenue, the moat is being tested faster than the AI line is growing.
INTU, Maybe the most underrated franchise here. Q3 FY26 revenue $8.56B (+10.4%), non-GAAP operating income $4.7B, FY26 FCF ~$7.4B with margin expanding 32.3%→35.1%, and ~100M customers across TurboTax/QuickBooks/Credit Karma. QBO Advanced plus Enterprise Suite grew 38%, it partnered with Anthropic for Enterprise Suite, and it cut 17% of staff in an explicitly AI-driven restructuring, eating its own cooking. The compliance-and-liability moat (someone must stand behind a tax filing) is precisely the kind of moat cheap models can’t dissolve; the long-run risk is AI doing the bookkeeping directly.
SAP, The system-of-record defense. Q1 2026 cloud revenue +27% constant-currency, cloud backlog €21.9B (+25%), FY26 FCF ~€10B, with Joule agents scaling to 50 assistants/200 agents by Q3. The bull argument for SAP over horizontal SaaS: agents still have to write to SAP’s ledger, you can compress seats, but you can’t bypass the ERP. Less seat-exposed than CRM/WDAY, but still carries the SaaS-era pricing model into an agentic world.
WDAY, The most seat-exposed franchise in this report. FY26 subscription revenue guided to $8.815B (+14%) at a 29% non-GAAP operating margin with ~$2.65B FCF, serving 11,000+ organizations including 60%+ of the Fortune 500. Its Illuminate agents are gaining real traction, AI net-new contract value more than doubled and >75% of new deals include AI, but HR and finance are exactly the departments agents shrink, and the SaaSpocalypse hit it accordingly.
AI-native / model-agnostic
PLTR, Palantir
What it does: Palantir sells the ontology layer, a live software model of an organization’s people, assets, and processes that AI models plug into. Crucially, it does not sell a model. Footprint: 1,007 customers, anchored by an up-to-$10B US Army consolidation deal and a £240M UK MoD contract, with US commercial revenue up 133%. Earnings power, remarkable at this scale: Q1 2026 revenue $1.6B, up 85%; GAAP net income $871M (a 53% GAAP net margin); adjusted FCF $925M (57% margin); a “Rule of 40” score (growth rate + margin) of 145, essentially unheard of. How it wins when models cheapen: Palantir’s platform hot-swaps any model, OpenAI, Anthropic, open-weight, under the customer’s workflows. Its moat is orthogonal to model quality, so model commoditization makes its input cheaper while its differentiation is untouched. This is the clearest single embodiment of the entire report’s thesis. Innovation angle: the ontology abstraction is genuinely novel, nobody else has productized “the org chart and physical world as an API.” Verdict: the purest cheap-model beneficiary in software. Key risk: the valuation already knows, a forward P/E of 75-90x (~450% premium to peers) plus government concentration. Thesis-right, price-demanding.
APP, DUOL, SHOP, tight profiles
APP, An AI ad engine printing cash. Q1 2026 revenue $1.84B (+59%), FCF $1.29B, adjusted EBITDA $1.56B at an 84-85% margin, $2.76B cash, $1B buyback. Its Axon ML engine monetizes a proprietary conversion-data moat, and cheaper compute means more ad-model experiments per dollar; e-commerce self-serve opened globally in June 2026 with 5,306 advertisers aboard. Two flags: an open SEC investigation (reported May 2026) and an e-commerce pivot less than a year old.
DUOL, The cleanest direct demonstration that cheap AI drops straight into gross margin. Q1 2026 revenue $292M (+27%), gross margin up 190bps to 73% explicitly from lower per-unit AI costs, FCF $147.8M (50.6% margin), 137.8M MAU / 56.5M DAU (+21%) / 12.5M paid subscribers. AI now generates essentially all content, DuoRadio scaled from 2 to 25+ courses at a 99% cost reduction. Watch items: bookings (+14%) lag revenue (+27%), and it leans on OpenAI’s models rather than a swappable stack.
SHOP, The merchant-side rails for agent-driven commerce. Q1 2026 revenue $3.17B (+34%), GMV >$100B (+35%, second straight quarter above the mark), FCF $476M (15% margin), ~5.5-6.9M merchants in 175+ countries. AI-driven traffic to Shopify stores is up 8x YoY and AI-search orders up 13x, and its agent-agnostic Universal Commerce Protocol (with Google) aims to be the standard a shopping agent transacts through. Thin FCF and unproven agent-commerce scale are the honest caveats, plus the risk hyperscalers build competing rails.
Data, dev & observability
SNOW, Snowflake
What it does: The independent data cloud, where enterprises store and query the data every AI application feeds on. Footprint: 13,912 customers, roughly half now using its AI features weekly; net revenue retention of 126% (existing customers spend 26% more each year); RPO $9.21B, up 38%. Earnings power: Q1 FY27 revenue $1.39B, up 33% and re-accelerating, rare at this scale, with product revenue +34%; but margins remain thin: 12% non-GAAP operating margin, adjusted FCF $265.5M (19%), with ~$2.95B in cash. Management called Cortex AI the “largest driver” of its raised forecast. How it wins when models cheapen: Snowflake is fully consumption-priced, no seats, no licenses, just usage. Every new AI app someone builds means more queries and more storage on the platform. That makes it one of the most purely Jevons-levered assets in software: cheap models mean more apps, and more apps mean more Snowflake consumption, with no seat ceiling to compress. It was a named beneficiary in this week’s software rotation. Verdict: the flagship consumption-model AI winner. Key risk: thin margins, and a two-front war against Databricks and the hyperscalers’ native data stacks.
DDOG, Datadog
What it does: Observability, the monitoring layer that tells engineering teams whether their software (now including their AI agents) is working, fast, and not on fire. Footprint: 32,700 customers, 4,550 paying over $100K, ARR above $4B, net retention in the low-120s. Earnings power: Q1 2026 revenue crossed $1B for the first time ($1.006B, +32%), with a 22% non-GAAP operating margin, $289M FCF (29% margin), and $4.8B in cash, growth and cash generation, the rare combination in this section. How it wins when models cheapen: this is the cleanest pure toll-taker on AI sprawl. Datadog doesn’t care which model you use, whether you rent it or self-host it, or whether it’s open or closed, every AI app and agent put into production needs monitoring, and Datadog’s LLM and Agent Observability products charge by usage. More cheap models → more deployed AI → more telemetry → more Datadog. Verdict: the most model-agnostic, usage-priced business in the entire report. Key risk: hyperscaler-native monitoring and point solutions (LangSmith, Arize) nibbling from below.
MDB, GTLB, tight profiles
MDB, The database AI-native apps default to. Q1 FY27 revenue $688M (+25%), Atlas (its cloud product) +29%, 18% non-GAAP operating margin, FCF $197.5M, RPO $1.46B up a striking 88%, $2.4B cash, 67,700 customers including 19 of the top 20 US banks. Native vector search (the storage format AI memory/retrieval uses) makes it a system-of-record candidate for AI apps, consumption-levered through Atlas. Risk: free and hyperscaler-native vector alternatives (pgvector, Pinecone) compress the differentiation.
GTLB, Genuinely two-sided. Q1 FY27 revenue $264M (+23%), 10,682 base customers, >50% of the Fortune 100, adjusted FCF of $147M (flattered by collection timing), but net retention decelerating 122→117 over four quarters. Tailwind: an explosion of AI-written code still needs testing, security scanning, and deployment. Threat: GitHub Copilot’s default distribution plus Cursor-class tools, a seat-priced model in a seat-compressing era, and a Duo Agent Platform that management itself says brings “no material FY27 revenue.”
Self-host on-ramp
DOCN, DigitalOcean
What it does: The simple, cheap cloud for small businesses and individual developers, and now the literal on-ramp for “everyone runs their own model.” Footprint: 650K+ customers and 6M+ developers, exactly the long tail that hyperscaler pricing and complexity exclude. Earnings power & growth: Q1 2026 revenue $258M (+22%), FY26 guided to $1.13-1.145B (25-27% growth), adjusted FCF margin 9-12%. The acceleration is the story: AI ARR hit $170M, up 221%; Q2 preliminary growth ~29% (from 14% a year ago); 2027 growth guidance raised to 50%+; RPO guided above $800M, up 10x YoY. How it wins when models cheapen: its Gradient AI Agentic Cloud, GPU droplets, and 1-click open-model deployment (on AMD MI350X/MI355X hardware) let a 20-person company download an open-weight model and run it for the price of a server, the purest small-cap expression of this report’s title. Verdict: the highest-leverage pure-play on democratized self-hosting. Key risk: scale, 31MW of new capacity is a rounding error against hyperscalers; this wins the long tail or nothing.
Short callout, the high-torque AI-cloud pure-plays
CRWV (CoreWeave) and NBIS (Nebius) are the leveraged ends of the GPU-landlord trade. CoreWeave: Q1 2026 revenue $2.08B (+112%), but a $740M net loss, $31-35B of debt-funded FY26 capex, a $99.4B backlog, and Microsoft at ~62-71% of revenue. Nebius: revenue $399M (+684%), $1.92B ARR, positive adjusted EBITDA ($129.5M), 3.5GW of power, and $17.4B Microsoft plus $27B Meta contracts with investment-grade counterparties, but $20-25B capex running ahead of revenue. Highest torque to the Jevons thesis in this report; highest fragility if it wobbles. Also note: Confluent (CFLT) was acquired by IBM for ~$11.6B (closed March 17, 2026) and is no longer a standalone tradeable equity.
Who does NOT win
Honesty section. Four groups lose or carry real two-sided risk:
The model labs themselves. This is the commoditized middle. Chinese models undercut Western API pricing by up to 9x; OpenAI and Anthropic are cutting prices in response; the blunt industry phrase in our digest is “margins disappear.” When your product’s open-source substitute is a few benchmark points behind and falling in price 50x a year, you are the coal, not the railroad. (None are directly tradeable, but their economics leak into everyone contracted to them, see Oracle’s OpenAI-heavy backlog and its BBB- downgrade.)
Pure training-capex plays, IF efficiency outruns demand. Jevons has won every test so far, post-DeepSeek capex rose to ~$725B tracking for 2026, but the July 17 semis bear market shows the market re-testing it. If model efficiency ever compounds faster than usage growth, the debt-funded GPU landlords (ORCL negative $23.7B FCF, CRWV $740M quarterly loss) get hurt first and worst, and NVDA’s 80%+ growth expectation deflates.
Seat-based SaaS exposed to agent-driven seat compression. The “SaaSpocalypse” already wiped ~$285B of SaaS market cap in February 2026, and Gartner sees seat-based share of software spend falling from 21% to 15% by 2030, $234B at risk. Flagging our own “distribution winners” honestly: CRM (down 38% YTD, the market’s chosen victim), NOW (agents can bypass tickets), and WDAY (HR/finance is where agents shrink headcount first) all carry this two-sided exposure, as do GTLB and to a lesser degree SAP. Their offset is real consumption-AI revenue (Agentforce $1.2B ARR +205%; Now Assist toward $1.5B), but the race between the new line growing and the old line compressing is genuinely unresolved. The bull counter deserves stating, because it’s fair: history says automation grows total employment rather than shrinking it (AI-application-engineer postings are up ~143% YoY even amid 2026 tech layoffs, and enterprise net revenue retention still holds near 118%). But total headcount is the wrong axis for this risk. What actually bites is narrower: seats licensed for the specific tool (an agent that closes a ticket end-to-end removes a ServiceNow fulfiller seat whether or not the company hires AI engineers elsewhere), and the per-seat pricing model itself, which Salesforce, ServiceNow, and Workday are already walking away from toward consumption/outcome pricing. Vendors abandoning per-seat is the tell that the model is under pressure. The genuinely open question is whether consumption revenue replaces seat revenue dollar-for-dollar per account, and that is company-specific: ServiceNow looks best-positioned (tenant-level consumption pools decoupled from users), Workday weakest, since its bear case uniquely rests on total customer headcount falling, exactly the axis the bull counter defends.
Richly-valued names where the thesis is pre-paid. PLTR at a 75-90x forward P/E (~450% premium) and NET at ~28-35x sales can be right on fundamentals and still deliver pain if growth merely meets expectations. Being correct and being early-at-any-price are different disciplines.
A discipline note on the semis correction (SOX −20%, SMH in a drawdown). The great traders converge on patience here. Paul Tudor Jones kept a three-word rule above his desk: “Losers average losers.” A sector in an active downtrend is not a bargain until the selling exhausts and it bases, averaging down into a falling knife is a different trade from buying the strongest name after the tape confirms. The value counter is real too, Warren Buffett’s “be fearful when others are greedy, and greedy when others are fearful”, but note even Buffett buys quality on sale, not the whole basket. The tension is the whole point: the Jevons fundamentals above argue the semi selloff is likely an overreaction; the tape has not yet confirmed a bottom. Respect both, let the fundamentals pick the names and the price action pick the moment.
My take, the tiers
Weighted to the operator’s screen: product footprint, customer base, earnings power, revenue growth, FCF firepower.
Tier 1, Best footprint × FCF (the franchises): MSFT, GOOGL, META, AAPL, AVGO.
MSFT, the largest enterprise base on earth with a direct COGS cut coming.
GOOGL, only player owning distribution and its own inference silicon.
META, monetizing cheap AI in ads today while weaponizing open weights.
AAPL, 2.5B devices, $62B net cash, zero training-capex exposure, the safest way in.
AVGO, 46% FCF margins on the exact mechanism (custom ASICs) that makes models cheap to serve.
Tier 2, Purest cheap-model beneficiaries: PLTR, DUOL, DOCN, AAPL (again).
PLTR, model-agnostic ontology, 57% FCF margin, Rule of 40 = 145, the thesis incarnate.
DUOL, only company with margin expansion explicitly attributed to falling AI costs.
DOCN, the literal “everyone runs their own” on-ramp, AI ARR +221%.
Tier 3, Consumption-toll compounders: DDOG, SNOW, ANET, MU, MDB, NET.
DDOG, every deployed agent pays the monitoring toll, cleanest usage-priced model here.
SNOW, no seat ceiling, re-accelerating at 33%.
ANET, every inference node needs Ethernet.
MU, HBM sold out regardless of who wins.
MDB, the AI-app system of record, RPO +88%.
NET, the edge inference toll, though see Tier 5.
Tier 4, High-growth, higher-risk: AMD, MRVL, SHOP, APP, CRWV, NBIS.
AMD, second-source upside on MI450 execution.
MRVL, Broadcom’s thesis, earlier and riskier.
SHOP, agent-commerce rails with thin FCF.
APP, elite economics under an SEC cloud.
CRWV / NBIS, maximum torque, debt/concentration fragility.
Tier 5, Priced-for-perfection: PLTR, NET, NVDA. Fundamentals excellent, entry demanding: 75-90x forward earnings, 28-35x sales, and 80%+ embedded growth respectively, Friday showed what a sentiment air-pocket does to that.
Tier 6, Two-sided / at-risk: CRM, NOW, WDAY, GTLB, ADBE, ORCL. Each pairs a real AI growth line against a real structural threat: seat compression (CRM/NOW/WDAY/GTLB), model-native disintermediation (ADBE), single-customer credit risk on negative FCF (ORCL).
Most innovative names, explicitly: Palantir’s ontology (a genuinely new abstraction, the enterprise as a model-swappable API), Google’s TPU/Ironwood stack (the only hyperscaler whose inference costs fall on its own silicon curve), Cloudflare’s AI Gateway + Monetization Gateway (building the toll booths and payment rails of the agent economy before it fully exists), Meta’s open-weights-as-weapon (giving away a billion-download model family to burn the rival’s business model), and honorable mention Amazon’s Trainium-plus-Bedrock neutrality (own the casino, not the gambler).
Clearest-benefit layers: consumption-priced infrastructure and observability (DDOG/SNOW/ANET/MU), edge/on-device (AAPL), and model-agnostic orchestration (PLTR), these win under almost any model-war outcome. Most-debated layers: seat-based app software (the SaaSpocalypse question is live) and training-heavy silicon/GPU-landlords (Jevons vs. efficiency is being re-tested this very week).
Confidence: High that value migrates away from the model layer toward distribution and usage-priced infrastructure, the post-DeepSeek capex evidence and this week’s TSMC -7%-on-+77%-profit dislocation both support structure-over-sentiment. Moderate on the seat-compression timeline for app software. Low on anything valuation-timing related, which is exactly why we don’t handicap price paths.
Bottom line
Cheap models are not the death of the AI trade, they are its industrialization. When intelligence becomes a commodity ingredient, the profits flow to whoever owns the distribution that serves it (Microsoft, Google, Meta, Apple), the meters that bill its usage (Datadog, Snowflake, Arista, Micron, Broadcom), and the model-agnostic layers that get cheaper to run every year (Palantir, Duolingo, DigitalOcean), while the model labs and the seat-priced middle fight over shrinking margins. The single most important idea: own the tolls and the storefronts, not the commodity in between.
Coming tomorrow: the distribution wrinkle
One more angle deserves its own piece. AI is driving the cost of building software toward zero, but the cost of winning a paying human hasn’t fallen a cent, so distribution and marketing become the scarce moat. Tomorrow we map the companies that sell the coming AI-software flood to real customers, Reddit, HubSpot, The Trade Desk, Klaviyo, Zeta, Twilio, and why the incumbents (Adobe, Publicis) are already buying that layer. Coming to SignalDeck.live tomorrow.
📊 Read this on SignalDeck.live.
Not financial advice. For informational and educational purposes only. Do your own research.

