2026-04-28 · 8 min read

Three Monetization Swings I'd Run on Claude This Quarter

Three Monetization Swings I'd Run on Claude This Quarter

Key takeaways

Decision filter. Before pricing any feature: Is there a named buyer? A dollar value? A real reference point to anchor it?

Companion tool: the willingness-to-pay math in this essay started as a spreadsheet. I rebuilt it as the AI Unit Economics lab, so you can model cost per task, cache leverage, and breakeven yourself. It runs in your browser.

Why the pricing is wrong now

Anthropic's revenue went from $1B to $14B in fourteen months, and it still prices Claude per token. Claude Code is running at $2.5B. Forty percent of enterprise production API usage runs on Claude. The $30B Series G says the market believes the trajectory is real.

But the public numbers from the first quarter of 2026 also expose three structural pressures the current pricing architecture was never designed to handle. All three are things the pricing team could move on this quarter.

The margin bridge. Anthropic's reported 2025 gross margins came in around 40%, ten points below internal estimates, because inference costs ran 23% over plan. The Series G projections require margins north of 75% by 2028. That's a 35-point bridge to close in 24 months while burning more than $6B a year, and Dario himself has called a 12-month delay a bankruptcy risk. So every monetization move from here should answer one question before it answers anything else: what does this do to gross margin?

The variable-workload problem the OpenClaw cutoff exposed. Boris Cherny's framing on April 4 was honest: subscriptions were never designed for continuous, automated agent demand. But the fix punted those customers entirely. What the OpenClaw decision really revealed is that the current pricing architecture is wrong for the agent era, and that Anthropic was willing to defend margin at the cost of developer goodwill while it worked out the right one. I read that as an open problem, not a closed door.

The value-metric mismatch. Anthropic charges per token on the API, per seat on Team, per capacity multiple on Max (5x and 20x). None of those line up with what enterprise buyers actually measure. A VP of Customer Experience, Engineering, or Operations measures cost per ticket, cost per PR, cost per claim. The token bill is denominated in a different currency than the value it delivers. That gap is money left on the table.

The three swings below go straight at each pressure. Together they make up one answer to the gross-margin question the public market will ask at S-1.


Run these three swings

Swing 1: Claude Outcomes (pricing the work, not the tokens)

Anthropic ships a new B2B SKU that charges per completed business outcome instead of per token. A resolved ticket, a merged PR, a processed claim, a generated report. You price the thing the customer's CFO already counts, and you capture ten to fifty times more revenue per workload doing it.

The pattern is already proven. Intercom's Fin product priced customer service at $0.99 per resolved ticket and it worked. The swing here isn't inventing outcome pricing. It's generalizing it across every agentic workflow Claude can run to a verifiable spec: PR review, document generation, claim processing, compliance review, anything Cowork can deliver.

The willingness-to-pay math is clean. A fully loaded customer service ticket costs a company $5 to $15 today. A senior engineer's code-review pass costs $50 to $150. A processed insurance claim costs $8 to $30. Claude's token cost on any of these is somewhere between five and fifty cents. So Anthropic prices at 20% to 40% of the human-labor baseline, call it $1 to $3 a ticket or $5 to $15 a PR, and both sides win. The customer saves 60% to 80% against what they pay now, and Anthropic books ten to a hundred times the revenue per workload at the same inference cost.

That last point is the whole case on margin: same cost to serve, far more captured per unit of that cost. The competitive logic is counter-positioning. OpenAI can't easily ship outcome pricing without exposing the margin gap on the consumption-priced enterprise contracts it has already sold.

Swing 2: Compute-Aware Gain-Share (get paid for saving the customer money)

Anthropic ships a Compute Optimization Service for enterprise: an actively managed routing layer that lowers a customer's Claude spend, billed as a 20% share of the savings it produces. FinOps consultancies like CloudHealth and Apptio have run this play for years on cloud bills. This is the same play, at AI scale.

Here's why it matters right now. The margin bridge is Anthropic's number-one financial pressure, and today the routing is wrong in both directions. The customer who defaults to Opus on every query is quietly subsidized. The customer who runs mostly Haiku-class work is quietly overcharged. Compute-aware routing fixes both: auto-routing on Pro and Max, an "Auto" SLA-bound endpoint on the API, gain-share contracts on Enterprise. Of the three swings, this is the only one that improves blended gross margin directly, without waiting on a slow new market to materialize.

The numbers work. FinOps cost optimization is a $5B market growing 35% a year, and it charges 15% to 25% of savings. A customer spending $5M a year on Claude who saves $1M throws off $200K of high-margin recurring revenue. Nobody in AI prices on cost optimization today, so this is a genuinely new pricing primitive, not a copy of someone else's SKU. Realistically I'd expect a 5 to 10 point lift in blended gross margin on the inference layer in year one, which at Anthropic's scale is hundreds of millions of dollars. The power is twofold: scale economies, because better routing means better unit cost, and process power, because the routing intelligence is an operational capability that compounds.

One thing has to be true for this to work, and it has to ship with the SKU rather than after it. The customer needs per-query model attribution they can see, a quality SLA a third party can verify, and a one-click override back to manual selection. The trust architecture is not a follow-up. It's part of the product.

Swing 3: Spot Inference (solve OpenClaw the right way)

Anthropic opens a Spot Inference tier: inference at 40% to 70% off pay-as-you-go for work that isn't urgent and can be interrupted. Overnight batch jobs, background agents, training-data generation, research workloads. The price floats on a transparent supply-and-demand market, the same way AWS Spot prices idle EC2 capacity.

The timing is the point. OpenClaw exposed that a subscription can't price continuous agent demand. A market can. Right now the customers Anthropic showed the door are running on OpenAI or self-hosted Llama. Spot Inference brings them back without recreating the problem that got them cut off in the first place.

The economics are almost too good. AWS Spot runs 50% to 70% below On-Demand and accounts for 15% to 20% of EC2 compute, so the pattern is durable: give people with deferrable work a discount and they'll trade latency for it. Anthropic's marginal cost on an idle GPU is close to zero, which means every spot dollar is incremental margin, not revenue cannibalized from somewhere else. This is pure utilization. Idle capacity becomes revenue at any price above zero. It's scale economies again, and it quietly counter-positions OpenAI, whose Azure-locked compute is less flexible for spot-market dynamics.


What Anthropic has already shipped, and how each swing extends it

The strongest version of each swing isn't built on a blank slate. Each one extends something Anthropic has already validated. Three public datapoints from the last quarter make the case that the insight is already proven. What's missing is the productized pricing primitive on top of it.

Spot Inference extends the Batch API and the fungible-compute strategy. Anthropic already sells latency for a discount. The Batch API runs at 50% off and returns in hours instead of seconds, and prompt caching cuts cached input by 90%. In his "cone of uncertainty" interview on Invest Like the Best, CFO Krishna Rao described fungible compute: workloads moved between training, inference, and internal use on short horizons, with daily allocation meetings to decide where the capacity goes. So Anthropic already runs the interruptible-capacity flywheel internally, and it has already proven that customers will trade latency for price. Spot Inference just puts a market price on it. Let customers bid for the interruptible capacity Anthropic is currently reshuffling by hand.

Compute-Aware Gain-Share extends internal allocation and plugs an external leak. That same interview is, in effect, a public description of compute-aware allocation as standard operating practice. The capability already exists inside the company. It just isn't pointed at customers yet. Meanwhile the value is leaking off-platform in real time. Third-party routers like OpenRouter and the open-source Claude Code Router cut Claude costs by anywhere from 50% to 99%, by sending the 60% to 70% of prompts that don't need a frontier model down to cheaper tiers. That's margin arbitrage Anthropic is currently handing to middlemen. Gain-Share brings the capture back in-house and, crucially, points it at the customer's interest instead of against it.

Claude Outcomes is the one true greenfield. No internal precedent, no first-party signal, and Intercom Fin still the only external proof point. It's the highest-novelty, highest-counter-positioning swing of the three, and the one least exposed to the "we're already doing a version of that" objection.

One caveat on the margin framing. The margin math still holds. Anthropic revised 2025 gross margin from around 50% to around 40% and targets roughly 77% by 2028. But the story around it has moved from defense to confidence. Rao, and separately Gavin Baker, frame inference as profitable on its own, with the bridge closing through per-generation efficiency gains and fungible compute rather than through pricing innovation. So I wouldn't pitch Compute-Aware Gain-Share as the answer to a margin crisis. I'd pitch it as additive to the efficiency work already underway, and as the fix for value leaking out to third-party routers. The swing survives the reframe. The framing just needs to match what the team is already saying in public.


Why this portfolio holds together

The three swings aren't variations on one idea. Each does a different job.

Outcomes opens a new revenue frontier: you charge for the value, not the consumption. Gain-Share improves margin on the business that already exists: you capture the routing intelligence as recurring revenue instead of giving it away. Spot Inference improves utilization: you turn idle capacity into revenue and win back the customers you defected.

Each one also leans on a different moat. Counter-positioning for Outcomes, process power for Gain-Share, scale economies for Spot. Each is either a 90-day pilot or a 12-month SKU. Each survives the willingness-to-pay-first test I'd hold any pricing idea to: a named buyer, a dollar value, and real reference points to anchor it. Put together, they give Anthropic three independent routes to the gross-margin trajectory the public market will want to see at S-1.

Price the outcome, not the tokens.


(Note: the willingness-to-pay numbers and margin estimates here are illustrative ranges, built from public reference points and outside observation. I'd want to validate them against Anthropic's actual customer data before committing to specific targets.)