The Financial Architecture of the AI Operating Model
Paper 2 of the AI Operating Model series — Why the money behaves counterintuitively as a firm crosses each operating model, and why the metric everyone watches is the wrong one
The first paper in this series asked which of four operating models an enterprise is really running. This one asks the question a CFO asks immediately after: we bought the AI, we can watch the teams move faster, so where is it on the income statement?
For most firms in 2026 the honest answer is uncomfortable. It is in the cost line, not yet the profit line. The copilot plateau that the last paper described as an organizational condition — real speed at the task level, no change to the structure — is also a financial one. A copilot is bought on top of the existing organization, so it is a cost added before any cost is removed. That is not a failure of the technology. It is the financial signature of buying AI without redesigning around it.
Which sets up the finding that should reframe how every operator reads their own numbers over the next three years. When redesign does happen — when AI stops being a tool bolted onto human workflows and becomes the workflow — the first thing that moves on the income statement is gross margin, and it moves down. Not because the business is deteriorating, but because the cost of delivery is migrating from people into inference, from a line buried in SG&A to a line sitting in cost of revenue. Read naively, that looks like decline. Read correctly, it is the transition working.
Four operating models, four financial models
Each of the four archetypes from the first paper is, underneath, a distinct financial model, defined not by how much AI a firm has adopted but by where AI cost lands and how revenue is priced against it. Read them in order and one thing moves steadily: revenue decouples from labor hours.
In the Hierarchical firm, labor is both the largest cost and the coordination mechanism. In labor-led services, payroll runs 45 to 60% of revenue, and revenue is coupled directly to hours — time and materials, full-time equivalents and seats. The balance sheet is light and the risk is wage inflation and utilization. Inference cost is near zero, because there is almost no inference.
The AI-Augmented firm — the copilot archetype, where the overwhelming majority of “AI transformation” actually lives — adds cost rather than removing it. Internal copilots land in SG&A and R&D, customer-facing ones land in COGS, and the base grows before it shrinks. The reason the plateau does not move the P&L is worth stating plainly: task-time saved is not EBITDA. It becomes EBITDA only when the freed capacity is either taken out as cost or redeployed into sold output, and in this archetype it rarely is — the hours are freed and then quietly reabsorbed. Revenue, meanwhile, is still seat- and subscription-based, and the market pays weakly for generic “AI features.” This is the trap archetype, the one where a firm can spend indefinitely and change nothing structural.
The move to AI-Native is a break in the shape of the income statement. Delivery cost migrates firmly into COGS, because the product is now the AI. Gross-margin percentage resets structurally lower — the AI-native band settles around 50 to 60%, against 80 to 90% for classic software — and the payoff appears one line down, as human coordinators are removed and SG&A leverage emerges. This is the first archetype where revenue genuinely decouples from headcount, and the numbers are stark: AI-native firms are posting revenue per employee of $2.7 to $3.3 million, five to ten times the incumbent-SaaS level.
The move to Outcome-Native is a break in kind, not degree. It is not more AI; it is a different contract, in which the provider is paid for a delivered result and therefore absorbs delivery risk onto its own balance sheet. Revenue indexes to the outcome — a resolved ticket, a recovered payment or a guaranteed level of uptime — rather than to tokens, seats or hours. This produces the stickiest revenue in the set and, done carelessly, the largest liabilities. It is the archetype where the strategic prize and the financial hazard are both greatest.
The margin falls as the model works
The single most important accounting fact in this transition is that gross-margin percentage can fall while gross-profit dollars, EBITDA and free cash flow all rise. It is not a theory. It is disclosed, in the filings of companies living it now. ServiceNow’s subscription gross margin fell from 85 to 83.5% while its free-cash-flow margin rose from 31.5 to 35%. Duolingo’s gross margin slipped roughly 100 to 200 basis points while adjusted EBITDA margin expanded about 330. MongoDB’s gross margin fell from 75 to 73% in the same period it posted its first GAAP operating profit. The mechanism is identical in each case: generative-AI inference is booked into cost of revenue and compresses the margin percentage, while the coordination and overhead cost it displaces falls faster one line down.
So the pivotal instruction for any operator, and any investor reading an operator’s numbers, is this: do not read a declining gross-margin percentage as declining earnings. Watch gross-profit dollars and the trajectory of SG&A instead. At ServiceNow, Duolingo and MongoDB, the reverse of the naive reading is true. The percentage is the wrong scoreboard. The dollars are the game.
The trough you should budget for
None of this happens on the timeline the pilots imply, and the historical record is unusually clear about the shape. General-purpose technologies do not pay off on installation. They demand large complementary investments — process redesign, retraining, reorganization and data remediation — that are expensed early and depress measured productivity before they lift it. The result is a J-curve: things get worse on paper before they get better in fact.
Enterprise resource planning is the cleanest firm-level analog we have. In a typical ERP program, hardware and software were under 20% of a roughly $20-million install; the rest was the intangible work of changing how the organization operated. Payback ran 1.7 to 3.2 years, and 41% of programs delivered less than half the benefits they had anticipated. The lesson is not that the technology failed. It is that the firms which treated it as a purchase rather than a redesign fell into the long half of that distribution.
The most useful precedent, though, is the automated teller machine, because it demolishes the assumption that automation means immediate headcount reduction. When ATMs arrived, tellers per branch fell from about 20 to 13, and the number of bank branches rose 43%, so total teller employment held and even grew for years. Capacity was freed long before it was harvested. That is the pattern to budget for: capacity comes free first, and converting it into either lower cost or new output is a separate act that takes time.
And here the honesty the first paper insisted on has to be stated in financial terms, because AI’s central bet differs from every analog above in one way that cuts against the timeline. ERP and cloud substituted infrastructure; robotic process automation and offshoring substituted task labor. None of them substituted coordination — and coordination is exactly what AI targets. The coordination layer is also historically the stickiest thing in the organization to remove: delayering was confidently predicted in 1988 and largely failed to happen for thirty years. This is the place where the value thesis is strongest and the precedent thinnest at once. The disciplined response is to model the trough from ERP, not from the optimistic RPA case — roughly twelve to eighteen months of accumulation and a further three to nine of conversion dip — and to beat that timeline only where the evidence earns it.
Paying for the outcome
The fourth archetype changes not just the cost structure but the contract, and the economics deserve precision, because this is where the largest prize and the largest hazard both live.
Outcome pricing is not an AI invention. Rolls-Royce sold Power by the Hour — jet-engine uptime rather than engines — beginning in 1962, and energy-service companies have been paid out of guaranteed savings since the 1980s. AI did not invent paying for the result; it collapsed the cost of observing and coordinating the result, which put the model within reach of firms that could never before afford the telemetry. The agent version is simply priced per unit of delivered outcome, a resolution or a recovery rather than a seat.
But assuming delivery risk onto your own balance sheet is a serious act, and the discipline is simple to state and hard to hold: define the billable outcome narrowly and audit it inside your own system, freeze an agreed baseline before any value is shared, and treat the downside as a priced choice rather than a default. The spectrum of that downside runs from bounded penalties capped at fees paid, through shared savings, to unlimited exposure — and the wind-turbine manufacturers who guaranteed availability on immature technology, and booked billion-dollar-plus warranty provisions for it, are the standing warning against guaranteeing an outcome your capability cannot yet reliably deliver. The prudent path is upside-only gainshare first, graduating to two-sided risk only once quality is demonstrably stable.
How to move forward
The first paper argued for knowing which operating model you are running and moving deliberately between them. Reading the financial architecture correctly comes down to three disciplines of its own.
Watch the right line. Gross-margin percentage is the metric almost everyone will use to judge an AI transformation, and for the firms actually doing the work it will fall exactly when things are going right. Track gross-profit dollars and the trend in SG&A as a share of revenue instead. Falling SG&A at stable or better service levels is the clearest evidence that redesign, and not merely spend, is underway.
Instrument capacity conversion. The single most important number in this transition is also the one most often missing: the share of freed capacity actually converted into cost taken out or output sold, rather than left idle. Do not let a business case book task-time savings as EBITDA until that conversion is measured. It is the metric that separates a genuine AI-Native transition from a well-funded stay on the plateau.
Price for the outcome, or watch the value leak away. In services above all, the client who benefits from your efficiency is also the one able to reclaim it at renewal through price. Savings you cannot hold onto are not savings, and moving toward outcome and gainshare pricing is what lets a firm keep the value its own redesign created.
The performance gap the first paper described — between the firms that redesign and the firms that merely equip — will show up in the accounts before it shows up anywhere else, and in a form most operators are not trained to read. The gross-margin line will fall for the firms doing it right and hold for the firms doing nothing, and by the time the profit and retention numbers make the difference obvious, the gap will be structural. The question is no longer whether AI creates value. It is whether your operating model is built to capture it.
Sources
- E. Brynjolfsson and L. Hitt, “Beyond Computation: Information Technology, Organizational Transformation and Business Performance,” Journal of Economic Perspectives (1998).
- J. Bessen, on bank tellers and the ATM, Learning by Doing: The Real Connection between Innovation, Wages, and Wealth (2015).
- Firm-level margin and cash-flow figures (ServiceNow, Duolingo, MongoDB) from company filings and disclosures.
- ERP program economics (install cost mix, payback range, benefit-realization rates) from the enterprise-systems implementation literature.
- Rolls-Royce “Power by the Hour” (from 1962); energy-savings performance contracting (from the 1980s); wind-turbine availability-warranty provisions from company disclosures.
- Organizational-delayering history draws on the flat-organization literature and Block/Sequoia, “From Hierarchy to Intelligence” (2026).
Note: company-specific figures trace to the underlying research base for the series, where each carries its evidence tier. Realized results are stated as realized; forward or third-party estimates are flagged as such.