SOLTEN & CO. RESEARCH | PUBLIC RESEARCH BRIEF | 25 AUGUST 2026 | RA-0009
| KEY THESIS The correct unit of analysis is the accepted workflow—not the seat, token or API call. Falling model prices help, but durable economics appear only when value captured per accepted workflow improves faster than total workflow intensity, rework and non-model variable costs. |
Research snapshot
| Field | Detail |
|---|---|
| Research question | How should investors underwrite AI-native application software when marginal cost, output quality and customer value all vary with usage? |
| Evidence cut-off | 25 August 2026 |
| Audience | Family offices, venture/growth investors, lean investment teams and strategy leaders |
| Reading time | 12–15 minutes |
| Scope | Application-layer economics; not a company rating, financing analysis or valuation opinion |
Executive summary
Software investors learned to value businesses whose cost of serving an additional user approached zero. AI reintroduces a meaningful variable-cost stack: model inference, tool calls, retrieval, orchestration, evaluation, human review and support. The familiar SaaS shorthand—annual recurring revenue, seats and gross margin—still matters, but it can hide whether greater adoption makes the product economically stronger or merely more expensive to operate.
The market’s first instinct has been to treat inference prices as the answer. That is too simple. Stanford’s AI Index found that the cost of querying a model with GPT-3.5-level benchmark performance fell from about $20 per million tokens in November 2022 to $0.07 by October 2024, a decline of more than 280 times.[1] Yet an agent that makes five times as many model calls, uses longer contexts or triggers more tools can consume the entire saving. An 80% decline in unit price combined with a fivefold increase in workload leaves total model cost unchanged.
Quality makes the equation harder. In a field study of 5,179 customer-support agents, generative AI raised issues resolved per hour by 14% on average and by 34% for novice and lower-skilled workers.[5] In a separate experiment involving 758 consultants, people working inside the technology’s capability frontier completed more tasks, worked faster and produced higher-quality output; on a task outside that frontier, AI users were 19 percentage points less likely to reach the correct answer.[6] The economic value of AI therefore depends on the workflow, the user and the cost of detecting failure—not merely on access to a capable model.
Our conclusion is that investors need a new primary unit: contribution margin per accepted workflow. “Accepted” means the customer or a defined control process accepts the output without material rework, reversal or undisclosed human completion. Revenue should be measured against every variable cost required to produce that accepted result, including failed attempts and escalations. This approach makes unlike pricing models comparable and exposes companies whose apparent automation is funded by invisible labour or unchecked risk.
What changed
- Intelligence became a metered input. API calls, context length, output length, tools and agent loops can all scale with use.[8], [13], [14]
- Quality became part of unit economics. Evaluation, monitoring and human review are operating costs, especially where errors are consequential.[7], [20], [21]
- Pricing moved beyond seats. Vendors now combine subscriptions, credits, usage and outcomes, shifting cost and quality risk between buyer and seller.[9], [10], [12], [23]
- Model choice became an optimisation layer. Routing, caching and asynchronous processing can reduce cost, but add engineering and governance complexity.[15]–[18], [24], [25]
- Gross-margin dispersion widened. Selected high-growth private AI companies report economics far below mature public SaaS benchmarks.[2]–[4]
Exhibit 1. Benchmark orientation: the margin gap is real, but the samples are not comparable

Sources: Salesforce FY2025 Form 10-K; ServiceNow FY2025 Form 10-K; Bessemer State of AI 2025. Salesforce figure calculated as 1 − $6.198bn / $35.679bn. Bessemer figures are selected private-company averages, not audited market benchmarks.[2]–[4]
The first mistake: treating token deflation as margin expansion
Model prices have fallen quickly, and providers offer cheaper cached input, asynchronous batches and smaller models.[13]–[15], [24], [25] Those mechanisms can improve cost. They do not determine total workflow economics. A workflow may become more ambitious as models improve: longer contexts, more retrieval, parallel agents, repeated reasoning, tool calls and verification. In agentic systems, model calls can branch or loop. Both Anthropic and the major cloud platforms advise using the simplest architecture that meets the task because added autonomy trades cost and latency for performance.[16]–[18], [26], [27]
Exhibit 2. Unit-cost deflation versus workload expansion

Source: Solten & Co. illustrative model. Each cell is (1 − unit-price decline) × workload multiplier. It excludes non-model costs and is not a forecast.
The implication is practical. Management should report both the price paid per unit of compute and the number of units required per accepted result. If model cost per token falls 50% while tokens per accepted workflow triple, model cost per accepted workflow rises 50%. A company that highlights the first number and omits the second is not showing its economics.
The second mistake: measuring automation without acceptance
An automation rate can overstate value if it counts attempts rather than accepted results. A workflow may appear automated while customers reopen cases, employees rewrite the output or a hidden operations team completes exceptions. NIST treats confabulation and downstream reliance as material risks; Google’s deployment guidance describes evaluation as a core operating process and notes that manual evaluation can become a bottleneck.[7], [20]
The quality-adjusted automation rate should therefore be calculated as accepted outcomes divided by attempted workflows. The definition of acceptance must be tied to the customer’s job: a support resolution that is not reopened, code that passes tests and is merged, a claim that survives audit, or a document approved without material rewrite. Intercom’s published approach separates AI involvement from resolution and charges for defined outcomes rather than unsuccessful attempts, illustrating the direction of travel even though the exact commercial definition remains vendor-specific.[10], [11]
Exhibit 3. A four-layer underwriting hierarchy

Source: Solten & Co. framework.
The six metrics that matter
| Metric | Definition | Why investors need it |
|---|---|---|
| Contribution margin / accepted workflow | Revenue less all variable model, tool, infrastructure, review, support and risk costs, divided by accepted workflows | Reveals whether usage creates economic value |
| Quality-adjusted automation | Accepted outcomes ÷ attempted workflows | Prevents failed attempts and hidden labour from appearing automated |
| Workload elasticity | Change in model calls, tokens and tools per workflow relative to unit-price change | Tests whether cost deflation survives product ambition |
| Cost-to-serve distribution | P50, P90 and P99 variable cost by customer/cohort/workflow | Exposes heavy users and tail-risk contracts |
| Retention after novelty | Cohort retention and expansion after initial experimentation | Separates durable workflow ownership from curiosity |
| Provider dependency | Share of cost and quality tied to one model/provider; time and quality loss to switch | Measures bargaining power and resilience |
Pricing is a risk-allocation decision
Seat pricing gives customers budget certainty but leaves the vendor exposed when usage varies widely. Pure usage pricing transfers much of that risk to the buyer, though it can weaken adoption. Outcome pricing aligns with value only where the result is observable and reversals are measurable. Hybrid pricing—base subscription plus included credits or usage—often provides the best bridge because it combines predictability with a limit on unbounded cost. GitHub Copilot’s plans, for example, combine subscription tiers with included AI credits; Stripe documents subscription, usage, outcome and hybrid patterns for AI companies.[8], [9], [12], [23]
What investors should ask
- What exactly is an accepted outcome, and how often is it reversed, reopened or materially edited?
- Show revenue, model cost, tool cost, human review and support per accepted workflow—monthly and by cohort.
- How do P90 and P99 customers compare with the median? Which contracts are contribution-negative?
- If model prices fall 80%, how much more inference does the product consume? If prices rise or discounts disappear, what reprices?
- How much labour sits outside cost of revenue? What happens to gross margin if evaluation and review are classified consistently?
- Does retention persist after experimentation? Is expansion caused by greater customer value or merely by more metered usage?
- Can the company route across models without losing quality? What proprietary context, distribution or integration survives model commoditisation?
Counter-thesis
The strongest counterargument is that this framework may overstate the break from SaaS. Inference prices could fall faster than workloads expand; products could standardise; customers could continue to prefer simple subscriptions; and many AI features may become low-cost functionality inside conventional software. If those conditions hold, mature AI applications could converge toward classic software margins. The framework would then remain useful mainly during the transition.
That counter-thesis is falsifiable. We would reduce the weight placed on accepted-workflow economics if several conditions persist across cohorts: model and tool cost per workflow falls faster than usage expands; human escalation approaches zero without worsening outcomes; P90 cost-to-serve converges toward the median; retention remains strong after novelty; and vendors achieve public-software gross margins without restricting useful consumption. Until then, the burden of proof should remain on the unit economics.
What to watch
- Contribution margin per accepted workflow, especially for mature cohorts.
- Tokens, model calls and paid tool calls per workflow as capability expands.
- Reopen, reversal, escalation and material-edit rates—not only vendor-reported automation.
- P90/P99 cost-to-serve and the share of negative-contribution customers.
- Pricing migration: unlimited seats toward credits, usage caps, outcomes or hybrids.
- Multi-model routing, caching and batch adoption, with quality held constant.
- Retention and expansion after pilots convert into production workflows.
- Gross-margin reconciliation: whether human operations, evaluation and support sit in cost of revenue.
Conclusion
AI-native software can become an exceptional business. It can also look like software while carrying the economics of a service, an infrastructure layer or an insurance contract. The distinction is not visible in ARR alone. It appears at the level of the accepted workflow: what the customer values, what it costs to deliver reliably, how often automation truly succeeds and who bears the tail risk. The companies that learn to improve that equation—through product design, routing, proprietary context, distribution and disciplined pricing—will turn cheaper intelligence into durable margin. The rest may discover that marginal cost did not disappear. It merely moved.
Sources and evidence
Evidence cut-off: 25 August 2026. Vendor pricing and implementation pages are used to document commercial structures and technical mechanisms, not to validate vendor performance. Private-company benchmarks are attributed and treated as directional. Adjacent citations are separated for legibility.
- Stanford HAI, AI Index 2025: inference cost decline for GPT-3.5-level performance. Source
- Bessemer Venture Partners, The State of AI 2025: private-company growth and gross-margin archetypes. Source
- Salesforce FY2025 Form 10-K: subscription and support revenue and cost. Source
- ServiceNow FY2025 Form 10-K: subscription gross-profit percentage. Source
- NBER, Generative AI at Work: field study of 5,179 customer-support agents. Source
- Dell’Acqua et al., Organization Science, Navigating the Jagged Technological Frontier. Source
- NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative AI Profile. Source
- Stripe, AI companies and usage-based billing. Source
- Stripe, AI pricing models. Source
- Intercom, Fin AI Agent outcomes and outcome pricing. Source
- Intercom, Fin AI Agent automation rate. Source
- GitHub, Copilot plans and included AI credits. Source
- OpenAI API pricing. Source
- Anthropic API pricing and prompt-caching structure. Source
- Google Gemini API, context caching. Source
- Anthropic, Building effective agents. Source
- AWS Bedrock, Intelligent Prompt Routing. Source
- AWS, Generative AI Lens — Well-Architected Framework. Source
- Google Cloud Architecture Framework, AI and ML cost optimisation. Source
- Google Cloud, deploy and operate generative AI applications. Source
- OpenAI API reference, Evals. Source
- AWS Bedrock, cost allocation by request. Source
- Stripe, token-based billing implementation patterns. Source
- OpenAI, Prompt Caching. Source
- OpenAI API reference, Batch API. Source
- Google Cloud, choosing a design pattern for an agentic AI system. Source
- AWS Prescriptive Guidance, cost optimisation for agentic AI. Source
Research disclosure
This report is independent research for informational purposes. It is not investment, legal, tax or accounting advice and does not recommend a security or transaction. The analysis relies on public information available by the evidence cut-off. Illustrative scenarios are analytical tools, not forecasts. Private-company operating data are incomplete; readers should obtain company-level cohort, billing, inference and support records before making an investment decision.
