
In one product we measured, the data advantage was spent within five examples.
A single capability, contract review, can now be built in more distinct ways than a diligence team could test in a quarter. The model underneath can be any of a dozen frontier systems. Each can be aimed through prompting, retrieval, fine-tuning, or tool use. Each sits inside a product that wraps it in workflow. The combinations run into the thousands, and almost none carry an independent measurement of what they do.
This is the practical problem behind most AI assessment today. A buyer, an investor, or an operator is not choosing between two clear options. They are looking at one point in a large and shifting space, described by the party with the strongest incentive to describe it well.
Independent benchmarks exist, but they measure models in isolation, not the product a buyer actually pays for. The distance between the two is where advantage is won and lost.

Every reported measure of an AI business blends two different things. One is the market it sells into: demand, competition, switching costs, regulation. The other is the product itself: what it can actually do, and how hard that is to reproduce.
The market side is well served. Customer conversations, channel checks, and comparable analysis read it reliably, and the methods are mature.
The product side is the harder one, and it has become harder still. What a model can and cannot do is now changing on a timescale of months. The people closest to the product, its own users, cannot see this clearly, because the capability sits below the surface they interact with.
So the intrinsic strength of the product, the part most exposed to AI, is also the part least visible to a conventional read. One half of the assessment has no reliable instrument.

Reading the product side becomes tractable once the product is broken into parts. Whatever the application, the same four components are present, and each can be measured on its own.
This frame carries everything that follows. It locates, for a specific product, where each part contributes and how reproducible that part is. The advantage can sit in any one of these four components, or in several at once.

Units. The functional blocks that do the work (a model, a classifier, a solver). A model is now the baseline unit, available to rent.
Data and knowledge. Everything the product carries: raw data, structured records, encoded know-how (templates, labelled sets, domain rules).
Connections. How the parts are wired together and to outside systems (retrieval, integrations, APIs, workflow steps).
Logic. How the model is aimed and orchestrated (prompts, decomposition, tool routing, control logic).
A compound, a product's overall performance, tells you nothing about its individual parts. To read one component, hold everything else fixed, move only that one component, and observe the change in outcome.
This is the standard method of experimental science. It is the one thing that assessment based only on the outward result can never do, because from outside you observe only the combined effect.
The moment a component can be isolated and varied independently, its contribution stops being a matter of opinion and becomes a measured fact.

Each component has its own way of being tested, but the method is the same: isolate it, vary it, read the result against independent ground truth. What is measured for each is set out below.
Units: how much the unit adds over the best available rented baseline, and how quickly that baseline is catching up.
Data and knowledge: the marginal value of each additional increment of data, typically steep at first, then bending, then flat.
Connections: how dependent the outcome is on the wiring, and how easily that wiring could be reconstructed by a competitor.
Logic: how sensitive the outcome is to how the model is aimed, holding the model and the data fixed.

A base model is now the default unit of capability. It is the basic competence a product gets out of the box, before any wiring, data, or tuning. For most tasks, that baseline is already strong.
Two things followed. The quality of the baseline unit rose sharply, and the price for equivalent quality fell by roughly ten times a year. A capability that cost about sixty dollars per million tokens in 2021 costs a few cents today.
When the baseline unit is rentable by anyone at a collapsing price, owning a generic unit stops being an advantage. A unit regains value only when it stops being the generic baseline.

Figma's strong unit was a design-surface capability, a browser-native multiplayer canvas for interface design. As the baseline model became able to generate and manipulate design surfaces directly, that capability lost proximity to the frontier.
Atlassian's strong unit was a workflow capability, issue tracking and knowledge structuring for engineering teams. As the baseline model absorbed the drafting and reasoning steps in those workflows, the standalone workflow layer sat closer to the baseline than to a defensible frontier.
Adobe's strong unit was a file-format and editing capability, mastery of image, vector, and document formats and the tools to edit them. Baseline generative models began producing and editing those artefacts directly, compressing the value of the editing tool. Adobe's forward multiple fell from roughly thirty times to twelve.
These were well-engineered companies with strong products and real customers. What they shared was proximity to the baseline unit, and that is what drove the repricing. A product whose advantage sits in data or connections is a separate case.

The collapse in the baseline price does not make units worthless. It moves the value to units that a competitor cannot simply rent.
A unit regains durable value in three main ways. Each turns the unit from a shared commodity into something specific to the product, and each is measurable.
Holds unique data in itself. A unit trained on data no one else has stops being generic. The weights carry the advantage, not just the architecture.
Is specialised or fine-tuned. A model tuned for a narrow domain can beat a larger general one on that domain, at lower cost. Small, domain-tuned models are the clearest case.
Runs where the general model cannot. On-device, in a secure environment, or under latency and cost limits a frontier model cannot meet.
The question for any product built on a model is simple: is its unit still the generic one, or has it become one of these? The answer sets how exposed the product is to the next model release.

Data and knowledge is the second component. It is widely treated as a strong layer of protection. The real question is the point at which additional data stops adding value.
Compounding case. An email-deliverability product keeps improving as volume grows, because genuine infrastructure and feedback learning turn each new send into signal. Scale keeps compounding.
Flat case. A product without that infrastructure effect flattens early, because the underlying advantage was shallow. More data buys almost nothing past the first climb.
The plateau location is specific to each product and has to be measured, not assumed.

The study. We took one contract-review product, the kind seen in a legal-technology deal, and measured how much its performance improved as we supplied more worked examples, scoring each result against answers expert lawyers had marked correct.
The first five examples moved the score by 0.286. The next forty-five moved it by 0.041. The first five were worth roughly seven times the next forty-five combined.
The advantage was front-loaded and then flat. Past a handful of examples, more data bought almost nothing. A moat described as compounding was, on measurement, a stock already spent.
Similar tools likely carry only a limited, front-loaded data advantage that is largely spent within a small number of examples.

This is the third and fourth component read, on the same product, from the same study. With the unit and the available data held fixed, we rebuilt the work one layer at a time.
Retrieval and definitions moved the score by very little. The connections, the parts that looked like the engineering moat, moved it barely at all. What appears to be defensible plumbing often is not.
Each layer has to be tested on its own before it can be credited with the advantage. This is the same isolation method: hold the rest, move one, read the outcome.

Three products Caliper Lab has evaluated, anonymised here and read through the same four-component method. The same analysis places the advantage in a different layer each time, and rarely where the pitch would point.
Contract-review legal tool. A tool that reads contracts and extracts key terms, built on a rented general model. Advantage in: Logic. The advantage is in how the model is aimed. The model is generic and the data runs out fast, so the only defensible part is the earned way the product directs the model at the task.
Research and citation platform. A tool that surfaces and cites source material, built on a rented general model. Advantage in: Data and knowledge. The advantage is in the data. The model is undifferentiated, but the proprietary corpus and citation graph are deep and hard for a competitor to reconstruct.
Forecasting and planning tool. A tool that forecasts on tabular and time-series data, built on a purpose-built model. Advantage in: Units. The advantage is in the unit. The model is specialised for tabular data and beats any rented baseline, so the value sits in the model itself rather than the wrapper.

The four-component read is not only a way to understand a market. It is a way to make specific decisions with less guesswork, on both the investment side and the operating side.
In each case the question is the same: what actually carries this product's performance, how durable is it, and what would a competitor have to reproduce.
For an investment decision:
For an operating decision:
About the author: Dhruv Gulati is the founder of Caliper Lab, an independent AI research and evaluations firm that reads products component by component against independent ground truth. thecaliperlab.com