Understanding innovation · Part IV · Organisation, impact, measurement
15 · Measuring innovation
13 min read
In brief
- Innovation indicators fall into three families (inputs, intermediate outputs, market outputs), and none is sufficient on its own: inputs measure effort, not results; patents and publications measure the production of knowledge with known distortions; only market outputs measure value, and they arrive late.
- The most cited piece of evidence in the discipline is negative and liberating: at company level, R&D spending does not correlate with economic performance. What counts is not how much is spent but how: selection, process, and alignment with strategy.
- In the early phases, when revenue is not yet informative, the correct measurement is innovation accounting: actionable metrics, tied to the hypotheses to be validated and capable of changing decisions, as against vanity metrics that grow anyway and decide nothing.
- Every metric turned into a target becomes corrupted (Goodhart's law): whoever designs measurement systems has to expect that they will be optimised, and defend them with a few outcome metrics, revised, balanced by counter-metrics.
- Definitions also count for tax purposes: the Frascati and Oslo Manuals fix what is R&D and what is innovation for statistical purposes, and in practice certifications and tax incentives follow from those definitions. Measuring well is also a requirement of access to the resources of chapter 09.
Section 1
The three families of indicators, and the limits of each
Inputs measure the resources devoted: R&D spending (absolute, or as a ratio to revenue: R&D intensity), research staff in full-time equivalents, risk capital invested at system level. They are the most available and comparable indicators, and for this reason they dominate statistics and comparisons between countries; their limit is constitutive: they measure effort, not what it produces. High R&D intensity is compatible with a process that burns resources on the wrong projects, and comparing companies on this indicator tells you which one spends more, not which one innovates better.
Intermediate outputs measure the production of knowledge: patents (counts, families, citations received as a proxy for quality) and scientific publications. They are real measures of inventive activity, with documented distortions that anyone using them must know: the propensity to patent varies enormously by sector (chapter 11: where secrecy or speed protect better, firms patent little while innovating a great deal, and services and software are under-represented by construction); the value of patents is distributed according to the usual power law (a few are worth a very great deal, the mass almost nothing, so counting rights without weighting them adds up zeros and giants); and all these indicators are manipulable when they become targets (section 3). Used with the necessary caution, they remain valuable: above all as an instrument for reading the competitive landscape (chapter 08) rather than as an internal report card.
Market outputs measure realised value: the share of revenue from new products (typically: introduced in the last three years), time to market, the success rates of launched projects, and ultimately the margins and growth attributable to innovation. They are the indicators that count, with two mirror defects: they arrive late (they measure the innovation of three to five years ago, not the health of today's pipeline) and attributing the result to innovation requires judgement. A mature measurement system combines the three families in a coherent chain: inputs for the sustainability of the effort, intermediate outputs for the productivity of the process, market outputs for the verdict, in the awareness that each family looks at a different stretch of time.
Section 2
The uncomfortable result: how much you spend does not predict how you do
If a single empirical result in the discipline deserved to be posted in meeting rooms, it would be this one. The Global Innovation 1000 series of studies (conducted for over a decade by Booz & Company, then Strategy&, on the thousand listed companies with the highest R&D spending in the world) verified the same question year after year: does R&D spending, absolute or in intensity, correlate with economic performance (growth, margins, shareholder returns)? The answer, stable across all the editions: no. No statistically appreciable correlation, at any level of spending; the companies with the best results were not those that spent most, and the big spenders did not show superior performance. What distinguished successful innovative companies, in the same analyses, were variables of quality and not of quantity: the alignment between innovation strategy and company strategy, the capacity to select projects, the quality of the process from idea to launch, and deep listening to customers in the early phases.
The result is liberating in the precise sense of the term: it moves the management question from "how much should we spend?" (the comfortable question, because it is answered with a number) to "does our process transform resources into value well?" (the real question, which chapters 02 to 08 of this essay answer). And it is consistent with everything the essay has built: if most of the value depends on the selection of problems, on the discipline of stopping and on the quality of validation, then doubling the fuel of a process that selects badly doubles the waste. R&D spending remains an indicator to monitor (below a threshold, you are simply not playing); it stops being an indicator to display.
Section 3
Measuring the early phases: innovation accounting and Goodhart's law
The problem of the early phases. For a project or a startup in its early phases, the indicators of section 1 are silent: revenue does not exist or is not informative, patents are few and recent, and the risk is falling back on the metrics that grow anyway: cumulative registered users, views, followers, mentions. Ries called them vanity metrics, and the criterion for recognising them is operational: a metric is a vanity metric if no decision would change as its value varies. The correct accounting of the early phases (innovation accounting) reverses the criterion: you measure the quantities tied to the hypotheses being validated (chapter 07), chosen because they are capable of changing decisions. The typical families, in order of phase: validated learning metrics (conversion rates in experiments, the outcomes of structured interviews) when validating the problem; engagement and retention by cohort (how many of January's users are active in March?) when validating the product, because retention is the hardest measure of perceived value to fake; unit economics (contribution margin, acquisition cost, payback time) when validating the model; growth efficiency when scaling. The progression is that of the F8/F9 boundary of chapter 07: each phase has its own metrics, and using those of the next phase (growth) before having passed those of the current one (retention, unit margin) is the accounting signature of premature scaling.
Goodhart's law. Any measurement system, at any phase, is subject to the regularity formulated by Charles Goodhart: when a measure becomes a target, it ceases to be a good measure, because whoever is assessed on it optimises it directly, disconnecting it from the phenomenon it was supposed to represent. The examples in the measurement of innovation are a catalogue: targets for the number of patents that produce fragmented, marginal patents; targets for the number of ideas in the pipeline that fill the funnel with candidates nobody will stop (the zombies of chapter 08); targets for the share of revenue from new products that are met by reclassifying a restyling as new. The known defences are matters of design, not of vigilance: few metrics, of outcome rather than of activity; paired counter-metrics that make unilateral optimisation expensive (volume with quality, speed with defect rates, patents with citations or with actual use); periodic revision of the set, because every metric wears out with use; and the preservation of judgement: measurement informs the decision, it does not replace it, and it is this subordination that removes the reward for manipulation.
Section 4
The official definitions, and why they are worth money
The measurement of innovation has two reference normative texts, both from the OECD, that the reader of this essay will sooner or later meet in a very concrete form.
The Frascati Manual (2015 edition) defines what counts as research and development for statistical purposes, with five criteria that an activity must satisfy jointly: it must be novel (aimed at discoveries or non-obvious solutions), creative (based on original concepts and hypotheses), uncertain (about the outcome, the costs, the time), systematic (planned and documented) and transferable or reproducible (the results must be able to be transferred or reproduced). The manual also distinguishes basic research, applied research and experimental development. The Oslo Manual (2018 edition, already encountered in chapter 01) defines innovation and governs its statistical measurement among firms, distinguishing product innovation from business process innovation.
The reason these definitions are worth money is that tax systems have incorporated them, directly or by derivation: the R&D tax incentives of chapter 09 apply to activities that qualify as R&D (or as technological innovation, a category that is typically broader and less incentivised), and the qualification in practice follows the Frascati criteria. In Spain, the distinction between investigación y desarrollo and innovación tecnológica determines very different rates of deduction, and the main route to giving certainty to the qualification is the certification of the activities by accredited bodies (ENAC) and the consequent Informe Motivado Vinculante, which binds the tax administration as to the nature of the certified activities. The operational consequence for anyone managing projects: the systematic documentation of R&D work (objectives, hypotheses, uncertainties addressed, experimentation, results: Frascati's "systematic" criterion) is not bureaucracy added to the research, it is the condition that makes the research certifiable, and therefore financeable with tax instruments. Measuring and documenting well is, literally, a requirement of access to capital.
The chapter closes with a rule of hygiene that the preceding pages have made possible to justify: on the measures that go outwards (investors, partners, the public), the discipline is to declare method and perimeter, to distinguish measurements from estimates and estimates from targets, and to resist the temptation to publish numbers that you are not yet able to support. A results page left empty until there are closed results to report communicates more credibility than a page filled with projections: in the measurement of innovation, as in validation, honesty about uncertainty is a competitive advantage, because it is rare.
In the Volcano method
This chapter gives the frame of the method's measurement system. The two thermometers of the path (TRL for the technical domain, CRL for the commercial one: chapter 06) are the measurement of maturity; the declared metrics of the commercial phases (F8: initial recurring revenue, an identified channel, a measured acquisition cost; F9: cohorts that persist, a positive unit margin, observable organic growth; F10: unit metrics that hold as customers multiply) are the innovation accounting of section 3, with the exact progression that vaccinates against premature scaling: growth metrics arrive only after those of retention and margin. The choice of the page "Nuestras cifras", declaredly empty "until there are closed operations to recount", with the model's figures always qualified as indicative and not as projections, is the rule of hygiene of section 4 applied to communication. And ENAC certification with the Informe Motivado Vinculante, the third lever of the model, is the point at which this chapter joins up with chapters 06 and 09: systematic documentation according to the Frascati criteria makes R&D activities certifiable by third parties, certification gives certainty to the tax qualification, and the tax qualification unlocks the capital that finances the most uncertain stretch of the path. Measurement, maturity and financing are the same system seen from three sides.
Readings
Further reading
- OECD, Frascati Manual 2015: the opening chapters with the five criteria and the definitions; the rest is consulted as needed.
- OECD/Eurostat, Oslo Manual 2018: the current definition of innovation and the logic of the surveys; already a reference in chapter 01.
- E. Ries, The Lean Startup (2011), the chapters on innovation accounting: the distinction between actionable metrics and vanity metrics in the original source.
- B. Jaruzelski and the co-authors of the Global Innovation 1000 series (Booz & Company/Strategy&, various editions in strategy+business): the evidence on the absence of correlation between R&D spending and performance, with the analyses of what really distinguishes successful innovators.
Frequently asked questions
Frequently asked questions
If R&D spending does not correlate with results, why does everyone communicate it?
Because it is available, comparable and reassuring: a single number, verifiable in the accounts, that signals commitment. As a signal of effort it is legitimate; the error begins when it is read as a predictor of results, which the evidence rules out, or used as a target, which under Goodhart's law would corrupt it (spending more is the easiest target in the world to hit). The mature reading: spending tells you whether you are playing the game, the process decides whether you win it.
What is the most reliable metric for a project in the early phases?
The one tied to the riskiest hypothesis being validated: that is the very definition of an actionable metric, and it changes with the phase. If you are forced to choose a single cross-cutting one for the product phase, retention by cohort is the best candidate: it measures whether the product retains those who have tried it, it is hard to fake (unlike cumulative figures, which grow by construction) and it predicts future economics better than any volume metric. The test for unmasking a vanity metric always remains the same: which decision would change if this number were different? If the answer is none, the metric is furniture.
How do you prevent people from optimising the metrics instead of the work?
You do not prevent it entirely: you expect it and make it unprofitable. The design defences: a few outcome metrics instead of many activity ones; paired counter-metrics (every volume with its quality); rotation and revision of the set before it wears out; and human judgement above the measure, declared and practised, because manipulation pays only where the number decides on its own. The signal that the system has been corrupted is typically a disconnection: the metrics improve and reality (customers, cash, outcomes) does not.
Is the documentation required for R&D certifications not time taken away from the research?
It is in part the same work under another name: Frascati's "systematic" criterion (objectives, hypotheses, uncertainties, experimentation, documented results) coincides with the validation discipline that chapters 03 and 07 recommend for internal reasons (justified decisions, honest retrospective learning, a reusable stock of concepts). Done as you go and with method, documentation costs a fraction of what it costs to reconstruct it after the event, and it pays: it gives access to tax incentives, it stands up to audits, and it turns even stopped projects into documented assets. The time taken away from the research is the time of documentation done badly, that is, late.
Disclaimer. This chapter has a training purpose: it describes measurement systems and criteria, including the definitions relevant for R&D certifications and tax incentives, in their general features. It does not constitute tax or legal advice; the qualification of specific activities and the applicability of the incentives depend on the legislation in force and on the specific case, to be verified with qualified professionals and with the competent certifying bodies.
John F. Kennedy, 1962