KPI and Reporting Architecture
Choosing the five KPIs that actually run the business
Choose five weekly numbers: one demand measure, one throughput or delivery measure, one margin measure, one cash measure and one people measure. Each needs a written definition, a named owner and a stated source system. In a leveraged acquisition the cash measure is coverage of debt service, and the margin measure has to survive a chart of accounts you did not design. Model determines the other three.
Who this is for. You are two or three months into running a business you bought, you have a general ledger you do not fully trust and a dozen numbers people mention in passing, and you have to decide which handful get reviewed every single week.
Five slots, not five metrics
The question is usually asked as 'which five KPIs should I track', which has no answer, and is better asked as 'which five slots must be filled'. The slots are demand, throughput, margin, cash and people. Demand tells you what is coming. Throughput tells you whether the business can convert it. Margin tells you whether conversion is worth doing. Cash tells you whether you survive the conversion. People tells you whether next quarter still has the capacity to do any of it.
Filling the slots is model-specific and the slots themselves are not. That is what makes this framing portable across a trades company, a distributor and a professional-services firm, and it is why a generic list of operational metrics is useless to someone who just inherited a specific business with a specific constraint.
Three artifacts get confused here and they are different objects. A dashboard is continuous and self-serve — you look at it when you want to know something. A scorecard is weekly, reviewed in a meeting, and every row has an owner. A report is periodic and narrative. The scorecard is the one that changes behaviour, because it is the only one where a human is accountable for a number in front of other humans on a fixed schedule.
At least one of the five should be a leading measure. A scorecard composed entirely of lagging outcomes tells you accurately what has already happened and gives nobody anything to do differently on Wednesday. This is the single most transferable idea in the entire cadence literature and it is the one most often dropped when a team builds its first scorecard, because lagging numbers are easier to get.
The rules that decide whether a scorecard survives month three
One owner per number, and the owner is a named human rather than a department. Shared ownership of a metric reliably produces no ownership of it, and it is the most-violated rule in the category. It is worth being structural about: if you cannot name the person, the row does not go on the scorecard.
Every number needs a written definition before it needs a target. Two people will otherwise report the same metric differently and the disagreement will surface in month three, in a meeting, about a number that mattered. The definition dictionary is a dull artifact that quietly determines whether the scorecard is still trusted in the second quarter. It also protects you from a specific trap: the same word meaning different quantities in different traditions.
Map each number to its source system and budget the manual entries explicitly. Some numbers come out of the ledger, some out of the field-service or inventory system, some out of the customer database, and some out of a person typing. Every manual entry is a future failure, and knowing which rows are manual tells you where the scorecard will decay first.
Adoption is a separate problem from design and it is the one nobody publishes. Every framework steward publishes the artifact; none of them publishes the failure modes. What makes a team fill a scorecard in is that the number is reviewed in a meeting the owner attends, that the result is publicly visible, and that missing it has a consequence — in that order. Cascading the scorecard down to departments and individuals is a later move, and below roughly thirty headcount it usually produces ceremony rather than signal. That threshold is an editorial judgement, not a sourced finding, and it is offered as one.
One dependency runs underneath all of this: the chart of accounts. Margin by service line cannot be produced from a chart that grew by accretion under the previous owner. If the margin slot on your scorecard is going to mean anything, the account structure has to be redesigned to feed it, and that work belongs earlier in the year than most operators put it.
Which numbers fill the slots, by business model
A trades business fills demand with booked-call rate, throughput with completed jobs per technician per day, margin with gross margin by job type, cash with collections timing, and people with technician availability. A professional-services firm fills demand with pipeline coverage, throughput with utilisation, margin with the effective bill rate against loaded cost, cash with collection realisation, and people with bench cost. A distributor fills throughput with lines shipped per labour hour and margin with gross margin dollars against average stock held. A shop fills throughput with cycle time against takt and margin with estimated-versus-actual job cost. A small subscription business fills demand with new bookings and margin with gross revenue retention.
Three same-word traps live inside those lists and each one has cost somebody a quarter. Utilisation means billable hours over available hours in one benchmarking tradition and a dollars-based ratio in another, so two firms reporting the same utilisation figure may not be comparable at all. Overhead rate is a multiple of direct labour in architecture, engineering and professional-services benchmarking, and a percentage of revenue in the trades — same phrase, incompatible scales. And a fifty per cent markup is a thirty-three per cent margin, which is the most expensive arithmetic confusion in trades pricing and the one that quietly breaks a flat-rate pricing conversion.
This is why the definition dictionary is not administrative overhead. It is the thing that stops a number from meaning one thing in the ledger, another in the field system and a third in the weekly meeting.
The target question, answered honestly
The next question is always what the number should be, and here the category performs a trick. Take a metric whose value you want, follow the figure back, and the chain runs: a field-service-software blog states a number, a listicle cites the blog, a consultant cites the listicle, and the number becomes 'the industry average'. There is no study at the bottom. Publishing that chain break is more useful to you than publishing another link in it.
The distribution of citable data is also perverse. Free, rigorous, methodology-disclosed benchmark corpora do exist — an architecture and engineering study covering 896 firms for financial year 2025, and an accounting-firm survey covering 1,073 firms for financial year 2024 — and they cover categories that are acquired comparatively rarely. The two models most commonly bought by operator-acquirers, the trades and distribution, are precisely the two where no such corpus is reachable: the trade association's financial survey is 2020-vintage and paywalled, and every primary distribution source we attempted returned a 403 or a 404.
So the standing rule on these pages is simple. A metric definition and its formula are safe to publish. A metric's typical value is not, unless it traces to a government or academic source we actually retrieved. Where it does not, the page ships the formula, the interpretation, and a plain statement that no verified benchmark exists — plus, where we traced it, where the circulating number came from. That is a more useful product than a fabricated target, and it is one no vendor-funded competitor can copy, because their marketing depends on the numbers.
The practical substitute for a benchmark is your own baseline plus a survivability calculation. Take three to four quarters of your own history on a stable definition and treat the trend as the target. Then compute how much revenue decline the business survives with debt service treated as a fixed cost — break-even, re-derived for a leveraged owner. That number is specific to your business, it is checkable, and it is the one that actually governs decisions.
The formulas, and the benchmarks we will not print
Each metric below publishes what can be computed and refuses what cannot be sourced. A median with no traceable population is not a benchmark; it is a number someone repeated. Where a figure would go, this page says why it is absent.
Labour as a percentage of revenue
fully loaded labour cost ÷ revenue
The dominant cost line in nearly every acquirable service business, and the one where an inherited compensation structure hides. Read it as a trend against your own prior quarters and against the seasonality of your own business, not against an outside figure.
No benchmark, because: Honest benchmarking of this ratio needs wage data by industry code and county, and the statistical agency that publishes it returns HTTP 403 to every request, domain-wide. Until that data is retrieved, we publish the formula and the interpretation and no band, because a labour-cost target sourced from a payroll vendor's blog is not a target, it is advertising.
Revenue per employee
revenue ÷ full-time-equivalent headcount
Useful strictly against your own history. Cross-industry comparison of this figure is meaningless — a distributor and a professional-services firm are not on the same scale — and within-industry bands are what you would actually need.
No benchmark, because: Those within-industry bands are not retrievable for the same reason as above. The only adjacent figures we can publish are computed from the Census employer-firm tables — an average of 16.4 employees per establishment and 21.2 per employer firm — and they must be labelled as computed from that source, never presented as wage-survey figures.
Debt service coverage
cash available for debt service ÷ total debt service
The number your lender watches and the one that constrains every other decision in a leveraged acquisition. Track the trend and the headroom, and know the threshold your own loan documents actually specify, because that is a contractual fact rather than an industry convention.
No benchmark, because: The widely repeated agency minimum is the first item on our do-not-publish list. It is not verifiable from any of the agency's HTML sources; its standard operating procedure is distributed only as a word-processor document in which the clause was not located, and the public programme page states only a requirement of reasonable ability to repay. Read your own credit agreement; do not take a threshold from a blog, including ours.
How many rows a weekly scorecard can carry
rows a team will actually read each week — an editorial judgement, offered as roughly fifteen
Past that point attention degrades and the review becomes a recital. Five weekly rows plus a longer monthly list is a more robust structure than fifteen weekly rows.
No benchmark, because: Framework stewards converge on the phrase 'a handful' and none publishes a failure threshold or a study behind one. We state roughly fifteen as our own editorial judgement and label it as such rather than dressing it as a finding.
When to cascade the scorecard below the leadership team
cascade when a department has both its own controllable numbers and a manager who owns them
Cascading works when the layer below has real decision rights over the number. Where it does not, the cascade produces reporting rather than management.
No benchmark, because: The roughly-thirty-headcount inflection commonly cited for this is an editorial judgement, not a sourced finding, and it is presented as one.
Five slots filled for an acquired commercial plumbing company
A 38-headcount commercial plumbing business, bought with acquisition debt. The operator wants five weekly numbers. Slot discipline decides the list; the model decides the contents.
1. Demand — booked calls per week, by channel
Leading. Owned by the dispatcher. Source: the field-service system, automatically. It moves before revenue does, which is the whole point of putting it in the demand slot.
2. Throughput — completed jobs per technician per working day
Owned by the service manager. Source: the field-service system. Written definition required, because 'completed' has to mean invoiced-ready rather than left-the-site.
3. Margin — gross margin by job type
Owned by the controller. Source: the ledger, but only after the chart of accounts has been rebuilt to separate service from install. Before that rebuild this row is not computable and should be left off rather than faked.
4. Cash — cash available for debt service against total debt service
Owned by the operator personally. Source: the ledger plus the loan schedule. This is the row that distinguishes an owner-operator's scorecard from a manager's.
5. People — unapplied labour hours
Owned by the service manager. Source: payroll against the field-service system. Leading rather than lagging: it moves weeks before turnover does.
6. What is deliberately not on it
Callback rate, membership attach, average ticket and truck utilisation are all real and all useful. They belong to the monthly review, because a weekly list past roughly fifteen rows stops being read — an editorial judgement, stated as one.
The headcount and the slot assignments are illustrative. What is not illustrative is the method: five slots, one named owner per row, a written definition per row, a stated source system per row, and no target number that we could not trace to a retrieved source.
What to take away
Fill five slots, not five metrics: demand, throughput, margin, cash and people. The slots are portable across business models; the contents are not.
Every row needs one named human owner, a written definition and a stated source system. Rows failing any of the three do not belong on a weekly scorecard.
At least one row must be a leading measure. A scorecard of pure lagging outcomes is accurate and useless.
A dashboard, a scorecard and a report are three different objects. Only the scorecard changes behaviour, because only it puts a named person in front of a number on a fixed schedule.
Margin by service line requires a deliberately rebuilt chart of accounts. Until that exists, leave the row off rather than reporting a figure the ledger cannot support.
The two business models operator-acquirers buy most — the trades and distribution — are exactly the two with no reachable benchmark corpus, while the categories that do have free rigorous data are acquired far less often.
Where no benchmark is verifiable, publish the formula, the interpretation and the reason. Your own three-to-four-quarter baseline, plus how much revenue decline the business survives with debt service as a fixed cost, is a better target than a number you cannot trace.
Sources
Clarity Architecture & Engineering Industry Study, 47th annual — 896 firms, FY2025
Deltek · A2
Used for: The existence, size and vintage of a free, methodology-disclosed professional-services benchmark corpus — 896 firms, financial year 2025 — as one half of the benchmark asymmetry.
National MAP Survey 2025 executive summary — 1,073 firms, FY2024
AICPA · A2
Used for: The accounting-firm survey covering 1,073 firms for financial year 2024, as the other half of the asymmetry.
US Census Bureau · A1
Used for: Employer-firm employment averages computed from the Census size-band table, used as the only publishable adjacent figures for revenue per employee.
Occupational Employment and Wage Statistics
US Bureau of Labor Statistics · A1 · blocked
Used for: Documenting that the wage series required to benchmark labour cost honestly is unreachable, which is why no band is published.
SOP 50 10, Version 8, effective 2025-06-01 (.docx only)
US Small Business Administration · A1
Used for: Establishing that the coverage-ratio threshold in circulation is not locatable in the governing document.
The EOS Model — Six Key Components: Vision, People, Data, Issues, Process, Traction
EOS Worldwide · B1
Used for: The scorecard's stated purpose — reducing the organisation to a handful of objective numbers — quoted as what the framework prescribes.
The 4 Disciplines of Execution (McChesney, Covey, Huling)
FranklinCovey · B1 · to-verify
Used for: The lead-versus-lag distinction that requires at least one leading row.
Permanent Equity · B2
Used for: Practitioner treatment of instrumentation inside operating companies, attributed as a practitioner view.
Blog — trades operating and margin content
ServiceTitan · C1
Used for: Evidence of what a field-service-software vendor publishes about trades margins — cited to demonstrate the origin of circulating figures, never as evidence that they are accurate.
Financial survey — 2020 vintage, paywalled at $325
Air Conditioning Contractors of America · C1 · gated
Used for: Evidence that the trade association's financial survey is 2020-vintage and paywalled, which is why trades figures attributed to it cannot be verified.