The CEO of a PE-backed company should put the first AI money of the hold on one task the company already counts per person each week, compare a group with the tool against a matched group without it over the same quarter, and skip projects whose only proof would be staff estimates.
An NBER working paper published in March 2026, not peer reviewed, surveyed nearly 750 corporate executives between November 11, 2025 and mid-January 2026, 603 of them on The CFO Survey panel that Duke University and the Federal Reserve Banks of Richmond and Atlanta run each quarter. Asked directly, they put 2025 labor productivity growth attributable to AI at 1.8 percent on average. Worked out from their own estimates of what AI did to revenue and employment, it was 0.6 percent, a gap the authors say likely reflects a delay in revenue realizations. In a supplemental first-quarter 2026 CFO Survey question, answered by 183 CFOs from the paper's main sample, companies with at least 500 employees expected to put, on average, 55 percent of their 2026 AI spending into operations, meaning subscriptions, services and training, and smaller companies 64 percent.
A study published in The Quarterly Journal of Economics in May 2025 followed the staggered introduction of a generative AI conversational assistant, using data from 5,172 customer-support agents. Issues resolved per hour rose 15 percent on average, with substantial heterogeneity: less experienced and lower-skilled workers improved both speed and quality, while the most experienced and highest-skilled saw small gains in speed and small declines in quality.
A support desk's resolved tickets or an accounts payable team's posted invoices per person each week are the same kind of measure. The COO picks one such task done by at least 20 people, Nine-67's floor, and splits them in half. Both do the same work, matched on tenure and on their last 13 weeks of output per person, held in the company's system so the baseline needs no new tooling. The COO gives the tool to one group and holds the other back for 13 weeks. Each week the controller pulls both groups' output per person with errors per 100 items, such as reopened tickets or invoice errors, and reports both to the CFO and CEO at week 13. The continue test is Nine-67's own threshold, which neither source sets. Over the 13 weeks, the tool group's output per person must rise at least 5 percentage points more against its own 13-week baseline than the other group's, with errors per 100 items no higher than the other's. If both hold, the CEO should continue, giving the second group the tool in week 14. The company pays from the operating budget of the department doing the task under a 13-week contract that renews only on the CEO's decision.
The CEO should skip, for the hold, any project whose case rests on staff estimates of hours saved with no before and after count of the work, since an estimate cannot show whether the time saved became output, and a count can. An operating partner can ask each portfolio company for each group's output per person over its 13 weeks and the 13 before, and its errors per 100 items, and should compare the gap between groups only across companies on the same task.
The same survey found little evidence that AI has had, or in the near term will have, large effects on the total number of employees or on costs, with one moderate exception: companies with at least 500 employees expect to reduce headcount by 0.7 percent in 2026 due to using AI.