Agentic AI

Cognitive Horsepower

A working framework from Signal Fidelity Group for measuring the work that survives human review — not the language that AI generates.

Key takeaways

What to remember

  • AI made drafting cheap; it did not make approved meaning cheap. Verification, not generation, is the new bottleneck.
  • Measure cost per approved output, not cost per token — raw volume can be negative leverage when it creates review escalations.
  • cHP separates raw, verified, and constrained throughput, and formalizes review debt through an expected-cost model.
  • A proposed methods framework for regulated knowledge work — open to challenge, calibration, and empirical validation.

Most AI productivity metrics measure generation: tokens, latency, model scores, cost per call, and draft volume. In high-stakes organizations, the bottleneck is what survives review.

Communications, regulatory, medical affairs, public affairs, and disclosure teams need outputs that preserve approved meaning, cite evidence, avoid prohibited claims, and pass review without creating new liability. The cHP framework separates raw generation from verified and constrained throughput — and asks a more commercially useful question: how much approved work does a human-agent workflow produce per unit of time, cost, and risk?

The evidence: why generation metrics fail

AI made drafting cheap. It did not make approved meaning cheap.

The invisible cost of generative AI is the labor of supervision. Three independent findings make the case.

The productivity mirage: In a 2025 METR randomized controlled trial, experienced developers were 19% slower when allowed to use AI tools — even though they forecast a 24% speedup and, after finishing, still believed AI had sped them up. A central driver was review burden: they accepted fewer than half of AI suggestions and spent significant time checking and correcting the rest. [1]

Workload creep: An eight-month Berkeley Haas ethnographic study of 200 employees, published in Harvard Business Review (2026), found that generative AI did not give time back. It expanded scope, blurred the line between work and rest, and drove constant context-switching — with exhaustion and decision paralysis emerging by month six. [2]

The domain error rate: Hallucination stays severe on exactly the high-stakes tasks regulated teams perform. A Stanford RegLab study found leading models hallucinated on 69–88% of specific legal queries, and OpenAI's own reasoning models scored 33% (o3) and 48% (o4-mini) on a person-fact accuracy benchmark. Grounded, source-constrained summarization, by contrast, often runs under 3% — evidence that constraint, not raw capability, produces trustworthy output. [3]

What cHP is — and what it is not

What cHP is

  • A proposed operational metric for constraint-adjusted human-agent labor.
  • A framework for comparing workflows by approved output per unit of risk.
  • A tool for measuring review debt and semantic-fidelity failure modes.
  • A domain-bounded measurement framework for regulated knowledge work.

What cHP is not

  • A validated universal industry standard.
  • A claim that AI and human cognition are physically identical.
  • A replacement for domain expert review in high-liability contexts.
  • A finished standard — it is a methods proposal, open to challenge and calibration.
The vocabulary

Raw cHP

Drafting speed — the rate of fluent candidate output, before any verification.

Verified cHP

Output that passes evidence, factuality, and semantic-preservation checks.

Constrained cHP

Output completed under a defined policy, risk, provenance, and approval profile.

Review debt

The time, attention, and risk absorbed verifying probabilistic outputs before release.

Semantic drift

A paraphrase that quietly changes scope, numbers, risk balance, or claim from the source.

Cognitive Work

Source information transformed into lower-load, higher-structure output — the work cHP measures.

As agents mediate how people discover information and act on it, value migrates to the layer between the person, the knowledge, and the action. Generation is now commoditized — every model writes. What does not commoditize is verifiable, compliant, decision-grade signal, and it grows more valuable as agents take on more autonomous action under regulators that have created no AI exception.

Fluent language is abundant. Trustworthy, evidence-bound, policy-safe meaning remains scarce.

That scarcity is where the next generation of AI infrastructure will be built. The cHP paper is published as a methods proposal — meant to be challenged, calibrated, and tested against real task batteries: claim extraction, citation mapping, risk-balanced summaries, FAQ drafting, social adaptation, and drift detection.

Sources

[1] Becker, J., Rush, N., Barnes, B., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. arXiv:2507.09089.

[2] Ranganathan, A., & Ye, X. M. (2026). AI Doesn't Reduce Work—It Intensifies It. Harvard Business Review.

[3] Stanford RegLab / Stanford HAI legal-domain hallucination study (2024); OpenAI PersonQA benchmark, o3 / o4-mini (2025); Vectara Hallucination Leaderboard (grounded summarization).

Working paper

Cognitive Horsepower: A Working Metric for Constraint-Adjusted Human-Agent Labor in the Inference Economy

Download the Signal Fidelity Group preprint below. (Currently under SSRN peer review.)

Basu, A. (2026). Cognitive Horsepower: A Working Metric for Constraint-Adjusted Human-Agent Labor in the Inference Economy. Signal Fidelity Group. Preprint.

Download PDF

Frequently asked

What is Cognitive Horsepower (cHP)?

Cognitive Horsepower is a proposed operational metric for measuring constraint-adjusted human-agent labor — the rate at which a human, an AI agent, or a human-agent team produces verified, policy-safe, evidence-bound work, not just fluent output.

What is review debt?

Review debt is the accumulated time, attention, and risk absorbed by experts who must verify probabilistic AI outputs before they can be released. As AI lowers drafting cost, verification — not generation — becomes the bottleneck.

How is cHP different from counting tokens or model benchmarks?

Token counts and benchmarks measure generation and general capability. cHP measures what survives review under a defined policy, evidence, and liability profile — cost per approved output rather than cost per token.

What is the difference between raw, verified, and constrained cHP?

Raw cHP is drafting speed. Verified cHP is the rate of output that passes evidence and quality checks. Constrained cHP is the rate of output completed under a defined policy, risk, provenance, audit, and approval profile — the number that matters for regulated workflows.

Is cHP a finished industry standard?

No. cHP is a proposed measurement framework and methods proposal, intended to be challenged, calibrated, and validated against task batteries — not a validated universal standard.

Signal Fidelity Group

We work with organizations that need their meaning to arrive intact.

Start a conversation