Tokens Aren't Output. Features Out Are. article visual
Orrbit Technologies← All insights
AI & Strategy

Tokens Aren't Output. Features Out Are.

Hours, lines, commits, and token burn are activity. The only adult scoreboard is what shipped, held, and can be restored.

8/14/2026 12 min read

There is a cartoon going around that ranks engineers by tokens spent. Everyone laughs. Then someone puts token burn on a dashboard next to "productivity" and the joke gets a budget.

We have been here. Lines of code rewarded the person who would not delete. Commits rewarded the person who sliced a thought into confetti. Hours rewarded the person who stayed visible. Tokens are the same sin in a new unit: they measure how hard the machine was breathing, not whether a customer got a feature that survived the week.

Token throughput is a real number. It belongs next to cloud cost, not next to someone's worth. High burn often means thrashing — the agent and the human circling a wrong abstraction while the bill climbs. TechCrunch's tokenmaxxing season made the pattern legible: more pull requests, more spend, not a matching leap in value. Volume is what generators are for. Value is what you still have to choose.

Commits merged are closer to the truth and still a trap. A merged PR can be a feature, a revert, a ritual, or a thousand-line shrug that will be rewritten in nine days. Churn is the tell. If AI adoption made the codebase churn like a washing machine, you did not get faster. You got a blender. Durability — what is still there at 30, 60, 90 days — is a better adult.

Measure features out. Not story points dressed as features. Not "the agent finished the ticket." A feature out is in production, behind the flag you can actually turn off, observed, owned. Pair it with the unfashionable DORA cousins: how often you ship, how long it takes a change to matter, how often a change needs a fire, how fast you restore. In an agent era, change-failure and restore time matter more, because you can now ship a mistake at generator speed.

Effort still exists. Calibrate it. How much human attention did this take to steer and verify? A two-day feature that took twenty minutes of judgment and a lot of tokens can be excellent. A two-hour feature that nobody understands is a liability with a green build. Attribute the agent's share so you do not confuse leverage with luck.

Do this at team level. Ranking individuals on tokens or diffs is how you train contractors of appearance. The scoreboard that does not rot: what left the building, what held, what you could explain, what you could undo. Everything else is telemetry. Use it when a goal is missed. Do not let it become the goal.

Finance will still want a graph. Give them cost per shipped outcome, not cost per million tokens as a personality. Engineering will still want a pulse. Give them durability and restore time, not a leaderboard of who typed the loudest prompt. Product will still want a date. Give them a feature that is out, observed, and owned — or admit it is not out.

If a metric can be gamed by generating more, it will be. Tokens, lines, and commits can all be gamed by a model that never gets tired. Features that survive contact with users are harder to fake. That is the point.

We do not measure the hours. We do not measure the lines. We measure the features you have out — and whether they still deserve to be.

Written by

Brandon Nkawu