AI is that employee who sends emails at midnight
The problem isn’t the worker. It’s the manager who mistakes activity for value.
We’ve all had a colleague who sends late night emails.
They are always available, and always responsive. Their green dot on Slack never disappears. They give you thirty options when you only needed three. Their calendar is full, but they keep working after everyone else has stopped.
Are they good at their job?
Perhaps.
They might be exceptional. Their force of will might be the only thing keeping an organization together. Carrying impossible work that their manager doesn’t even understand, and their colleagues never see.
They might be completely overwhelmed. Inefficient. Anxious. They might be producing a stream of work that makes more work for all the rest of us.
And the midnight email doesn’t help us know which one.
We know that manager who mistakes dedication for performance.
They see the hours at your desk. They look at the messages you send, the meetings you attend. The tickets closed and the lines of code written. They count the documents you produce.
Did that message change an outcome? Was that outcome worth the effort?
Now we have AI. We’ve built the employee who sends emails at midnight.
We haven’t changed its manager.
It’s permanently available. It never gets tired. It responds immediately, every time. It can outmatch any human worker in its capacity to generate reports, plans, code, images, summaries, strategies - every document or artifact you could ever want.
If we’re measuring by presence, effort, and production then AI is the greatest employee who ever lived.
That’s probably not helpful.
The things a manager sees
When I was leading the Carbon Design System at IBM, I brought adoption targets to Phil Gilbert, the GM of IBM Design. He told me to stop focusing on adoption. It wasn’t the right measure of success.
Adoption was interesting. It wasn’t inherently valuable.
That conversation was years before we started asking the same question about a machine.
There are six different questions we can ask about work.
Presence: Was someone at their desk?
Effort: Did they put in the work?
Production: Did their work create outputs?
Performance: Were those outputs any good?
Outcome: Did they change anything?
Value: Was the change worth it?
Good management, good measurement, will look at all six but judge mostly through the final three.
Presence and effort matter. We like people who show up. And we usually want work to produce something.
Those things are necessary. But they don’t demonstrate success.
Performance is an assessment of the quality of the effort, not just the volume. Outcome lets us ask what changed in the world because the work was done. Value asks whether that change was worth the cost and risk.
Those are harder measures. At a minimum, they need someone to define what “good” means.
Bad management - maybe even average management - reverses that importance.
The employee who stays late looks more committed than the employee who finished the important work and left at 4pm. The person who sent you a fifty-page document looks more substantial than the one who sent you a two-line email that identified the decision that mattered. One team closes a hundred tickets, and looks productive. Another team prevents the creation of a hundred tickets, and looks quiet.
And it’s not just managerial foolishness. It’s not malice.
We understand bums in seats. We understand people working late. We understand the production of stuff. It’s all trackable. It all shows willing.
Value emerges more slowly. There are dependencies. More people are involved. Value might be stopping some work from ever happening. It needs an understanding of quality.
And we’re not great at that nuance. So organizations reward the visible. Employees learn to be visible.
If you’ve sent an email after 6pm when you could have sent it before 6pm, you know what I mean.
That’s about showing that you’re working. Not about showing that the work matters.
The proxies used to mean something
We valued presence, effort, and production because they’re connected with human limits.
You have to be there. You get tired. You spend your own time making stuff.
If you’re a person who made twenty substantial reports, beautifully presented as PDFs, you’ve done a bunch of work. That doesn’t mean the reports are good, but shows you’re invested.
Stuff that takes time, that takes effort, has scarcity.
Until AI came along.
AI can be there every hour of the day. Its effort doesn’t come with fatigue. It can make a lengthy report while you’re making your coffee.
Human limits secured that proxy.
Now? Delivery is decoupled from investment.
AI isn’t padding its hours to optimize a number. It genuinely is always available, always responsive. It’s making a lot of stuff. Nothing is being faked. The activity is real, and it’s honest.
But whether it’s valuable...that’s another question.
Here’s the kind of evidence we’ve seen for AI success:
The system generated ten thousand lines of code.
The system answered five hundred questions.
The system created forty campaign concepts.
The system summarized every customer interview.
These might all be true.
Was the code reliable, maintainable, or necessary?
Were the answers useful? Correct? Did anyone act on them?
Were the forty concepts any better than the three that the team developed without that tool?
Were the summaries accurate? Did they reflect what was really important to users?
They aren’t, by themselves, value claims. They’re diagnostics.
Is the machine operating?
The industry has automated its weakest managerial instincts and called it “observability”.
Production is not productivity
Generative AI is seductive. It makes stuff.
Blank page to full page. Empty backlog to populated. Busy inbox to drafted replies.
It did something. And we, as humans, like something more than nothing.
And we’re still adjusting to the sense of movement.
Before we had generative AI, making that report was - at a minimum - demonstration of effort: research, organization, and writing. The size of the artifact was evidence that we’d worked hard.
AI completely breaks that relationship.
That long report might come from a one-line prompt. It might never provide value.
The artifact, when it had human cost, looked significant. AI strips that cost. But as humans we still respond to the length, fluency, and completeness. We want the work to be consequential.
The AI stayed late. Look at everything it made.
AI gives you execution speed, no question. The correction tax is an efficiency cost.
Who’s going to read the report? Which developer will be reviewing the code? Am I going to have to choose between the fifty options it gave me? I’ll need to check the citations. I’ll go through the error logs to find the gaps in a convincing answer.
That isn’t a productivity increase. It’s an increase in paperwork.
We don’t demo the cost part of AI, just the speed.
The work management avoided defining
I wrote recently about auditing the AI skills I had built for my own work and life. The audit showed me that measuring the machine was just the beginning.
We only know if the AI-generated product brief is good if we actually know what a good product brief is. To evaluate a coding agent, we need to decide if we care about delivery time and defects, or lines written.
We can’t rely on “I’ll know if it’s good if I see it.” That’s a managerial dodge.
And it’s not about one universal definition of quality.
We need stated definitions for success, quality, and value, in the context the work is delivered. And those definitions should be owned by somebody willing to defend them.
Organizations have left this vague for years. Effort is reassuring. Good managers - able to apply their contextual understanding - filled gaps with their own judgment on impact.
AI makes those gaps harder to ignore.
AI can generate more activity than we can inspect. Faster to production also means faster to shipping mistakes.
If production is unlimited, what does production tell us?
This is why evals matter.
It’s not about reducing every aspect of work to a score. Quantitative measurement doesn’t remove judgment.
Who receives the benefit, and who absorbs the cost? A serious eval makes those kinds of judgments explicit.
Evals shouldn’t just be tests for the AI systems.
They’re tests of whether management knows what it is asking the system to accomplish.
The performance review
AI looks amazing under the weakest measures of productivity.
It is always at its desk. It works with astonishing speed. It’s never tired. It produces more than anyone else.
Perhaps it is also doing excellent work. Perhaps it is changing outcomes that matter and creating value far beyond its cost.
But the first three things don’t prove the final three.
The midnight email never proved value.
Evals move the performance review beyond production.
It means defining what good looks like. Being able to sift through volume for value. Not simply admiring the output.
AI worked through the night. We already know that about AI.
We need to know whether the work matters.
Further reading:
Design system adoption numbers...just a vanity metric? - the same argument, before AI: adoption is interesting, not valuable.
I know kung fu. I might remember how to throw one punch. - the skills audit referenced above.
Collins, S. We Hired AI to Work Less. Instead, Our Workload Jumped 346%. Here’s Why. Activated Thinker, Mar 2026.
Schawbel, D. How AI turned your best work into the bare minimum. Fortune, Aug 2026.
