I know kung fu. I might remember how to throw one punch.
I'm really, really good at AI. Right? I audited myself to find out.
This March I built another Claude skill suite - ten commands for helping to manage my own fitness. There was a daily check-in, a biweekly retrospective, and so on. A couple of hours work, designed to be used every single day.
Last week I audited my entire toolkit of AI skill suites. Telemetry had its own verdict.
I didn’t use the fitness suite. At all.
I’ll come back to that verdict.
In The Matrix, Neo jacks in, downloads a file, his eyes open: “I know kung fu.”
Using my deep research skill to define a strategic plan for a topic, and turning that into an AI skill suite, is the closest thing I’ve come to feeling like that for real. Research goes in, a tool comes out.
Capability acquired.
I’ve had that feeling multiple times. I have twenty-two domains for work and life - 193 skills in total. Product specs, competitive analysis, balcony gardening, wardrobe management.
Do I know kung fu? Or do I just have a big folder of downloads?
I’m a bad person to answer my own question
In 2025 METR ran a randomized trial. They gave experienced developers frontier AI tools for their own codebases, and measured what happened. The developers believed the AI tooling made them 20% faster.
They were 19% slower.
METR ran a follow-up in late 2025. The slowdown headline went away - newer tools, more practiced people. More interestingly, their measurement broke. Developers wouldn’t even submit tasks they’d have to do without AI. And time-tracking was unreliable when they applied it to agentic multitasking.
The people who measure this as their job concluded that their own results weren’t a good “proxy for the real productivity impact.”
So the durable finding isn’t the 19%. It’s the gap between what we feel and what we can verify - and that gap is getting harder to close, not easier.
This isn’t about developers, so much as it’s about our own testimony. Fluency feels fast. My tooling makes me feel capable, which makes me less reliable. With nearly two hundred tools, I’ve constructed myself into the least reliable witness possible.
So I decided to see if I could pull the data.
Invocation - what actually ran, when, and where.
Artifact trail - whether the tools produced anything, and whether that went anywhere further.
Research lineage - what was fed downstream.
And, fair to say, the telemetry itself isn’t perfect. It’s definitely missed skills that I know I use. Capture is hard.
But the numbers are more honest than my own feelings...my own vibes.
What did the audit say?
In the window I measured, 23% of the skills I’ve built were used. And only twelve skills carry half of all the activity.
That sounds terrible! 193 skills built, and I only really use twelve of them. What a failure!
But I don’t read it that way.
That’s not failure. It’s portfolio selection.
The marginal cost of building the skills with AI was tiny. And when the cost is tiny, the rational strategy is “build a bunch of stuff, and then let reality select.”
While the cost of building might be negligible, the cost of not deciding what to do with it is not.
My sin wasn’t building too much. It was that it took me this long to run the curation.
Living, dead, and somewhere in-between
When I look across my whole portfolio, I can break it down into six different states.
Regular use. My daily workhorses. A dozen skills doing half the work - meeting debriefs, deep research, spec writing. No notes, these are clearly things I actively reach for.
Retired, with honors. My first delivery suite - twenty commands - I used heavily for months. But it’s silent because I built a successor and ran an A/B between them. Three of six predictions were wrong - including one I’d filed explicitly as “what would particularly surprise me.” The new suite won, the old one retired. That’s a system that works.
Obsoleted by drift. I had sprint-cycle skills. We moved from cycles to Kanban. The tools weren’t wrong, they just didn’t apply any more.
Dormant. I have some skills around tax preparation. They’ve done nothing all summer. That’s correct - they shouldn’t be doing anything in summer.
Graduated. An interesting one. My wardrobe and style suite logged 46 outfits, cataloged a 222-item inventory, and scored what worked. It went quiet in June. But I think that’s because I’ve taken on those behaviors myself. The tool taught the pattern and then became unnecessary. I need to validate - I’m suspicious of that conclusion - but it’s a case where the kung fu downloaded into me, and didn’t stay in the file.
Presumed dead. The fitness suite I mentioned earlier. Ten commands for regular use, zero recorded invocations.
...except when I went to look back, even the fitness suite’s logbook told a different story. There are dated entries over several months. Including some that are inside my measurement window. So for some reason the telemetry was blind to it.
The audit result then suggests that there are zero confirmed “dead” skill suites. Silence by itself isn’t enough to declare them dead.
That telemetry issue needs investigating. Measurement brings discipline to judgment. It doesn’t replace it. Apparently a death certificate needs a second witness.
Research - my busiest, most unread, most vital skill
By volume, my deep research skill is the most productive by a lot. It’s made hundreds of artifacts over the last few months.
I almost never reread them.
But that’s the wrong model. If I use a skill to generate a competitive battlecard, its value is clear: someone opens it and uses it.
If it’s a research brief, its value comes from being transformed into something else. After that, it might never be opened again. The value moves forward.
I ran a trace on my ninety research runs. Eighty percent of them led to something real. Thirty-three of them became skills or skill suites themselves. Some of my most-used skills come from the research run that designed them. Others fed identifiable decisions. And some of them led nowhere - like fully drafted skills I never moved to install. No entire suite was confirmed dead. Plenty of individual work was.
The 20% dead rate helps me believe the 80%.
We have to learn to say goodbye
This goes beyond just skills creation.
Build-on-demand works. Twice this summer a real need appeared and I could implement a working Claude skill suite in days, one that immediately became integral.
For personal tooling like this, production isn’t the difficult bit.
But what’s missing - for my practice, but really in every AI-use conversation I see - is the other half of the loop.
Scheduled, honest selection. Kill criteria written at build time, before we get emotionally attached. A quarterly pass where every tool gets a verdict: active, dormant, retired, obsolete - or given a funeral if it’s actually dead.
And I still can’t empirically claim that this makes me faster or better than I’d be if I didn’t have it. I’m not sure anyone can claim that about themselves based on feelings alone. That’s the thirty-nine-percentage-point gap between perception and measured performance in the METR study.
Designing the experiment to measure that is what comes next.
So...I know kung fu. I have the logs to prove it.
But the only kung fu that’s mine is the punch I can still throw after the file closes.
Right now, that’s just one punch. I’ve counted it. Twice.
Further reading:
Becker, J, et al. We are Changing our Developer Productivity Experiment Design. METR, Feb 2026.
New data: AI’s impact on engineering velocity is more modest than expected. DX, April 2026.
Counts, L. AI promised to free up workers’ time. UC Berkeley Haas researchers found the opposite. UC Berkeley Haas, Feb 2026.
Article photo by Getty Images on Unsplash.
