<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Robin Cannon: Field Notes]]></title><description><![CDATA[Professional writing from the edges of product, design, and digital systems. Drawing on my role as VP of Product at Knapsack, and years leading design systems and product strategy at IBM and J.P. Morgan — the patterns, decisions, and dynamics that don't fit neatly into case studies. Systems thinking, strategy, and leadership from inside the work.]]></description><link>https://www.robin-cannon.com/s/field-notes</link><image><url>https://substackcdn.com/image/fetch/$s_!maYW!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2c62c87-7ba3-444c-ad20-4a4cf617a8f7_1024x1024.png</url><title>Robin Cannon: Field Notes</title><link>https://www.robin-cannon.com/s/field-notes</link></image><generator>Substack</generator><lastBuildDate>Mon, 24 Aug 2026 14:46:09 GMT</lastBuildDate><atom:link href="https://www.robin-cannon.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Robin Cannon]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[shinytoyrobots@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[shinytoyrobots@substack.com]]></itunes:email><itunes:name><![CDATA[Robin Cannon]]></itunes:name></itunes:owner><itunes:author><![CDATA[Robin Cannon]]></itunes:author><googleplay:owner><![CDATA[shinytoyrobots@substack.com]]></googleplay:owner><googleplay:email><![CDATA[shinytoyrobots@substack.com]]></googleplay:email><googleplay:author><![CDATA[Robin Cannon]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[AI is that employee who sends emails at midnight]]></title><description><![CDATA[The problem isn&#8217;t the worker. It&#8217;s the manager who mistakes activity for value.]]></description><link>https://www.robin-cannon.com/p/ai-is-that-employee-who-sends-emails</link><guid isPermaLink="false">https://www.robin-cannon.com/p/ai-is-that-employee-who-sends-emails</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 18 Aug 2026 15:02:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/95843b06-258f-41bb-884a-fa0cc1533aef_4997x1877.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We&#8217;ve all had a colleague who sends late night emails.</p><p>They are always available, and always responsive. Their green dot on Slack never disappears. They give you thirty options when you only needed three. Their calendar is full, but they keep working after everyone else has stopped.</p><p>Are they good at their job?</p><p>Perhaps.</p><p>They might be exceptional. Their force of will might be the only thing keeping an organization together. Carrying impossible work that their manager doesn&#8217;t even understand, and their colleagues never see.</p><p>They might be completely overwhelmed. Inefficient. Anxious. They might be producing a stream of work that makes more work for all the rest of us.</p><p>And the midnight email doesn&#8217;t help us know which one.</p><p>We know that manager who mistakes dedication for performance.</p><p>They see the hours at your desk. They look at the messages you send, the meetings you attend. The tickets closed and the lines of code written. They count the documents you produce.</p><p>Did that message change an outcome? Was that outcome worth the effort?</p><p>Now we have AI. We&#8217;ve built the employee who sends emails at midnight.</p><p>We haven&#8217;t changed its manager.</p><p>It&#8217;s permanently available. It never gets tired. It responds immediately, every time. It can outmatch any human worker in its capacity to generate reports, plans, code, images, summaries, strategies - every document or artifact you could ever want.</p><p>If we&#8217;re measuring by presence, effort, and production then AI is the greatest employee who ever lived.</p><p>That&#8217;s probably not helpful.</p><h3>The things a manager sees</h3><p>When I was leading the Carbon Design System at IBM, I brought adoption targets to Phil Gilbert, the GM of IBM Design. He told me to stop focusing on adoption. It wasn&#8217;t the right measure of success.</p><p>Adoption was interesting. It wasn&#8217;t inherently valuable.</p><p>That conversation was years before we started asking the same question about a machine.</p><p>There are six different questions we can ask about work.</p><p><strong>Presence:</strong> Was someone at their desk?</p><p><strong>Effort:</strong> Did they put in the work?</p><p><strong>Production:</strong> Did their work create outputs?</p><p><strong>Performance:</strong> Were those outputs any good?</p><p><strong>Outcome:</strong> Did they change anything?</p><p><strong>Value:</strong> Was the change worth it?</p><p>Good management, good measurement, will look at all six but judge mostly through the final three.</p><p>Presence and effort matter. We like people who show up. And we usually want work to produce something.</p><p>Those things are necessary. But they don&#8217;t demonstrate success.</p><p>Performance is an assessment of the quality of the effort, not just the volume. Outcome lets us ask what changed in the world because the work was done. Value asks whether that change was worth the cost and risk.</p><p>Those are harder measures. At a minimum, they need someone to define what &#8220;good&#8221; means.</p><p>Bad management - maybe even average management - reverses that importance.</p><p>The employee who stays late looks more committed than the employee who finished the important work and left at 4pm. The person who sent you a fifty-page document looks more substantial than the one who sent you a two-line email that identified the decision that mattered. One team closes a hundred tickets, and looks productive. Another team prevents the creation of a hundred tickets, and looks quiet.</p><p>And it&#8217;s not just managerial foolishness. It&#8217;s not malice.</p><p>We understand bums in seats. We understand people working late. We understand the production of stuff. It&#8217;s all trackable. It all shows willing.</p><p>Value emerges more slowly. There are dependencies. More people are involved. Value might be stopping some work from ever happening. It needs an understanding of quality.</p><p>And we&#8217;re not great at that nuance. So organizations reward the visible. Employees learn to be visible.</p><p>If you&#8217;ve sent an email after 6pm when you could have sent it before 6pm, you know what I mean.</p><p>That&#8217;s about showing that you&#8217;re working. Not about showing that the work matters.</p><h3>The proxies used to mean something</h3><p>We valued presence, effort, and production because they&#8217;re connected with human limits.</p><p>You have to be there. You get tired. You spend your own time making stuff.</p><p>If you&#8217;re a person who made twenty substantial reports, beautifully presented as PDFs, you&#8217;ve done a bunch of work. That doesn&#8217;t mean the reports are good, but shows you&#8217;re invested.</p><p>Stuff that takes time, that takes effort, has scarcity.</p><p>Until AI came along.</p><p>AI can be there every hour of the day. Its effort doesn&#8217;t come with fatigue. It can make a lengthy report while you&#8217;re making your coffee.</p><p>Human limits secured that proxy.</p><p>Now? Delivery is decoupled from investment.</p><p>AI isn&#8217;t padding its hours to optimize a number. It genuinely is always available, always responsive. It&#8217;s making a lot of stuff. Nothing is being faked. The activity is real, and it&#8217;s honest.</p><p>But whether it&#8217;s valuable...that&#8217;s another question.</p><p>Here&#8217;s the kind of evidence we&#8217;ve seen for AI success:</p><ul><li><p>The system generated ten thousand lines of code.</p></li><li><p>The system answered five hundred questions.</p></li><li><p>The system created forty campaign concepts.</p></li><li><p>The system summarized every customer interview.</p></li></ul><p>These might all be true.</p><ul><li><p>Was the code reliable, maintainable, or necessary?</p></li><li><p>Were the answers useful? Correct? Did anyone act on them?</p></li><li><p>Were the forty concepts any better than the three that the team developed without that tool?</p></li><li><p>Were the summaries accurate? Did they reflect what was really important to users?</p></li></ul><p>They aren&#8217;t, by themselves, value claims. They&#8217;re diagnostics.</p><p>Is the machine operating?</p><p>The industry has automated its weakest managerial instincts and called it &#8220;observability&#8221;.</p><h3>Production is not productivity</h3><p>Generative AI is seductive. It makes stuff.</p><p>Blank page to full page. Empty backlog to populated. Busy inbox to drafted replies.</p><p>It did something. And we, as humans, like something more than nothing.</p><p>And we&#8217;re still adjusting to the sense of movement.</p><p>Before we had generative AI, making that report was - at a minimum - demonstration of effort: research, organization, and writing. The size of the artifact was evidence that we&#8217;d worked hard.</p><p>AI completely breaks that relationship.</p><p>That long report might come from a one-line prompt. It might never provide value.</p><p>The artifact, when it had human cost, looked significant. AI strips that cost. But as humans we still respond to the length, fluency, and completeness. We want the work to be consequential.</p><p>The AI stayed late. Look at everything it made.</p><p>AI gives you execution speed, no question. The correction tax is an efficiency cost.</p><p>Who&#8217;s going to read the report? Which developer will be reviewing the code? Am I going to have to choose between the fifty options it gave me? I&#8217;ll need to check the citations. I&#8217;ll go through the error logs to find the gaps in a convincing answer.</p><p>That isn&#8217;t a productivity increase. It&#8217;s an increase in paperwork.</p><p>We don&#8217;t demo the cost part of AI, just the speed.</p><h3>The work management avoided defining</h3><p>I wrote recently about auditing the AI skills I had built for my own work and life. The audit showed me that measuring the machine was just the beginning.</p><p>We only know if the AI-generated product brief is good if we actually know what a good product brief is. To evaluate a coding agent, we need to decide if we care about delivery time and defects, or lines written.</p><p>We can&#8217;t rely on &#8220;I&#8217;ll know if it&#8217;s good if I see it.&#8221; That&#8217;s a managerial dodge.</p><p>And it&#8217;s not about one universal definition of quality.</p><p>We need stated definitions for success, quality, and value, in the context the work is delivered. And those definitions should be owned by somebody willing to defend them.</p><p>Organizations have left this vague for years. Effort is reassuring. Good managers - able to apply their contextual understanding - filled gaps with their own judgment on impact.</p><p>AI makes those gaps harder to ignore.</p><p>AI can generate more activity than we can inspect. Faster to production also means faster to shipping mistakes.</p><p>If production is unlimited, what does production tell us?</p><p>This is why evals matter.</p><p>It&#8217;s not about reducing every aspect of work to a score. Quantitative measurement doesn&#8217;t remove judgment.</p><p>Who receives the benefit, and who absorbs the cost? A serious eval makes those kinds of judgments explicit.</p><p>Evals shouldn&#8217;t just be tests for the AI systems.</p><p>They&#8217;re tests of whether management knows what it is asking the system to accomplish.</p><h3>The performance review</h3><p>AI looks amazing under the weakest measures of productivity.</p><p>It is always at its desk. It works with astonishing speed. It&#8217;s never tired. It produces more than anyone else.</p><p>Perhaps it is also doing excellent work. Perhaps it is changing outcomes that matter and creating value far beyond its cost.</p><p>But the first three things don&#8217;t prove the final three.</p><p>The midnight email never proved value.</p><p>Evals move the performance review beyond production.</p><p>It means defining what good looks like. Being able to sift through volume for value. Not simply admiring the output.</p><p>AI worked through the night. We already know that about AI.</p><p>We need to know whether the work matters.</p><div><hr></div><h4>Further reading:</h4><ul><li><p><em><a href="https://www.robin-cannon.com/p/design-system-adoption-numbersjust">Design system adoption numbers...just a vanity metric?</a> </em>- the same argument, before AI: adoption is interesting, not valuable.</p></li><li><p><em><a href="https://www.robin-cannon.com/p/i-know-kung-fu-i-might-remember-how">I know kung fu. I might remember how to throw one punch.</a></em> - the skills audit referenced above.</p></li><li><p>Collins, S. <em><a href="https://medium.com/activated-thinker/we-hired-ai-to-work-less-instead-our-workload-jumped-346-heres-why-3963ebf4d412">We Hired AI to Work Less. Instead, Our Workload Jumped 346%. Here&#8217;s Why.</a> </em>Activated Thinker, Mar 2026.</p></li><li><p>Schawbel, D. <em><a href="https://fortune.com/2026/08/04/your-best-work-ai-bare-minimum-no-free-time/">How AI turned your best work into the bare minimum.</a> </em>Fortune, Aug 2026.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@brunobd?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Bruno BD</a> on <a href="https://unsplash.com/photos/office-building-windows-at-night-with-workers-inside-wSxag1zUalw?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI governance stack only answers half the problem]]></title><description><![CDATA[You know who asked the AI. You don't know if the AI gave the right answer.]]></description><link>https://www.robin-cannon.com/p/your-ai-governance-stack-only-answers</link><guid isPermaLink="false">https://www.robin-cannon.com/p/your-ai-governance-stack-only-answers</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 11 Aug 2026 15:01:18 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c94abfa2-844a-4b0a-928f-5b781499f3f5_7547x5034.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Enterprise AI governance is getting good at answering a question.</p><blockquote><p><em>Should this person be allowed to generate that?</em></p></blockquote><p>That&#8217;s identity management and role-based access controls. Audit logs. Data-retention policies. Spend limits (now that we&#8217;re realizing tokens aren&#8217;t free!). Permissions for tools and connectors. And controls over the models available, what information they can reach, and what gets recorded.</p><p>This is good. Sensible and secure.</p><p>It&#8217;s also only half the problem.</p><p>There&#8217;s another question that gets missed.</p><blockquote><p><em>Is what gets generated acceptable?</em></p></blockquote><p>We&#8217;re still talking about governance. But the first question was about access. The second is about conformance.</p><p>Enterprise AI stacks are much better at the first one than the second.</p><div><hr></div><p>Anthropic has built in SSO, SCIM, role-based permissions, audit and compliance tools, data-retention controls, observability, spend management, and controls to configure how models connect to outside systems. Microsoft has a growing governance layer around Copilot. Okta is putting work into expanding identity governance so it captures agents as well as humans.</p><p>Enterprise companies know how to solve these problems.</p><p>I&#8217;ve worked at IBM and at J.P. Morgan. These are the types of companies who need to determine who someone is, what they can access, and what they did.</p><p>So it&#8217;s natural that they apply the same methods to generative AI.</p><p>Who invoked that model? Were they allowed to? What data and tools could it use? How much did it cost?</p><p>And these could all work perfectly.</p><p>The authorized employee signs into the approved AI product at their company. They have all the right permissions. The model is locked down - only the information they&#8217;re allowed to see. The interaction is logged, there are traces. Nothing leaks, no policy is breached.</p><p>The AI makes some kind of plausible artifact.</p><p>And it&#8217;s wrong.</p><p>Maybe not obviously. But wrong according to all of the organization&#8217;s own standards.</p><p>It violates accessibility. It ignores a code convention. It&#8217;s in conflict with an existing product decision. It invents a component that&#8217;s already in your design system.</p><p>Nobody catches this. It looks right, and it passed all the governance controls.</p><p>It ships.</p><p>The governance stack only solved access. It protected the company from misuse of AI.</p><p>Conformance governance would protect it from AI that makes the wrong thing.</p><p>We&#8217;re building the first one faster than the second.</p><h3>Not my first governance rodeo</h3><p>Before AI (and after), I&#8217;ve spent years working on design systems.</p><p>Design systems always had a governance problem - more than they had a component problem.</p><p>The component library is the simpler part. The hard questions come when hundreds or thousands of people start using it.</p><p>Introduce standards and training. Establish review procedures and processes.</p><p>And, if you&#8217;re not careful, you become the design systems police.</p><p>I wrote <a href="https://www.robin-cannon.com/p/dont-become-the-design-systems-police?utm_source=internal">about that problem four years ago</a>.</p><p>That central team discovers their job isn&#8217;t helping people make good decisions. It&#8217;s turned into catching people who make bad ones. It&#8217;s not governance, it&#8217;s enforcement. The system is a gate.</p><p>That wasn&#8217;t scalable when the people doing the work were all human.</p><p>AI&#8217;s made the same problem much larger.</p><p>A designer can create five credible options in the time it used to take them to make one. Now there are five times as many things that need checking.</p><p>Same on the engineering side. There&#8217;s much more code to review.</p><p>Did the ten pages of analysis that the product manager made this morning align with what the company knows?</p><p>The smaller the generation problem, the bigger the output problem.</p><p>And we&#8217;re only just starting to pay attention to that part of the productivity story.</p><p>We&#8217;re all striving for velocity. For me, that means a combination of speed and quality.</p><p>AI can reduce the cost of making an artifact from four hours to four minutes.</p><p>Then we have to spend thirty minutes reviewing that artifact. Another thirty to identify all the deviations. An hour of corrective work. Coordination with other people who also need to review.</p><p>The work&#8217;s moved, not disappeared.</p><p>AI without conformance means the generation cost is converted into correction cost.</p><p>And right now more of the correction remains human. So the faster generation makes things even worse.</p><p>I wrote earlier this year that <a href="https://www.robin-cannon.com/p/execution-is-cheap-coordination-is?utm_source=internal">execution is becoming cheap while coordination is not.</a></p><p>This is one of the consequences.</p><p>The bottleneck moves downstream.</p><h3>System discovery isn&#8217;t the problem</h3><p>In early evals work at Knapsack, we started to see a clear distinction.</p><p>A controlled evaluation. Test a coding agent doing design system tasks. A matched task set, no customer data. Give it two conditions. One where it has an MCP connection that provides guidance for that system. One where it doesn&#8217;t.</p><p>The MCP doesn&#8217;t help the agent find the system. It does that successfully, assuming it has access. It even imports the components at about the same rate.</p><p>The agent knows to look in <code>node_modules</code>.</p><p>But if it doesn&#8217;t have the MCP serving the context, it doesn&#8217;t know how to use that system.</p><p>Give it that context, and design fidelity improves 5-10%. Code quality improves by more than 20%.</p><p>Prompts that produced shippable code rose from 40% to over 60%.</p><p>TypeScript errors down significantly. Prop violations similarly reduced.</p><p>The cost of running any individual task was higher. But the cost per <em>shippable</em> output was thirty percent cheaper.</p><p>This agent didn&#8217;t have new access. It already had the components.</p><p>But now it had context about how the organization expected the components to be used. Props, variants, compositions, and the constraints around them.</p><p>The access problem was already solved. The conformance problem wasn&#8217;t.</p><p>Conformance improves the economics of usage. Even with our early, simple context provision the benefits are clear and measurable.</p><p>The organization&#8217;s intent becomes a participant in the generation.</p><h3>Plausible mistakes compound</h3><p>The more complex the task, the more the agents benefit from additional context.</p><p>On a multi-screen flow, attaching the MCP gave us gains that were nearly double what we saw on simple patterns.</p><p>This isn&#8217;t a neat scaling. We don&#8217;t have enough data. On some simpler template-based tasks, the additional context sees the agent second guess itself and reduces its reuse of the right components.</p><p>But high-complexity tasks, broadly, benefit the most.</p><p>As tasks get bigger, they involve more decisions. And multiple small plausible but wrong decisions start to compound into something globally wrong.</p><p>I always talked about it when it came to design systems. You can take a bunch of completely accessible components and assemble them into a really inaccessible experience.</p><p>A complete flow can violate how your company believes the experience should work.</p><h3>Conformance is broader than compliance</h3><p>&#8220;Conformance&#8221; isn&#8217;t just another word for regulatory compliance.</p><p>Conformance is a broader evaluation of whether an artifact satisfies the constraints your organization has already decided.</p><p>Yes, be compliant with accessibility. Dealing with private information. Legal restrictions your company needs.</p><p>But some of the constraints and context are just how the organization has chosen to operate.</p><p>Brand standards.</p><p>Design-system conventions.</p><p>Product strategy.</p><p>Architectural decisions.</p><p>Content guidelines.</p><p>Approved patterns.</p><p>These are things that teams throughout your company have learned painfully before. And they&#8217;d prefer not to learn them painfully again.</p><p>Any organization is full of these decisions.</p><p>But they&#8217;re scattered.</p><p>Design system sites. Wikis. Policy documents. An old presentation. Architecture records. Jira tickets. Multiple Slack threads. That developer&#8217;s memory. The shared knowledge of the people who were in the room when some decision was made.</p><p>AI isn&#8217;t magically better than humans at pulling all this context together when it&#8217;s time to execute.</p><p>Given incomplete organizational context, it&#8217;s probably worse. It&#8217;ll make more guesses, ask fewer questions, and generate something plausible from the context it does have.</p><p>Which is what we asked it to do.</p><p>And generic AI evaluation only gets us to a certain point.</p><p>A benchmark can tell me if Opus 5 is generally good at coding.</p><p>An eval can tell me if GPT 5.6 tends to complete some specific task successfully.</p><p>It doesn&#8217;t tell me if the code it&#8217;s producing is what IBM wants to ship. Or if the artifacts it creates follow Amazon&#8217;s design standards.</p><p>I say it a lot. AI is very good at making &#8220;plausible but wrong.&#8221;</p><p>Conformance is contextual to the environment.</p><h3>But we can&#8217;t just make another gate</h3><p>The default enterprise solution is probably going to be simple.</p><p>Review everything.</p><p>Create an approval process.</p><p>Put a human in the loop (...that poor human).</p><p>Build a better police force.</p><p>It won&#8217;t work.</p><p>Generative systems are supposed to dramatically increase the amount of plausible work produced. We couldn&#8217;t scale human review enough when humans were the only ones generating work. It&#8217;s even more impossible to grow proportionally with AI generated output.</p><p>Which defeats the economics of using that system to govern AI.</p><p>The best design system teams I&#8217;ve worked on weren&#8217;t successful because they were really efficient at rejecting bad work.</p><p>They made doing the right thing the path of least resistance.</p><p>Documented decisions. Context. Clear constraints.</p><p>Visibility to what already exists.</p><p>People were guided towards conformance <em>while they worked</em>. It wasn&#8217;t a case of coming to the end of the day and realizing they&#8217;d violated some rule they never knew about.</p><p>I&#8217;ve described this, slightly in jest, as <a href="https://www.robin-cannon.com/p/the-virtuous-design-system-panopticon?utm_source=internal">a virtuous panopticon</a>. Not watching everyone. Instead making the possible and the preferred more visible.</p><p>The same applies to AI. Conformance can&#8217;t come after, it needs to participate in the generation.</p><h3>Guide the way, don&#8217;t gate the path</h3><p>So that&#8217;s an interesting new enterprise AI infrastructure problem.</p><p>Not how to restrict what models can reach.</p><p>How do we make our organizational standards available to them while they work.</p><p>If the model is producing UI it needs to know components, accessibility requirements, voice and tone, patterns, and interaction conventions.</p><p>If it&#8217;s also writing code, it needs access to architecture decisions, security policies, dependency rules and engineering conventions.</p><p>Output needs to be evaluated against those constraints.</p><p>Block some failures before they&#8217;re made. Warn about others. Point to an alternative.</p><p>Or just make the relevant context more visible to the human-in-the-loop.</p><p>Governance can mean guidance as much as it means permission.</p><p>It&#8217;s not enough to have standards. It&#8217;s about having standards that are easy to follow.</p><p>Which, if you&#8217;ve worked in enterprise, you&#8217;ll know is very often not the case.</p><p>There&#8217;s a standard. Or six.</p><p>Policies and approved patterns.</p><p>And the people doing the work don&#8217;t know where to find it. Or how to choose which one they should use.</p><p>If it&#8217;s a human, they probably stop and ask someone.</p><p>Generative AI doesn&#8217;t stop. It produces some finished-looking artifact.</p><p>That can be a hard failure to notice.</p><h3>The missing half</h3><p>AI access governance needs to be sophisticated.</p><p>We need all that information about users, controls, audits, permissions and boundaries. Especially when we&#8217;re monitoring autonomous systems.</p><p>But those controls don&#8217;t tell us whether the thing those systems produced should exist.</p><p>That&#8217;s the conformance layer.</p><p>The layer that goes beyond just who can ask the question, but whether the answer that comes back is right.</p><p>Approval queues don&#8217;t work as governance at scale. They slow everyone down. They make people hate the expert teams reduced to policing them.</p><p>AI can actually give us an opportunity to make it better.</p><p>Make our standards more easily available when the work is being created.</p><p>That means encoding organizational knowledge at the right level. And evaluating our outputs directly against what the organization has already decided.</p><p>Guide people and agents in the right direction, before we have to stop them going in the wrong one.</p><p>Access protects the company from the AI. Conformance protects the company from what it asked the AI to do.</p><div><hr></div><h4>Further reading:</h4><ul><li><p><em><a href="https://www.robin-cannon.com/p/dont-become-the-design-systems-police?utm_source=internal">Don&#8217;t become the design systems police</a></em> - on showing up as a facilitator and not a blocker.</p></li><li><p>Nashawaty, P. &amp; Weston, S. <em><a href="https://www.efficientlyconnected.com/ai-output-governance-enterprise-blind-spot/">AI Output Governance: The Blind Spot in Enterprise AI Strategy.</a></em> Efficiently Connected, Aug 2026.</p></li><li><p>Sure, R. W. <em><a href="https://arxiv.org/pdf/2607.03516">The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence</a></em> (pdf). Independent Research Paper, Jun 2026.</p></li><li><p>Mugel, S. <em><a href="https://www.forbes.com/councils/forbestechcouncil/2026/08/05/enterprise-ais-governance-gap-runtime-safety-is-the-missing-layer/">Enterprise AI&#8217;s Governance Gap: Runtime Safety Is The Missing Layer.</a></em> Forbes, Aug 2026.</p></li></ul><p><em><span>Article photo by </span><a href="https://unsplash.com/@anniespratt">Annie Spratt</a><span> on </span><a href="https://unsplash.com/photos/a-metal-gate-leads-into-a-green-path-4aVDkPL313g">Unsplash</a><span>.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The "AI can do the UX" mistake]]></title><description><![CDATA[You might be cutting the wrong discipline]]></description><link>https://www.robin-cannon.com/p/the-ai-can-do-the-ux-mistake</link><guid isPermaLink="false">https://www.robin-cannon.com/p/the-ai-can-do-the-ux-mistake</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 04 Aug 2026 15:01:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7b93f8e4-235f-40be-b98f-117e70c1cf73_4849x3233.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There&#8217;s an experiment I would love to see.</p><p>Give a designer and a developer the same product brief.</p><p>Give them the same amount of time. The same access to users. The same AI tools.</p><p>The designer needs enough technical awareness to understand that architecture, performance, accessibility, and security are real constraints. The developer needs the design awareness to recognize familiar interaction patterns and make a coherent interface.</p><p>Let them work independently.</p><p>At the end of the time, don&#8217;t judge the codebase. Don&#8217;t count the features. Don&#8217;t ask which interface is more polished.</p><p>Put the products in front of users and see which one better solves their problem.</p><p>My bet is on the designer.</p><p>That wouldn&#8217;t have been the case a few years ago.</p><h2>The gates were never symmetrical</h2><p>Designers understand users. They can frame a problem. Structure the information they have into useful conclusions. Take all that, and determine what the experience should be.</p><p>But they couldn&#8217;t ship it.</p><p>They could make screens. They could have a click-through Figma prototype of the interactions. Explain the behavior with notes, annotations, tickets, meetings, and the all-important handoff.</p><p>So that someone else could make it real.</p><p>Developers had an opposite advantage. They can take an idea and make a working piece of software. What&#8217;s behind the interface - dependencies, data, performance implications, and how simple requirements get complicated very quickly when they meet reality.</p><p>But they didn&#8217;t have the judgement to decide how that software should work for a person.</p><p>Obviously these aren&#8217;t universal limitations. I know plenty of great designers who can code, and developers who have excellent design judgment.</p><p>But disciplines are training. They point our attention, so we notice different things.</p><p>Design trains people to see what&#8217;s confusing. What is incorrectly emphasized. The breaks between what a system allows and what a person is actually trying to do.</p><p>Engineers are trained to see fragility and bad abstractions. Understand the architecture and hidden dependencies. Bridge the gap between a convincing demo and a system to survive production.</p><p>Both those kinds of judgement are important.</p><p>AI hasn&#8217;t affected them equally.</p><h2>Didn&#8217;t we always want designers who code?</h2><p>AI gives the technically aware designer a remarkable amount of implementation capacity.</p><p>Scaffold an app. Connect it to APIs. Generate components or use the ones that exist. Explain unfamiliar code. Debug when things go wrong. Write a test suite. Take a clear spec around intended behavior and execute on it.</p><p>It will still need direction. A designer will need some technical understanding - enough to recognize if the system is making dangerous assumptions. To know if something is moving beyond their competence.</p><p>But it&#8217;s not the hard stop it used to be.</p><p>Designers don&#8217;t have to persuade a production chain to make something in order to discover whether it works.</p><p>They can make it.</p><p>AI also gives developers greater access to design production.</p><p>It can generate a clean dashboard. Use a familiar onboarding flow, or a plausible settings screen. It knows visible conventions of software very well. Sensible spacing, tidy cards, useful empty states - everything to make a product feel like it&#8217;s finished.</p><p>These upgrades aren&#8217;t symmetrical.</p><p>The system can increasingly perform implementation on the designer&#8217;s behalf. But design judgement can&#8217;t be acquired by a developer asking a system to produce design artifacts.</p><p>The artifact isn&#8217;t as valuable as the judgement. As the critique. As the understanding of the user.</p><h2>AI passes the &#8220;first look&#8221; test</h2><p>AI can make some very plausible interfaces. Many product generation tools lead with that capability. Here&#8217;s an immediate, visual, impressive interface.</p><p>The product appears before your eyes.</p><p>That looks like the design was the easy part.</p><p>It&#8217;s evidence of something else.</p><p>The presentation layer is the part of software applications where it&#8217;s easiest to manufacture the appearance of correctness, and hardest to verify if it&#8217;s actually correct.</p><p>If code is plausible but wrong, there are usually backstops. It doesn&#8217;t compile. The tests fail. There&#8217;s an error.</p><p>Engineering is great at detecting what&#8217;s incorrect because - if they don&#8217;t - the machine doesn&#8217;t execute its instructions.</p><p>Plausible-but-wrong design is dangerous because it can work perfectly.</p><p>The interface renders. The button works. The form submits cleanly.</p><p>It&#8217;s a great demo.</p><p>And there&#8217;s no alarm because the dashboard surfaced the wrong metric, and hid the one the user needed to make a decision.</p><p>There isn&#8217;t the same binary test when we cleanly and efficiently guide our user to the wrong outcome.</p><p>If we make a destructive action overly convenient, the app still builds.</p><p>We&#8217;ve got something usable.</p><p>We&#8217;ve got something wrong.</p><p>Those costs will come later - after it ships. Abandoned tasks. More support tickets. Bad business decisions because the information was wrong. Seemingly inexplicable churn.</p><p>It&#8217;s because AI is so fluent in the generalities of finished software that discrimination becomes more valuable, not less.</p><h2>It&#8217;s not a matter of taste</h2><p>This isn&#8217;t domain protection.</p><p>I&#8217;ve led design systems. Of course I think designers are important.</p><p>But the argument only works if design judgment means something tangible.</p><p>It&#8217;s not about choosing better typography, or adding visual polish. It&#8217;s not &#8220;make the logo bigger&#8221;. If it is, then AI has eaten up most of it already.</p><p>Some design execution can be automated - because production work in every discipline is getting easier and easier to automate.</p><p>Designers don&#8217;t have better taste.</p><p>Designers are best trained to determine if an experience represents what the user actually wants and needs to do.</p><p>What decision is the user trying to make? What information changes their decision? Which actions are reversible?</p><p>Does the structure of the product match the structure of the problem?</p><p>That&#8217;s judgement rather than taste. And, while not exclusive to designers, it&#8217;s what their discipline is supposed to train.</p><p>The design-aware developer knows that a destructive action may need confirmation.</p><p>The designer asks why the destruction is even an option.</p><h2>It&#8217;s different below the waterline</h2><p>The argument works on the surface. At the level the user sees.</p><p>Deeper than the presentation layer, and the results invert quickly.</p><p>Ask those same two people to create an infrastructure for high-scale transactions. To protect sensitive data. Recover from partial system failure. Avoid making an architectural decision that will constrain the company for a decade.</p><p>Now it&#8217;s the engineer&#8217;s judgement that&#8217;s the scarce thing.</p><p>AI produces plausible architecture, too.</p><p>So it&#8217;s not that designers now outrank developers. That&#8217;s dumb. And it reproduces the same mistake companies always seem to make - taking multidisciplinary product development and trying to turn it into a contest between the disciplines.</p><p>But AI shifts capability based on the nature of the missing skill.</p><p>AI is good at execution.</p><p>AI is bad at judgement.</p><p>If your historical limitation was execution, then AI gives you more benefit than the person whose historical limitation was judgement.</p><p>At the interface layer, that person is probably a designer.</p><h2>Our org charts were built for the world before AI</h2><p>Companies are making headcount decisions right now.</p><p>They&#8217;re compressing design teams, and ensuring engineering is the presumed center of their product creation.</p><p>And there&#8217;s plenty of logic to that. Hard engineering problems still exist. Software needs to work. And their AI investment is framed around making developers faster.</p><p>But that assumes the old distribution of capabilities.</p><p>The ability to ship is the decisive gate. Whole organizations have formed around that gate.</p><p>Now more people can go through it.</p><p>A designer who understands enough about software can increasingly go through it. Move from problem framing to working product, without a bunch of translation layers.</p><p>The reverse gate isn&#8217;t widening in the same way.</p><p>A developer can ask AI to produce a convincing interface. They can&#8217;t easily tell if that interface has the right understanding of the person who&#8217;ll be using it.</p><p>AI won&#8217;t be reliable at warning them.</p><p>Companies cutting design while concentrating on engineering investment &#8220;because AI can do the UX&#8221; might be reinforcing a discipline whose historical advantage is eroding the fastest. And reducing or removing the discipline whose central value AI is least suited to reproduce.</p><p>AI has made it much easier for designers to become builders.</p><p>It has not made it equally easy for builders to become designers.</p><div><hr></div><h4>Further reading:</h4><ul><li><p><em><a href="https://www.robin-cannon.com/p/execution-is-cheap-coordination-is?utm_source=crosslink">Execution is cheap. Coordination is not.</a></em> - on catching what&#8217;s organizationally wrong and codifying it.</p></li><li><p><a href="https://www.ideou.com/blogs/inspiration/ai-and-design-thinking">The Intersection of Design Thinking and AI: Enhancing Innovation.</a> IDEO U, Jun 2025.</p></li><li><p>Hill, M. <a href="https://www.aidataanalytics.network/data-science-ai/news-trends/half-of-developers-say-ai-can-code-better-than-most-people">Half of developers say AI can code better than most people.</a> AI Data &amp; Analytics Network, Aug 2025.</p></li><li><p>Morton, P. <a href="https://www.philmorton.co/why-is-ai-bad-at-design/">Why is AI bad at design?</a> Phil Morton, Jul 2026.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@eyrejune123?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Eyre June Bustamante</a> on <a href="https://unsplash.com/photos/woman-looking-at-lighted-neon-signage-Vz-S96BoIIY?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[I know kung fu. I might remember how to throw one punch.]]></title><description><![CDATA[I'm really, really good at AI. Right? I audited myself to find out.]]></description><link>https://www.robin-cannon.com/p/i-know-kung-fu-i-might-remember-how</link><guid isPermaLink="false">https://www.robin-cannon.com/p/i-know-kung-fu-i-might-remember-how</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Mon, 03 Aug 2026 04:30:45 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/dd06dbd4-17b8-4b9b-b8ad-3d3ee6417b2a_4635x3090.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This March I built another Claude skill suite - ten commands for helping to manage my own fitness. There was a daily check-in, a biweekly retrospective, and so on. A couple of hours work, designed to be used every single day.</p><p>Last week I audited my entire toolkit of AI skill suites. Telemetry had its own verdict.</p><p>I didn&#8217;t use the fitness suite. At all.</p><p>I&#8217;ll come back to that verdict.</p><p>In The Matrix, Neo jacks in, downloads a file, his eyes open: &#8220;I know kung fu.&#8221;</p><p>Using my deep research skill to define a strategic plan for a topic, and turning that into an AI skill suite, is the closest thing I&#8217;ve come to feeling like that for real. Research goes in, a tool comes out.</p><p>Capability acquired.</p><p>I&#8217;ve had that feeling multiple times. I have twenty-two domains for work and life - 193 skills in total. Product specs, competitive analysis, balcony gardening, wardrobe management.</p><p>Do I know kung fu? Or do I just have a big folder of downloads?</p><h3>I&#8217;m a bad person to answer my own question</h3><p>In 2025 METR ran a randomized trial. They gave experienced developers frontier AI tools for their own codebases, and measured what happened. The developers believed the AI tooling made them 20% faster.</p><p>They were 19% slower.</p><p>METR ran a follow-up in late 2025. The slowdown headline went away - newer tools, more practiced people. More interestingly, their measurement broke. Developers wouldn&#8217;t even submit tasks they&#8217;d have to do without AI. And time-tracking was unreliable when they applied it to agentic multitasking.</p><p>The people who measure this as their job concluded that their own results weren&#8217;t a good &#8220;proxy for the real productivity impact.&#8221;</p><p>So the durable finding isn&#8217;t the 19%. It&#8217;s the gap between what we feel and what we can verify - and that gap is getting harder to close, not easier.</p><p>This isn&#8217;t about developers, so much as it&#8217;s about our own testimony. Fluency feels fast. My tooling makes me feel capable, which makes me less reliable. With nearly two hundred tools, I&#8217;ve constructed myself into the least reliable witness possible.</p><p>So I decided to see if I could pull the data.</p><ul><li><p>Invocation - what actually ran, when, and where.</p></li><li><p>Artifact trail - whether the tools produced anything, and whether that went anywhere further.</p></li><li><p>Research lineage - what was fed downstream.</p></li></ul><p>And, fair to say, the telemetry itself isn&#8217;t perfect. It&#8217;s definitely missed skills that I know I use. Capture is hard.</p><p>But the numbers are more honest than my own feelings...my own vibes.</p><h3>What did the audit say?</h3><p>In the window I measured, 23% of the skills I&#8217;ve built were used. And only twelve skills carry half of all the activity.</p><p>That sounds terrible! 193 skills built, and I only really use twelve of them. What a failure!</p><p>But I don&#8217;t read it that way.</p><p>That&#8217;s not failure. It&#8217;s portfolio selection.</p><p>The marginal cost of building the skills with AI was tiny. And when the cost is tiny, the rational strategy is &#8220;build a bunch of stuff, and then let reality select.&#8221;</p><p>While the cost of building might be negligible, the cost of not deciding what to do with it is not.</p><p>My sin wasn&#8217;t building too much. It was that it took me this long to run the curation.</p><h3>Living, dead, and somewhere in-between</h3><p>When I look across my whole portfolio, I can break it down into six different states.</p><ul><li><p><strong>Regular use.</strong> My daily workhorses. A dozen skills doing half the work - meeting debriefs, deep research, spec writing. No notes, these are clearly things I actively reach for.</p></li><li><p><strong>Retired, with honors.</strong> My first delivery suite - twenty commands - I used heavily for months. But it&#8217;s silent because I built a successor and ran an A/B between them. Three of six predictions were wrong - including one I&#8217;d filed explicitly as &#8220;what would particularly surprise me.&#8221; The new suite won, the old one retired. That&#8217;s a system that works.</p></li><li><p><strong>Obsoleted by drift.</strong> I had sprint-cycle skills. We moved from cycles to Kanban. The tools weren&#8217;t wrong, they just didn&#8217;t apply any more.</p></li><li><p><strong>Dormant.</strong> I have some skills around tax preparation. They&#8217;ve done nothing all summer. That&#8217;s correct - they shouldn&#8217;t be doing anything in summer.</p></li><li><p><strong>Graduated.</strong> An interesting one. My wardrobe and style suite logged 46 outfits, cataloged a 222-item inventory, and scored what worked. It went quiet in June. But I think that&#8217;s because I&#8217;ve taken on those behaviors myself. The tool taught the pattern and then became unnecessary. I need to validate - I&#8217;m suspicious of that conclusion - but it&#8217;s a case where the kung fu downloaded into me, and didn&#8217;t stay in the file.</p></li><li><p><strong>Presumed dead.</strong> The fitness suite I mentioned earlier. Ten commands for regular use, zero recorded invocations.</p></li></ul><p>...except when I went to look back, even the fitness suite&#8217;s logbook told a different story. There are dated entries over several months. Including some that are inside my measurement window. So for some reason the telemetry was blind to it.</p><p>The audit result then suggests that there are zero confirmed &#8220;dead&#8221; skill suites. Silence by itself isn&#8217;t enough to declare them dead.</p><p>That telemetry issue needs investigating. Measurement brings discipline to judgment. It doesn&#8217;t replace it. Apparently a death certificate needs a second witness.</p><h3>Research - my busiest, most unread, most vital skill</h3><p>By volume, my deep research skill is the most productive by a lot. It&#8217;s made hundreds of artifacts over the last few months.</p><p>I almost never reread them.</p><p>But that&#8217;s the wrong model. If I use a skill to generate a competitive battlecard, its value is clear: someone opens it and uses it.</p><p>If it&#8217;s a research brief, its value comes from being transformed into something else. After that, it might never be opened again. The value moves forward.</p><p>I ran a trace on my ninety research runs. Eighty percent of them led to something real. Thirty-three of them became skills or skill suites themselves. Some of my most-used skills come from the research run that designed them. Others fed identifiable decisions. And some of them led nowhere - like fully drafted skills I never moved to install. No entire suite was confirmed dead. Plenty of individual work was.</p><p>The 20% dead rate helps me believe the 80%.</p><h3>We have to learn to say goodbye</h3><p>This goes beyond just skills creation.</p><p>Build-on-demand works. Twice this summer a real need appeared and I could implement a working Claude skill suite in days, one that immediately became integral.</p><p>For personal tooling like this, production isn&#8217;t the difficult bit.</p><p>But what&#8217;s missing - for my practice, but really in every AI-use conversation I see - is the other half of the loop.</p><p>Scheduled, honest selection. Kill criteria written at build time, before we get emotionally attached. A quarterly pass where every tool gets a verdict: active, dormant, retired, obsolete - or given a funeral if it&#8217;s actually dead.</p><p>And I still can&#8217;t empirically claim that this makes me faster or better than I&#8217;d be if I didn&#8217;t have it. I&#8217;m not sure anyone can claim that about themselves based on feelings alone. That&#8217;s the thirty-nine-percentage-point gap between perception and measured performance in the METR study.</p><p>Designing the experiment to measure that is what comes next.</p><p>So...I know kung fu. I have the logs to prove it.</p><p>But the only kung fu that&#8217;s mine is the punch I can still throw after the file closes.</p><p>Right now, that&#8217;s just one punch. I&#8217;ve counted it. Twice.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Becker, J, et al. <em><a href="https://metr.org/blog/2026-02-24-uplift-update/#wider-adoption-of-ai-has-made-it-more-difficult-to-measure-task-level-productivity">We are Changing our Developer Productivity Experiment Design</a></em>. METR, Feb 2026.</p></li><li><p><em><a href="https://getdx.com/news/new-data-ais-impact-on-engineering-velocity-is-more-modest-than-expected/">New data: AI&#8217;s impact on engineering velocity is more modest than expected</a></em>. DX, April 2026.</p></li><li><p>Counts, L. <em><a href="https://newsroom.haas.berkeley.edu/ai-promised-to-free-up-workers-time-uc-berkeley-haas-researchers-found-the-opposite/">AI promised to free up workers&#8217; time. UC Berkeley Haas researchers found the opposite</a></em>. UC Berkeley Haas, Feb 2026.</p></li></ul><p><em><span>Article photo by </span><a href="https://unsplash.com/@gettyimages">Getty Images</a><span> on </span><a href="https://unsplash.com/photos/strong-young-lady-with-boxing-gloves-punching-a-punching-bag-in-gym-alone-focus-is-on-hand-xhLiv-5ZfX4">Unsplash</a><span>.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Ideas are cheap. Invention is not.]]></title><description><![CDATA[AI can generate ideas forever. How do you decide which ones survive?]]></description><link>https://www.robin-cannon.com/p/ideas-are-cheap-invention-is-not</link><guid isPermaLink="false">https://www.robin-cannon.com/p/ideas-are-cheap-invention-is-not</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 28 Jul 2026 15:01:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/339894d8-7ccf-4aec-812d-7c6859d549db_6603x4280.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Recently, my five-year-old wanted to build a pillow fort.</p><p>But he&#8217;d decided it couldn&#8217;t just be any old pillow fort. It needed to be different. It needed to be more interesting than &#8220;put the cushions against the sofa and throw a blanket over it&#8221;.</p><p>He wanted me to suggest things.</p><p>A cave? No.</p><p>A spaceship? No.</p><p>An igloo? No.</p><p>A pirate ship? No.</p><p>And at a certain point, I hit the limits of my imagination. I was tired, and I think he&#8217;d just decided to say &#8220;no&#8221; to everything. He was a challenging client. So I said &#8220;why don&#8217;t we ask ChatGPT?&#8221;.</p><p>It gave me ten ideas immediately. Some of them I&#8217;d already suggested. But he latched on to the idea of building an animal hospital, and he was off and building.</p><p>AI is a good machine for getting unstuck - it doesn&#8217;t run out of ideas.</p><p>But if there&#8217;s a machine that can generate ideas indefinitely, ideas stop being a scarce resource. The scarce resource becomes something else: selection, validation, even taste.</p><p>It&#8217;s less &#8220;can I think of something?&#8221;</p><p>It&#8217;s more &#8220;which of these ideas deserves to survive?&#8221;</p><p>Everyone with access to an LLM has their own brainstorming machine.</p><p>Brainstorming fills up a page. It can be energizing. It produces a list of things. But the old refrain of &#8220;no bad ideas in a brainstorm&#8221; only lasts as long as the brainstorm does. There&#8217;s nothing to say that those ideas will turn into something people would use, pay for, or trust.</p><p>Idea generation is not invention.</p><h3>Building a pipeline, not a prompt</h3><p>I created a set of invention skills that, together, are a structured opportunity discovery and invention suite.</p><div class="callout-block" data-callout="true"><p><code>Discover &#8594; Generate &#8594; Stress-test &#8594; Validate &#8594; Brief</code></p></div><p>I need the sequence.</p><p>Too much AI ideation starts in the middle. Product ideas. New features. Startup concepts. Ten improvements.</p><p>We need to start before that.</p><p>What&#8217;s the opportunity? Who has the pain? What workarounds exist? What would make someone want to use this? What pushes them away from current solutions, or keeps them attached to what they already have?</p><p>Once there&#8217;s a problem worth exploring, the suite generates concepts that go through multiple methods.</p><p>This isn&#8217;t free association. This is structured invention.</p><p>The suite applies SIT patterns. Uses TRIZ-style contradictions. De Bono provocations. It will collide problems with another domain entirely. It&#8217;ll map a solution space morphologically, and build a Zwicky box of parameters and combinations.</p><p>And when it&#8217;s expanded the problem into potentially hundreds of combinations of ideas, it gets less generous.</p><p>Ideas are scored. Stress tested. It compares competing hypotheses against the evidence. Validation planning asks which experiments reduce uncertainty. We do a Mom Test.</p><p>We apply kill criteria.</p><p>And then the pipeline produces a brief.</p><p>One page that states the opportunity, the proposed solution, and the evidence that supports it - provenance is vital. What&#8217;s still uncertain? How will we validate? When do we stop?</p><p>If you can&#8217;t give an elevator pitch for your invention then it&#8217;s not finished.</p><p>That doesn&#8217;t mean it isn&#8217;t useful. But it&#8217;s not work yet.</p><h3>Pressure, not abundance</h3><p>Different methods apply different pressures.</p><ul><li><p>The opportunity scan will ask if there&#8217;s a real problem underneath an imagined solution.</p></li><li><p>SIT asks what happens if you remove something, divide it, give it new tasks, or add dependencies.</p></li><li><p>Collision takes a domain and forces it into contact with another. Sometimes that second domain might be ancient. Sometimes it&#8217;s ordinary. Sometimes it&#8217;s weird. The point is not to deliberately get exotic. The point is to apply distance and structure.</p></li><li><p>Morphological analysis maps the problem to configurations that might have been entirely overlooked.</p></li><li><p>Hypothesis testing weighs ideas against supporting and contradictory evidence.</p></li><li><p>Scoring forces those criteria into the open. It attempts to apply some objectivity, and to expose that judgment.</p></li><li><p>Validation planning asks what the world needs to show us for this idea to earn more confidence.</p></li><li><p>And the final brief ensures that the idea survives to the point of clear explanation.</p></li></ul><p>The suite isn&#8217;t magic prompts.</p><p>It&#8217;s a set of instrumentation, where each instrument marks and deforms the problem in different ways.</p><h3>Collision rather than metaphor</h3><p>The skill I have the most fun, for me, is collision.</p><p>In a recent piece, I wrote about how old institutions can be repositories of hard-won judgement. Guilds, courts, religious orders, astronomers and scribes. Not because they&#8217;re quaint, but because they solved problems around trust, readiness, drift and dissent centuries ago.</p><p>My invention suite entrenches that instinct. But it&#8217;s not limited to ancient history.</p><p>Collision uses anything structurally rich.</p><p>DJs building mix decks manage transition, mood and energy. Air traffic controllers sequence risk. Emergency rooms triage scarce attention and resources. I&#8217;ve collided ideas and inventions with sources ranging from prehistoric petroglyphs to jazz improvisation to standup comedy.</p><p>The source domain has to be useful.</p><p>Pure analogy is decoration. &#8220;The dashboard is like a city.&#8221; Great. Maybe. But something operational needs to flow from that comparison. Metaphors can give the impression of depth without making anything better.</p><p>My collision skill tries to avoid that trap by mapping domains structurally.</p><p>Who are the roles? What are the processes? What are the constraints? What feedback loops exist? How does the system fail?</p><p>After that mapping, it looks for isomorphisms: places where the bones of those structures match up.</p><p>I don&#8217;t want to know if one idea reminds me of another.</p><p>I want to know if one domain has used a mechanism to solve a problem, and whether I can apply that mechanism to a different domain.</p><p>Collision should provide functionality.</p><h3>Making a better dashboard</h3><p>I&#8217;ve been building an internal product flywheel. An AI-and-human pipeline turning customer signals into prioritized work. Scanning sources, routing work, and giving a product manager a draft queue to review.</p><p>There&#8217;s a dashboard, so that the product manager can see what the flywheel did.</p><p>It provided status, trust tracks, run history, signal counts, sizing. It showed what was under the flywheel&#8217;s hood.</p><p>But the question I was asking of the dashboard was a simple one:</p><blockquote><p><em>Do I need to do anything?</em></p></blockquote><p>My dashboard had the problem many dashboards have. It shows you everything you already know. It expects you to infer what matters.</p><p>I ran the invention suite against a question: what should the next version of the dashboard look like?</p><p>The invention suite came back with a cluster of solutions.</p><p>Pattern-driven concepts from SIT. Possible configurations from the morphological pass.</p><p>The collider put the dashboard into contact with petroglyphs and twelfth-century administrative registers. That sounds ridiculous - until it produces useful artifacts.</p><p>The collision focused on an idea missing from normal critique. Spatial hierarchy is semantic encoding.</p><p>If something urgent is visually subordinate, that&#8217;s confusing.</p><p>That wasn&#8217;t the whole answer. But it was an additional pressure brought to bear.</p><p>And the same recommendation began to coalesce. A headline-first dashboard, with an action queue separated from the observatory.</p><p>That&#8217;s a good idea. It might even seem an obvious idea. And the important piece isn&#8217;t that the AI produced it. It&#8217;s important that multiple methods converged on it.</p><p>SIT liked that there was no need to parse the deck first. Morphological analysis had already identified &#8220;headline plus detail&#8221; as a strong configuration. Collision liked it because spatial hierarchy encodes priority.</p><p>It was solving a problem painful enough to justify work. And the hypothesis testing didn&#8217;t come back with contradiction.</p><p>And so the suite of skills concluded with a brief.</p><p>Build a headline-first dashboard with a plain-English status sentence at the top, and in the browser tab title. Put a queue of actions immediately underneath it. Don&#8217;t remove all the existing dashboard, but put it below the fold - for observation rather than action.</p><p>With a set of validation tests.</p><p>Do PMs close the dashboard quickly if it&#8217;s all-clear? Do action items get resolved? Does the headline produce false &#8220;all-clear&#8221; messages?</p><p>Scoring, challenging, briefing and validating is invention.</p><h3>The Machine God isn&#8217;t omniscient</h3><p>This doesn&#8217;t make the machine wiser.</p><p>The invention suite can still produce nonsense. Scoring can be subjective, with a veneer of precision. Validation plans are useless unless someone runs the experiments.</p><p>But the suite provides a foundation for better judgement.</p><p>I don&#8217;t want any single method to be authoritative. I want it to disagree. I want it to consider a null hypothesis (&#8221;don&#8217;t do anything&#8221;). I want kill criteria.</p><p>Curiosity is a good reason to explore something.</p><p>The invention suite is intended to help us turn that curiosity into trust.</p><h3>Back at the pillow fort</h3><p>My son didn&#8217;t need the invention suite for his pillow fort.</p><p>He just wanted ten ideas, fast, from a machine that thought faster than a tired dad.</p><p>After that, it was just a source of childhood play. The big list of possibilities is great when the cost of choosing is just having to pick up some of the cushions.</p><p>That&#8217;s not most work.</p><p>Product ideas have cost. Engineering work has cost. A strategic choice has cost. AI-generated recommendations entering the workflow have cost.</p><p>And as AI becomes increasingly powerful at generating ideas, our judgement of those ideas becomes increasingly important.</p><p>AI makes ideation trivial.</p><p>That exposes how much of invention was never ideation in the first place.</p><p>Invention is about killing the idea you really wanted to like. Then turning the surviving thing into a brief that&#8217;s clear enough for someone else to decide what to do next.</p><p>Invention is not cheap.</p><div><hr></div><h4>Further reading:</h4><ul><li><p><em><a href="https://www.robin-cannon.com/p/why-i-want-my-ai-projects-blessed">Why I want my AI projects blessed by Jesuits</a></em> - on teaching AI what others learned centuries ago.</p></li><li><p>Lohr, S. <em><a href="https://www.nytimes.com/2023/07/15/technology/ai-inventor-patents.html">Can A.I. Invent?</a></em> New York Times, Jun 2023.</p></li><li><p>Schultz, B. N. N. <em><a href="https://www.si-labs.com/en/articles/morphological-box/">Morphological Box (Zwicky Box): Guide with CCA &amp; Example</a></em>. Service Innovation Labs, Feb 2026.</p></li><li><p>Rodriguez, G. R. <em><a href="https://www.forbes.com/sites/giovannirodriguez/2015/04/12/sit-israels-answer-to-design-thinking/">SIT: Israel&#8217;s Answer To Design Thinking?</a> </em>Forbes, Apr 2015.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@historyhd?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">History in HD</a> on <a href="https://unsplash.com/photos/white-metal-fence-on-white-sand-during-daytime-2MUqdhKBMzw?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for essays on design, technology, and culture - plus original fiction.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Your AI council needs better characters]]></title><description><![CDATA[The LLM Council is fine. But have you tried asking the contestants of Love Island instead?]]></description><link>https://www.robin-cannon.com/p/your-ai-council-needs-better-characters</link><guid isPermaLink="false">https://www.robin-cannon.com/p/your-ai-council-needs-better-characters</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 21 Jul 2026 15:01:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bbb1e7b4-0f88-47b6-b6d0-f9a66aeff917_5853x3902.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI councils are too respectable.</p><p>Boring.</p><p>Some combination of:</p><ul><li><p>The Strategic Advisor</p></li><li><p>The Risk Analyst</p></li><li><p>The Customer Advocate</p></li><li><p>The Technical Expert</p></li><li><p>The Innovation Lead</p></li></ul><p>A sensible group of faceless people who sound like they&#8217;re in a windowless office somewhere talking about operational synergy.</p><p>It doesn&#8217;t have to be like that.</p><p>It could be the Council of Elrond.</p><p>It could be a pirate crew.</p><p>It could be your younger self, your exhausted present self, your hypothetical 70-year-old self, and someone who thinks your entire definition of success is bonkers.</p><p>It could be the contestants of <em>Love Island</em>.</p><p>That&#8217;s the idea behind my <a href="https://github.com/shinytoyrobots/configurable-council">Configurable Council</a>, a Claude Code plugin that lets you run a question through a council of AI perspectives you define yourself.</p><p>The process might be serious.</p><p>The people don&#8217;t have to be.</p><h2>The council is a mechanism</h2><p>The original <a href="https://github.com/karpathy/llm-council">LLM Council</a>, by Andrej Karpathy, asks different language models the same question. They answer, review each other&#8217;s answers, and pass everything to a final model for a synthesized verdict.</p><p>The diversity comes from the models. Models from OpenAI, Anthropic, Google, and xAI can all make their individual points.</p><p>&#8220;LLM council&#8221; has become a bit of a broader pattern now. Run one model several times, give each instance a different role or perspective. Ask them to approach the problem from different angles.</p><p>The diversity is from the prompt, not the model.</p><p>The Configurable Council uses that second pattern.</p><p>Each member of the council answers separately. Then the answers are redistributed and they review them - anonymously. A chairman weighs the arguments and produces a recommendation.</p><p>But you can decide who&#8217;s on that council. And that changes more than just the report&#8217;s title.</p><h2>Characters are compact thinking models</h2><p>A default council might be sensible enough.</p><p>A Contrarian tries to find the fatal flaw. A First Principles Thinker asks whether you&#8217;re solving the right problem. An Expansionist looks for missed upside. An Outsider catches assumptions. An Executor asks what anyone is actually going to do next.</p><p>Professional. Entirely defensible.</p><p>But it&#8217;s not the Council of Elrond.</p><p>Gandalf challenges the framing. Aragorn focuses on practical execution. Boromir looks for power and opportunity. Frodo asks who&#8217;ll carry the burden. Galdor demands evidence. Elrond makes the final decision.</p><p>The process is the same.</p><p>The report is much more dramatic.</p><p>But the theme isn&#8217;t only decoration.</p><p>Any good character is a compact thinking model.</p><p>We already understand that Boromir will be drawn to power.</p><p>We understand why Frodo cares about the cost to the person doing the work.</p><p>We expect Gandalf to tell everyone they have misunderstood the problem.</p><p>So we already have a sense of their perspectives.</p><p>A &#8220;Risk and Governance Advisor&#8221; sounds important, but I don&#8217;t know how it really thinks.</p><p>I have a better idea of how Boromir thinks.</p><h2>It&#8217;s not about more responses</h2><p>Claude can already generate five answers to the same question.</p><p>What we&#8217;re trying to do with the council model is generate productive disagreement.</p><p>If everyone is kind and helpful, and they produce slightly different versions of the same answer, nothing really happens.</p><p>A council needs tension.</p><p>Someone should be arguing for the high-risk approach.</p><p>Another should think that&#8217;s irresponsible and dangerous.</p><p>Someone cares about speed. Someone else cares about precedent.</p><p>If one person wants to know if the plan will work, another might be asking if success creates a result worth having.</p><p>The Configurable Council lets you define explicit tensions. And that disagreement is vital.</p><p>A Contrarian isn&#8217;t there to add balance. They should be actively trying to kill the idea.</p><p>The Expansionist isn&#8217;t there to be a cheerleader. They should be trying to push you beyond the caution that undermines the idea.</p><p>That&#8217;s the work we want.</p><h2>There is no single council</h2><p>For a technical architecture question, you might create a council made up of a security engineer, an operator, a maintainer, a finance lead, and the person who gets paged at 3 a.m.</p><p>A product strategy question might need a customer, a salesperson, a competitor, a skeptical CFO, and a user who has already decided to leave.</p><p>When you&#8217;re making a career decision, maybe you need your ambitious younger self, your exhausted present self, your partner, your future self, and someone who thinks you are asking the wrong question entirely.</p><p>And some decisions might genuinely benefit from the contestants of <em>Love Island</em>.</p><p>Probably not because they know more about corporate strategy.</p><p>But expertise might not be what the decision is missing.</p><p>Maybe two departments keep saying they&#8217;re aligned, but they&#8217;re behaving like they&#8217;re waiting for someone better to walk into the villa.</p><p>Or one team thinks their relationship is an exclusive strategic commitment, and the other behaves like they still want to explore.</p><p>Different groups will notice different things.</p><p>Your organization probably already has the same groups of people coming together in the same rooms. People with similar backgrounds, incentives, and vocabulary.</p><p>Creating a bunch of AI personas that imitate those people doesn&#8217;t really give you much diversity.</p><p>A dwarf, a hobbit, an elf, a wizard, and a ranger might.</p><h2>It&#8217;s still your verdict</h2><p>A council is there to provide reasoning.</p><p>It&#8217;s not there to make the decision.</p><p>It&#8217;s still AI. A polished report, reviewed responses, a confident-sounding chair.</p><p>None of them have to live with the consequences.</p><p>Your AI council shouldn&#8217;t be an authority. It should expose the assumptions, objections, and missing perspectives around a decision.</p><p>It might give you a better answer. It might help you understand why you disagree with the answer it gives. Either will be useful.</p><p>The default council is fine.</p><p>The Council of Elrond might be better.</p><p>Sometimes what your product strategy really needs is a fire-pit conversation with eight people in swimwear trying to work out who&#8217;s here for the right reasons.</p><p>You don&#8217;t always need the same people in the room.</p><p>You don&#8217;t always need more experts in the room.</p><p>You might just need <em>different</em> people in the room.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>My <a href="https://github.com/shinytoyrobots/configurable-council">Configurable Council</a> on GitHub.</p></li><li><p>Krishnan, R. <em><a href="https://www.strangeloopcanon.com/p/llm-councils-show-groupthink">LLM councils show groupthink.</a></em> Strange Loop Canon, Jun 2026.</p></li><li><p>Champagne, S. <em><a href="https://www.today.com/popculture/tv/love-island-mental-health-effects-psychologist-rcna220785">What&#8217;s Really Behind &#8216;Love Island USA&#8217; Drama? A Psychologist Explains.</a> </em>Today, Aug 2025.</p></li><li><p>Jacobs, Prof A. <em><a href="https://blog.ayjay.org/why-gandalf-and-elrond-were-wrong/">Why Gandalf and Elrond were wrong.</a></em> The Homebound Symphony, Aug 2011.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@piiiiine?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Muhammadh Saamy</a> on <a href="https://unsplash.com/photos/woman-in-bikini-lying-on-beach-during-daytime-YTXDMf5UWzc?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for essays on design, technology, and culture - plus original fiction.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Actually, the shape of the work does change]]></title><description><![CDATA[Bolting agents onto sprints only makes the horse run faster.]]></description><link>https://www.robin-cannon.com/p/actually-the-shape-of-the-work-does</link><guid isPermaLink="false">https://www.robin-cannon.com/p/actually-the-shape-of-the-work-does</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 14 Jul 2026 15:01:59 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/13b3c85c-847c-4b28-ae37-6db9cde4d09e_5432x3621.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>So many of the approaches to AI in software delivery are about making our existing workflows faster. We take sprints, stories, gates and bolt agents into the seats. Work moves quicker.</p><p>The shape of the work doesn&#8217;t change.</p><p>It&#8217;s faster horses.</p><p>I know, because I built one of them.</p><p>I made a Claude skill suite called <code>delivery-team</code>. Thirteen role-shaped agents; scrum master, devs, QA, conservative and aggressive project managers - all moving stories through a multi-stage pipeline.</p><p>It&#8217;s good, too. Fast, tireless, it never lost the thread. It moved &#8220;vibe coding&#8221; to something with much more rigor around it. It wasn&#8217;t something I was looking to replace.</p><p>But a couple of months ago I had the chance to sit down with some of the staff at <a href="https://obvious.ai/">Obvious.ai</a> to talk about their autobuild solution. And an idea got stuck in my head.</p><p>It wasn&#8217;t &#8220;how do I make this faster?&#8221;</p><p>I wondered if we were looking at things the wrong way. That speeding up human processes is the least interesting things we can do with the power of AI.</p><p>Scrum exists because humans get tired, change their minds, need ceremony to keep coordinated. Agents have none of those problems.</p><p>So I built a second suite of skills. It&#8217;s called <code>flow</code>, and it keeps almost nothing from the first. It trades sprints for convergence. Stories for detailed, executable specs (something Obvious.ai highlighted), pass/fail gates for a Pareto front of competing implementations. No retros, but preserved dissents that reactivate when the conditions come true.</p><p>This piece is two things. I&#8217;ll talk about the practicalities of each suite - the repo is public, there&#8217;s a website explainer, you&#8217;re welcome to run either. But it&#8217;s also an argument about the underlying thoughts: about the difference between speeding up human workflow or asking what delivery can look like when humans don&#8217;t have to coordinate inside it.</p><h3>Respect to the horse</h3><p>I built the horse first. And I built it with care.</p><p><code>delivery-team</code> simulates a thirteen person Scrum team, with seven gated stages. They coordinate with documents - agents reading and writing artifacts with clear schema. QA can veto but not write code. Architects lay down high level skeletons that are sharded into stories, and that avoids drift. It uses recognized foundations: BMAD, MetaGPT, Team Topologies.</p><p>I can point it at a codebase, or work with it to build from scratch. It goes fast, and it&#8217;s good. It never gets tired, never waits for a calendar, and never forgets where it&#8217;s at.</p><p>But that&#8217;s also a problem.</p><p>It&#8217;s accelerating workarounds that we made for humans. It&#8217;s not doing AI-native delivery.</p><p>Hell, it still asks me &#8220;do you want this to be a 3-day, 5-day, or 10-day sprint?&#8221;, even when I know the work will be done in about an hour and a half.</p><ol><li><p><strong>Role abstraction is a crutch.</strong></p><p>The suite has thirteen fixed agents. Hand-tuned. But those roles exist because hiring humans is expensive and so we specialize and we commit. That&#8217;s payroll, not architecture.<br></p><p>Google and the University of Cambridge&#8217;s paper on Multi-Agent Design found that optimizing topology and prompts beats fixed roleplay by nearly 80% on agentic tasks. Anthropic&#8217;s own research system dispatches dynamically - one agent for something simple, ten or more for something hard. It&#8217;s not using the same org chart every time.</p></li><li><p><strong>Sprints are a backstop for risks in human commitment.</strong></p><p>We time-box sprints because calendar coordination is costly, and if we put things into two-week boxes then it&#8217;s safer. An agent doesn&#8217;t have any of those uncertainties, and it doesn&#8217;t need a calendar.<br></p><p>Sprint boundaries are arbitrary lines we use to manage people. Not the work.</p></li><li><p><strong>Heuristics for humans.</strong></p><p>Story files locked at eighty percent test coverage. That&#8217;s a number to trade human effort for human risk. It&#8217;s not a property of LLM code. But our stories start drifting the moment we start work.</p></li></ol><p>Rituals aren&#8217;t features. And if we put them in the AI loop, they start to become obstacles.</p><h3>What if nobody gets tired?</h3><p><code>flow</code> is designed <em>for</em> agents instead of trying to design around them.</p><p>I took some insights from Cognition AI&#8217;s 2025 engineering post.</p><ul><li><p>Actions carry implicit decisions, and conflicting decisions carry bad results.</p></li><li><p>Agents can read in parallel safely. But they cannot <em>write</em> in parallel safely, because they all made decisions and accumulate a pile of them that aren&#8217;t reconciled.</p></li></ul><p><code>flow</code> lets multiple agents read, score, search and dissent - but only one of them commits. Intelligence works in parallel, but writing is serial.</p><p>Not my invention, but it&#8217;s what I built around. It shapes everything that happens downstream of it.</p><ol><li><p><strong>Specs are the source of truth.</strong></p><p>I think this is a concept fast getting traction. Take the whenwords library, which shipped in February. More than seven hundred conformance tests, no hand-written code, and all the maintenance is in the spec.<br></p><p>The code is the build artifact of the spec.</p></li><li><p><strong>Stories become generations.</strong> <br>We build a wide population of implemented variants. They&#8217;re scored against an eval suite. It&#8217;s not one attempt that we choose to accept or send back.</p></li><li><p><strong>Acceptance gates become a Pareto front.</strong></p><p>We treat quality as a vector. We&#8217;re trading speed against simplicity, simplicity against security, and so on.</p></li><li><p><strong>Retros become diffs, reviews aren&#8217;t forgotten.</strong></p><p>We maintain a running diff between predictions and productions, rather than the ceremony. And dissents are saved as objects that get reactivated under certain conditions. If a reviewer warns &#8220;this will break if we introduce X,&#8221; they might be right later.</p></li></ol><p>The bit that still feels like AI being &#8220;magic&#8221; is what happens next. <code>flow</code> doesn&#8217;t build one implementation and refine it. It builds five (or six, or seven - it decides how many it needs) at once, all from the same spec. It scores against different evals, and each build focuses attention on a different metric. Then it converges, keeps what&#8217;s good, regenerates and converges again.</p><p>That&#8217;s five parallel takes being iterated in tandem. If that was a human team, that&#8217;s five engineers, five branches, a month of time we probably don&#8217;t have. An agent population does it in one afternoon.</p><p>It&#8217;s a weird way to work. It burns a lot of tokens. But watching it happen convinced me it&#8217;s a question worth chasing after.</p><h3>But we all know the price of gas</h3><p>None of this is perfect.</p><p><code>flow</code> costs about five times the generation tokens of delivery-team. That signal might be worth every cent. Or you might have spent fifteen dollars when three dollars would have told you what you need.</p><p>It fails in different ways, too.</p><p>Sometimes it will try to game its own metric. Generators tend to satisfy the eval as much as solve the problem. Goodhart&#8217;s Law lives in the same system I built to avoid Goodhart&#8217;s Law about old threshholds!</p><p>Eval suites can be wrong, and wrong evals aren&#8217;t a bug, they&#8217;re a problem with the spec. So we might confidently reward the wrong variant.</p><p>It&#8217;s not fully automated yet (although neither is <code>delivery-team</code>), which is already leaving opportunities for acceleration on the table.</p><p>Perhaps most importantly, <code>flow</code> needs that spec, so you need to be able to write it. If it&#8217;s doing exploratory work, or you don&#8217;t know what &#8220;correct&#8221; even means, it&#8217;s going to flail. And its flailing gets expensive.</p><p>The horse is still valid. <code>delivery-team</code> might still be the right tool. When the spec can&#8217;t be pinned down. When stakeholders need a recognizable word like &#8220;sprint&#8221; to trust the machine. You can reach for different skills at different times.</p><p>My curiosity has resolved into a much better set of questions. Not yet a verdict.</p><h3>Curiosity killed the cat</h3><p>Curiosity is a great reason to build something. But a bad reason to trust it.</p><p>So run the same effort through both. You can write the same problem statement. Define the same set of evals. And look at the metrics; time to ship, token spend, human review time, defect count, etc. <code>flow</code> should beat the <code>delivery-team</code> horse, and it certainly shouldn&#8217;t regress.</p><p>This is still a live test. It isn&#8217;t a launch.</p><p>Both suites are open-source. </p><ul><li><p>The repo is at <a href="https://github.com/shinytoyrobots/agentic-delivery-suites">github.com/shinytoyrobots/agentic-delivery-suites</a>.</p></li><li><p>Further documentation at <a href="https://shinytoyrobots.github.io/agentic-delivery-suites/">shinytoyrobots.github.io/agentic-delivery-suites</a>.</p></li></ul><p>Clone them, run them, apply them to your work. Tell me what&#8217;s wrong.</p><h3>The fork</h3><p>It&#8217;s not really about the specific skill suites.</p><p>It&#8217;s about the problem. Can you write the spec? Can you instrument the evals? Do stakeholders need a specific vocabulary to instill trust.</p><p>One of these is the better horse. The other one <em>might</em> be a car.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Zhou, H, et al. <a href="https://arxiv.org/abs/2502.02533">Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies</a>. Google, Feb 2025.</p></li><li><p>Yan, W. <a href="https://cognition.com/blog/dont-build-multi-agents">Don&#8217;t Build Multi-Agents</a>. Cognition, June 2025.</p></li><li><p>Breunig, D. <a href="https://www.dbreunig.com/2026/01/08/a-software-library-with-no-code.html">A Software Library with No Code</a>. dbreunig.com, Jan 2026.</p></li><li><p>Marwala, T. <a href="https://unu.edu/article/greatest-good-exists-not-extremes-through-exploration-middle-ground-pareto">&#8216;Greatest Good&#8217; Exists Not at the Extremes but Through Exploration of the Middle Ground &#8212; Pareto</a>. United Nations University, Feb 2024.</p></li></ul><p><em>Article Photo by <a href="https://unsplash.com/@octopus_photo?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Pete Godfrey</a> on <a href="https://unsplash.com/photos/a-large-flock-of-birds-flying-over-a-body-of-water-jKNR--HDA_A?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Why I want my AI projects blessed by Jesuits]]></title><description><![CDATA[The hardest problems in AI aren't in the code. They're problems of judgment. And someone already solved them.]]></description><link>https://www.robin-cannon.com/p/why-i-want-my-ai-projects-blessed</link><guid isPermaLink="false">https://www.robin-cannon.com/p/why-i-want-my-ai-projects-blessed</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 30 Jun 2026 15:00:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8eba08bc-305e-4a22-bcee-179be804d4b4_5948x3965.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In the early 18th century, Maharaja Jai Singh II built five Jantar Mantar complexes. Astronomical observatories with no lenses, no electronics, and no moving parts. They used them, in part, to ensure their astrological birth charts were more accurately cast.</p><p>Earlier this year this centuries old approach to astronomy and astrology solved a problem I had trying to fix a dashboard.</p><p>The dashboard&#8217;s mine. It&#8217;s on top of a system I&#8217;ll come back to later - an AI pipeline to turn customer signals into executable work. The dashboard tells me if I can still trust the machine&#8217;s judgement.</p><p>It had a problem every dashboard faces. It can&#8217;t see its own drift.</p><p>When the AI slowly gets worse, and if my own instinct on &#8220;good enough&#8221; also slides, the dashboard will have green lights all the way down. It&#8217;s not lying, but it&#8217;s quietly going wrong.</p><p>The astronomers at the Jantar Mantar solved the problem centuries ago.</p><p>They knew their instruments drifted. They knew the assumptions they had in their calendar would, over time, pull away from the actual stars in the heavens.</p><p>They also didn&#8217;t trust the running system to catch its own decay. They re-anchored, on a schedule, against a fixed external baseline. It&#8217;s a mathematical correction called Ayanamsha to anchor their astrology to the fixed stars, and account for the earth&#8217;s drift.</p><p>I didn&#8217;t learn this solution from a digital product blog about dashboards.</p><p>But I did get it on purpose. By colliding the problem in front of me with a domain that had nothing to do with it. And, once I really started doing that deliberately, I can&#8217;t stop noticing a pattern. The hardest problems in the newest technology - AI - are often old problems wearing new clothes. And ancient answers are better than ones we&#8217;re busy reinventing.</p><h3>The library five years shallow</h3><p>Most people building with AI are reasoning from five years of data at most. Last quarter&#8217;s framework, the most recent SaaS playbook, or the pattern that worked last time.</p><p>That&#8217;s not a knock. It&#8217;s a fast moving field. Five years ago feels like forever. But it means we&#8217;re solving ancient problems with a shallow frame of reference.</p><p>And it&#8217;s not just a story about code. That&#8217;s the one we&#8217;re telling the most. AI writes a function. AI reviews a pull request. AI ships a feature. That&#8217;s real, interesting, and only a slice of the whole.</p><p>I lead product. When I&#8217;m using AI it&#8217;s not primarily about writing code. I&#8217;m triaging customer signals, internal data, and predictions about the future. And using that data to recommend how to route, draft specs.</p><p>How to decide what&#8217;s worth building.</p><p>Those are judgement calls, and I need to know how much of that judgement I feel comfortable trusting.</p><p>Is something ready? Calibrated? True? Is that prediction honest?</p><p>These aren&#8217;t engineering problems. They&#8217;re the problems guilds, courts, councils, and churches have stress-tested over centuries - and we can use the answers they wrote down.</p><p>So I went looking for the answers on purpose. This is how I did that, and what I found when I started building with them.</p><h3>Reading old code</h3><p>The how is a method I built into a skill called <code>inv-collide</code>. Part of a broader suite of invention related skills.</p><p>It stems from an idea called bisociation, created by the author Arthur Koestler. Then it operationalizes that idea.</p><p>Take a problem in front of you, and then a domain that has nothing to do with it, and you force them through three steps.</p><ol><li><p>Map both as structures. The roles, processes, constraints, feedback loops, value flows, and failure modes. Not what they&#8217;re <em>about</em>, but about the bones of what they do.</p></li><li><p>Find the places where the bones are identical. Isomorphisms.</p></li><li><p>Generate concepts at that intersection. <em>e.g. if this domain solves problem P with mechanism M, and I also have a version of problem P, can I transplant M?</em></p></li></ol><p>It&#8217;s not the same as brainstorming. Brainstorming is free-association. Bisociation matches the bones of the thing. That discipline makes the output usable.</p><p>And that discipline requires a step people might skip.</p><h3>Disciplined enough to discard it</h3><p>It&#8217;s all very romantic. I found a poetic parallel between a dashboard and the stars in the sky. Well, the universe is really big and it&#8217;s easy to make up a metaphor.</p><p>If the method was just making nice-sounding coincidences, it would be a party trick.</p><p>It also generates a big discard pile.</p><p>When I collided the trust dashboard with the Jyotish astronomy and the Jaipur observatory, it surfaced twelve structural matches.</p><p>Astrological <em>muhurta</em> - an auspicious window where conditions align and you can move forward - mapped directly onto a promotion gate in my system. The point where AI capability became more trusted to act autonomously.</p><p>That got thrown out. It wasn&#8217;t bad, but my system already had a streak counter and encoded readiness gates. This added costume, not structure.</p><p>The collisions are only worth integrating when they survive an honest attempt to kill them. Re-anchoring survived. It runs in production.</p><h3>Running old code in my product flywheel</h3><p>Underneath the status dashboard is what I call my product flywheel. It&#8217;s an AI-and-human pipeline to turn customer signal and internal consensus into execution ready work, without a person writing every issue.</p><p>It reads from various sources - customer knowledge base, Slack, Zendesk. Then it classifies and routes what&#8217;s changed. It&#8217;ll draft a business case and independently assess it, scaffolding the approved work into Linear. Then it will recommend routing, priority, and provides a rate-limited queue for engineering to pull from.</p><p>This is product work. The product flywheel feeds execution, but it isn&#8217;t the execution. This is triage, routing, specs, and prioritization. Engineering pulls from its output in order to execute.</p><p>This isn&#8217;t a story about code review. These are rules for intelligence from medieval guilds, eighteenth-century astronomers, and the Talmud. And they&#8217;re applied to the most difficult parts of trusting AI with product judgement. Applying methods that were created a long time ago.</p><p>Here are three that are running.</p><h3>Earned trust, and the medieval masterwork</h3><p>One of the oldest problems in management. When do you let someone work unsupervised?</p><p>If you get it wrong early, you&#8217;re letting unqualified hands do a lot of damage in your name.</p><p>If you get it wrong late, you&#8217;re throttling the potential of someone who was ready.</p><p>Every craft tradition that&#8217;s lasted seems to have solved this in the same way. With a gate. A guild apprentice submitted a masterwork, and the sitting masters judged it. A Jesuit student reached a point where his judgement was, in their words, <em>formed</em>. Trust was domain-specific, slow to build, easy to lose.</p><p>AI tooling uses a flag. We set a permission mode, and we crank the autonomy up and down depending on how brave we&#8217;re feeling. But the system isn&#8217;t demonstrating anything.</p><p>The flywheel doesn&#8217;t have a dial. Its agents earn the right to act, and they don&#8217;t start with it.</p><p>I track trust per stream. Incoming customer signals, routing, drafting, the queue controller all have their own standing. Being good at one thing is no evidence of being good at anything else.</p><p>Each capability can climb through three tiers. The tier changes what the agent is allowed to do. An apprentice intake agent reports what it <em>would</em> have done, line by line, and needs a human to confirm every call. A journeyman provides a summary of what it would do, and asks for a single confirmation of that summary. A master executes and then reports back for review.</p><p>An agent moves from apprentice to journeyman on fourteen consecutive clean runs - confirmed, by a human, as correct. The journeyman to master takes thirty. And this isn&#8217;t an average. One bad run resets the streak to zero. And mistakes made by more &#8220;senior&#8221; agents mean demotion and a doubled threshold.</p><p>That&#8217;s the medieval guild structure, as a YAML file.</p><p>Trust is isolated by domain. It&#8217;s earned slowly, lost easily, and more expensive to win back a second time. That&#8217;s not how permission flags work, but every master craftsman who ever lived would recognize it.</p><p>That makes autonomy something earned - and easily lost.</p><h3>Calibrating my judgement against the stars</h3><p>Let me close the loop on that dashboard.</p><p>I&#8217;m not worried about the AI &#8220;breaking&#8221;. That will be loud and obvious. What I&#8217;m worried about is silent drift. Miscalibration as the AI degrades slowly, and the gap never shows up.</p><p>Those astronomers in Jaipur had an answer. Re-anchor on a fixed, external, baseline. On a schedule. Not trusting the instrument to audit itself.</p><p>My flywheel uses two versions of that.</p><ol><li><p>A deliberately imperfect target. My runs are clean if I override the AI&#8217;s call no more than fifteen percent of the time. Not zero. Zero overrides is a person not paying attention. We want a target that keeps someone in the loop.</p></li><li><p>You can&#8217;t re-baseline a scoring system - even if it&#8217;s an improvement - without re-scoring at least five historical projects against the new rules. Then recording the sign-off and noting the discontinuity. You can&#8217;t move a baseline and also erase the evidence that you moved it.</p></li></ol><p>That&#8217;s Jantar Mantar. Re-anchoring against a fixed point, on a cadence. And keeping its receipts.</p><h3>Adversarial truth-seeking, in the room and in the spec</h3><p>If we review for consensus we throw information away.</p><p>Two smart, competent people disagree about a tough problem. When that disagreement gets resolved, we move on.</p><p>But that signal tells you a problem has more than one shape. And that losing argument might be right under conditions that develop in future.</p><p>The Jesuits formed students through <em>disputatio</em>. A structured, adversarial defense. You don&#8217;t prove you&#8217;re competent. You prove you&#8217;re competent when a skeptic attacks your work.</p><p>The Talmud has preserved the minority ruling alongside the majority one, deliberately, for two thousand years. A defeated argument might become the right argument when the world changes.</p><p>Those memories need to be written down. If they don&#8217;t, we don&#8217;t remember them. A dissent gets aired in a meeting, a decision gets made, and the losing argument is - at best - noted in a retro document that nobody reads. That&#8217;s even worse for product decisions - routing, prioritizing, build this not that - than code, because the record&#8217;s thinner.</p><p>The flywheel runs these institutions against product judgement.</p><p><em>Disputatio</em> is in every routing decision. Recommendations don&#8217;t come from one agent, they come from a pair. One proposes the route and the priority, and the second is designed specifically to challenge it. Attack the recommendation before it&#8217;s committed. The proposal has to survive the examination. That&#8217;s Jesuit insight - that the defense is where the judgement is formed - applied to a decision about a customer request.</p><p><em>Chavruta</em> is the basis for how the system records the disagreements. Reviews produce dissents. The dissents don&#8217;t need to be resolved, but they are <em>committed</em>. And not just as a stale record. They&#8217;re structured objects that include conditions under which they wake up again. Recheck at the next live run. Resurface if a specific metric flatlines. When the world changes to match a trigger, that old losing argument comes back on its own and demands a second hearing.</p><p>Consensus is lossy. The Talmud knew that two thousand years ago. My flywheel acts on that memory.</p><h3>Others still on the bench</h3><p>There&#8217;s a couple more collisions which produced things that I haven&#8217;t shipped yet. The same method, but in the design stage.</p><ol><li><p><em>Drift lineage.</em> A text copied by hand for a thousand years has scribal drift - small errors that harden into official fact <em>because</em> the chain is trusted. Traditions that survive have apparatus for this - lay one manuscript next to another and see where the text mutated. I&#8217;m building the same thing for a claim - a walk backwards to the primary source, and flags where the numbers or the conclusions changed. Not whether those changes were right or wrong, but a lineage of provenance and drift from source.</p></li><li><p>Avoiding &#8220;<em>vaticinium ex eventu</em>&#8220;. This is prophecy after the fact. Re-reading the record once you know the outcome, and bending the record to show you were right. AI multiplies confident predictions, and decision journals are editable. Medieval clerks had the answer. A boring, comprehensive one. Witnessed, dated, tamper-evident entries. Predictions logged and stamped, and scored at outcome time. It&#8217;s about not lying to yourself about a prediction.</p></li></ol><h3>We keep reinventing what was proven</h3><p>It&#8217;s about trust.</p><p>Transferring trust to something that acts in your name. Keeping your own judgement from drift. Not letting disagreements die. And being honest about predictions.</p><p>None of them are about AI writing code.</p><p>It&#8217;s about trusting AI&#8217;s judgement, calibrating it against my own, auditing it, and keeping us both honest.</p><p>This match to old institutions is real, not a flourish. History is not just a charming source of cool metaphors. It&#8217;s where they already ran these experiments. Where they wrote down the answer in a language nobody speaks any more.</p><p>It&#8217;s the debugging history of our species, and we&#8217;re reaching past it for last quarter&#8217;s new framework.</p><p>When you hand an AI something that matters - a decision, a forecast, a release - find the institutions that already solved it.</p><p>Read their old code.</p><p>Get the thing blessed by Jesuits.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Sarda, S. <em><a href="https://www.bbc.com/travel/article/20220530-jantar-mantar-indias-mysterious-gateway-to-the-stars">India&#8217;s mysterious gateway to the stars.</a> </em>BBC, May 2022.</p></li><li><p>Popova, P. <em><a href="https://www.themarginalian.org/2013/05/20/arthur-koestler-creativity-bisociation/"><span>How Creativity in Humor, Art, and Science Works: Arthur Koestler&#8217;s Theory of Bisociation.</span></a><span> </span></em><span>The Marginalian, May 2013.</span></p></li><li><p>Farrell, A. P. <em><a href="https://www.educatemagis.org/wp-content/uploads/documents/2019/09/ratio-studiorum-1599.pdf">The Jesuit Ratio Studiorum of 1599.</a> </em>Conference of Major Superiors of Jesuits, 1970.</p></li></ul><p><em><span>Article photo by </span><a href="https://unsplash.com/@dagerotip"><span>George Dagerotip</span></a><span> on </span><a href="https://unsplash.com/photos/woman-holding-a-cup-of-coffee-with-red-nails-ub7fZv70bQI?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyTexthttps://unsplash.com/photos/a-red-building-with-a-spiral-design-on-it-ZjqmN5lvHhU">Unsplash</a><span>.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[What does Figma do next?]]></title><description><![CDATA[Figma solved the problem of making design multiplayer. It might still be solving that problem when the problem has changed.]]></description><link>https://www.robin-cannon.com/p/what-does-figma-do-next</link><guid isPermaLink="false">https://www.robin-cannon.com/p/what-does-figma-do-next</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 23 Jun 2026 14:01:43 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/66f42207-c948-4958-95da-1b58cbff9318_4032x2268.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Figma has a deep collection of useful features.</p><p>It also seems to have a problem: a strategic imagination still bound to the canvas.</p><p>I realize that&#8217;s a challenging thing to say about perhaps the most important product tool of the past decade. This is not a &#8220;Figma is dead&#8221; article.</p><p>Figma changed how digital product teams work. It made design a genuinely multiplayer activity. It made a design file a shared space. Collaboration, critique, exploration, and handoff in a browser-based canvas everyone could see.</p><p>Sketch looked comfortable before Figma came along. Users and workflows and plugins, and enterprise legitimacy. A whole ecosystem. InVision for prototypes, Zeplin to support handoff. Abstract for version control.</p><p>Then Figma came in like the Kool-Aid Man and made Sketch look obsolete almost overnight.</p><p>It wasn&#8217;t anything to do with Sketch&#8217;s design features. It could still draw a rectangle!</p><p>But Figma changed the whole basis of where the two products were competing. Not the design tool with the best interface, but making design collaborative.</p><p>I&#8217;ve never seen another product that created as much practitioner pressure for change as the internal demand at IBM to switch from Sketch to Figma. It overcame corporate inertia faster than I&#8217;d have imagined.</p><p>Figma just had better answers. Staying on Sketch meant being left behind.</p><p>Figma solved the coordination problem of its moment. Its risk is in continuing to solve the problem after the problem has changed.</p><p>There is a historical parallel. But it&#8217;s not as glib as &#8220;Figma is the new Sketch&#8221;. That&#8217;s too neat. Figma is clearly larger, more deeply embedded, and has a degree of strategic awareness.</p><p>But incumbents don&#8217;t usually look like they&#8217;re sleeping. Especially from the inside.</p><p>Figma is shipping a lot of stuff. And they&#8217;re telling a coherent story about the future that runs through them.</p><p>Are they building that future, or just extending the conditions that made them dominant before?</p><p>The center of gravity is moving from canvas to code.</p><p>That means from abstraction to execution. From static artifacts to live systems. And from design files to context that AI interprets and generates from.</p><p>Designers will still need visual tools. And teams will need shared spaces for critique and exploration.</p><p>But what does Figma do when the canvas is not the center of gravity?</p><h2>How Figma won in the first place</h2><p>Figma&#8217;s first great achievement was technical. They made the browser matter far more for design than anyone thought possible. Cross-platform access mattered. Performance mattered. Multiplayer mattered.</p><p>The product was excellent, and execution counts.</p><p>But the deeper shift was cultural.</p><p>Before Figma, collaboration was fragmented. It needed local files, redlines, PDFs, and those meetings where everyone asked &#8220;is this the right version?&#8221; Figma collapsed all that distance.</p><p>Figma wasn&#8217;t merely a better canvas. Figma was a better coordination model.</p><p>It made work around the design abstraction collaborative. Which was a huge step forward.</p><p>But an abstraction is still an abstraction.</p><p>The canvas is not the product. It&#8217;s a representation. The real product is in code.</p><p>The canvas was vital for helping us think before the reality of implementation got too expensive.</p><p>But it depends on a world where there&#8217;s a big gap between visual intent and working software. That&#8217;s where the abstraction lives.</p><p>AI is collapsing that distance.</p><h2>The canvas answers a translation problem</h2><p>The canvas makes sense.</p><p>Designers express intent. Engineers translate the intent into code. Product managers mediate priority and scope.</p><p>We use the thing we imagined to help us ship the thing that&#8217;s real.</p><p>And that model isn&#8217;t going to be going away any time soon. Many organizations will likely work this way for years to come, if they can get away with it.</p><p>But the direction of travel has changed.</p><p>Design-to-code is faster. Which is great. But it&#8217;s just collapsing the way we already work. Handoff, but faster. Translation, but faster.</p><p>What&#8217;s genuinely different is how structured design and product context, component code, and rules can be interpreted directly into coded, working interfaces. A prompt no longer has to start from nothing if it has access to the design system, APIs, patterns and engineering constraints.</p><p>And design becomes that context. A context for AI execution systems to use.</p><p>Teams are still going to need visual comparison and critique. They&#8217;ll need shared spaces to make business calls. The terminal window or an IDE is not a place for a lot of stakeholders to participate.</p><p>That doesn&#8217;t make the canvas central.</p><h2>Bring it back to the canvas</h2><p>When I look at Figma&#8217;s recent moves, they make sense. They build on its current strength.</p><p>More work should happen in Figma. More artifacts should come from Figma. Workflows should come back into Figma. More of the product development should be in the Figma ecosystem.</p><p>Reduced to its simplest form, the strategy seems to be:</p><p><em>Bring everything back to the canvas. Our canvas.</em></p><p>But the next era won&#8217;t be organized around that.</p><p>It&#8217;s why I thought &#8220;code-to-canvas&#8221; was pretty revealing. Make a real thing, then bring it back into Figma as editable frames.</p><p>That might solve a short-term collaboration problem. Directionally, it&#8217;s strange. Actually, it&#8217;s wrong. Wrong for the future, even if useful for Figma&#8217;s current position.</p><p>In that example, Figma is more worried about getting you back into their room - where they know how collaboration works. Less worried about whether that&#8217;s the right model of collaboration for the future.</p><h2>The canvas won&#8217;t be the source of truth</h2><p>Of course, Figma might be moving towards a more compelling future. One where Figma is a collaborative interface that reflects reality.</p><p>But it would be Figma as a lens.</p><p>Figma might be where you inspect your working systems. Compare variants. Annotate things that are real. See the design system drift. To steer and govern.</p><p>That might be valuable.</p><p>It also means accepting the canvas isn&#8217;t the center any more. And if it remains important, it only does so if it can be an interface to the truth.</p><p>The code, the runtime, what&#8217;s real, and what actually ships.</p><p>Figma&#8217;s danger seems to be trying to remain central by making everything pass through your old model.</p><p>That&#8217;s an incumbent trap.</p><p>That&#8217;s looking at what made you dominant in the first place, and only working to improve that thing. And that will be right...right up to the moment that the basis of competition changes.</p><p>Figma won against Sketch because it realized the center of gravity could change.</p><p>Now that center of gravity is changing again. And Figma is on the other side of the innovator&#8217;s dilemma.</p><h2>Execution is cheap. Coordination is not.</h2><p>AI makes execution cheaper.</p><p>Not free. But from a practitioner perspective, it can feel that way.</p><p>AI scaffolds the screens, uses the components, wires them up, refactors and gives us variants. It can create at a speed that changes all the old bottlenecks.</p><p>So the limiting factor is not &#8220;can we produce an interface?&#8221;. The limiting factor is &#8220;can you produce the <em>right</em> interface, with the right standards, for the right users, in a way our organization can trust?&#8221;</p><p>Coordination with AI assistance is not the same as collaboration in a canvas abstraction. We need structured and ranked context.</p><p>Which components are approved? Which patterns are deprecated? Which implementation is authoritative when the docs say one thing, the code says another, and Figma says a third? Which accessibility rules apply? Which regulatory constraints matter? Which engineering standards are non-negotiable?</p><p>That isn&#8217;t a canvas problem.</p><p>It&#8217;s an infrastructure problem.</p><p>Design systems are even more important in this world. Not as component libraries or asset stores, or even as docs for people to manually consult. They&#8217;re executable intelligence that tell AI systems how an organization builds.</p><p>The canvas is insufficient. It can arrange. It can invite critique. But unless it&#8217;s deeply connected to some control layer of product delivery, it risks becoming a pretty picture while the real thing lives elsewhere.</p><p>That&#8217;s a strategic problem.</p><h2>What Figma seems to believe</h2><p>From the outside, Figma seems to believe it can expand its canvas to contain the next era.</p><p>And, look, that may be unfair. It&#8217;s an external read of a company&#8217;s strategy. Figma is full of smart people, with every incentive to understand the shift. It may even be the smart commercial decision. That doesn&#8217;t make it the right product model for the next era of work.</p><p>Product strategy reveals posture. And Figma&#8217;s posture seems focused on a return to canvas.</p><p>Bring your generated work back. Bring your coded artifacts back. Bring your developers into Figma. Bring AI into the canvas.</p><p>Put more of your organization into the place Figma owns.</p><p>Which isn&#8217;t necessarily stupid. Enterprises have historically liked consolidation. People are familiar with Figma. And Figma has a gravitational pull from its market dominance.</p><p>Figma can keep adding useful capabilities.</p><p>Will those capabilities help Figma adapt to a world where the working artifact, and the organizational context, matter more than the design file?</p><p>Figma&#8217;s bet is: yes, because all of that will come back into Figma.</p><p>It&#8217;s a bet that the canvas is the core.</p><h2>And if we change the spaces where we work?</h2><p>I don&#8217;t think the next dominant product workspace will look like Figma with more AI features.</p><p>I don&#8217;t think it will look like a traditional design tool at all.</p><p>More likely, an IDE with some spatial collaboration. Or a browser-based product environment where live software is directly editable, inspectable, and deployable.</p><p>It will need to involve an AI orchestration layer that sits across design systems, repos, documentation, analytics, and product management tools.</p><p>Some integration of canvas, code editor, staging environment, governance and rules system.</p><p>It&#8217;s going to look bad at first.</p><p>Early versions of what&#8217;s right are going to look worse than mature versions of the past. Awkward, incomplete, and easy to dismiss.</p><p>Figma should understand this better than most. It won the last round because the future wasn&#8217;t just a better design tool, it was a different environment for the work.</p><p>The canvas may well remain essential. The canvas-as-abstraction will not.</p><p>The canvas needs to be a place to discuss reality, not flatten it.</p><p>That is a hard, interesting problem.</p><h2>What does Figma do next?</h2><p>I can think of three paths.</p><p>A defensive path is to continue to expand the canvas. Build to make more and more work happen inside Figma. That will definitely produce useful features. And it might produce strong revenue. Figma is dominant, and can become stickier and more embedded.</p><p>A second path is transitional. Make the canvas more code-aware, and more interactive. Better generation and better workflows. Better import and export. This seems to be where their current moves are. It&#8217;s really useful, but it still organizes around the canvas as the product environment.</p><p>Or it might accept that the canvas - and thus Figma - won&#8217;t be the center of truth. So they build to become one of the best collaborative interfaces into that truth.</p><p>That means treating code, product context, design systems, and live behavior as the actual work. And the canvas is just a view into that. A place where teams can reason in a visual way about a system that&#8217;s already alive.</p><p>I don&#8217;t know if Figma wants to make that pivot.</p><p>Strategic change isn&#8217;t necessarily about seeing into the future. It&#8217;s about having to give up on the assumptions that make the present business work.</p><p>Multiplayer design isn&#8217;t going to go away. It still matters.</p><p>The question is where that will live when we can generate, modify, review, and ship product much closer to code.</p><p>Figma understood the last change in the center of gravity. Now that center of gravity is moving again.</p><p>I&#8217;m curious whether Figma follows it.</p><div><hr></div><p><em>I&#8217;m not a neutral observer.</em></p><p><em>I&#8217;m VP of Product at Knapsack. We&#8217;re building in the place where structured design systems and product context meet AI-driven delivery.</em></p><div><hr></div><h4>Further reading:</h4><ul><li><p>Seiz, G. &amp; Kern, A. <em><a href="https://www.figma.com/blog/introducing-claude-code-to-figma/">From Claude Code to Figma: Turning production code into editable Figma designs</a></em>. Figma Blog, Feb 2026</p></li><li><p>Banfield, R. <a href="https://richardmbanfield.medium.com/digital-design-isnt-dead-it-just-got-way-more-interesting-befbdcf49324">Digital Design Isn&#8217;t Dead. It Just Got Way More Interesting</a>. Medium, Apr 2025.</p></li><li><p><a href="https://en.wikipedia.org/wiki/The_Innovator%27s_Dilemma">The Innovator&#8217;s Dilemma</a>. Wikipedia.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@krakograff?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Krakograff Textures</a> on <a href="https://unsplash.com/photos/a-close-up-of-a-wall-with-peeling-paint-FnDm9xq42bY?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[AI means your design system can't suck anymore]]></title><description><![CDATA[People patch the gaps. AI falls down the holes.]]></description><link>https://www.robin-cannon.com/p/ai-means-your-design-system-cant</link><guid isPermaLink="false">https://www.robin-cannon.com/p/ai-means-your-design-system-cant</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 02 Jun 2026 15:04:03 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8818ab73-cf3f-40a4-bd42-486e9fae6371_5786x3857.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Many so-called design systems are really component libraries with some human support wrapped around them.</p><p>A library is where you put your artifacts.</p><p>The humans provide the system.</p><p>People explain what the documentation misses. They can tell you why a component works that way. The problems with the official pattern, and why it hasn&#8217;t been fixed. They&#8217;re the emergency service for a team who can&#8217;t find the artifact or guidance that quite fits what they need.</p><p>Design system teams patch the gaps with critique. Slack threads. Office hours. Design reviews. The accumulated judgment of the organization.</p><p>We got away with that for a good while now.</p><p>It&#8217;s not ideal. Not efficient. But it&#8217;s workable.</p><p>AI makes it a lot less workable.</p><p>That&#8217;s <strong>not</strong> because AI needs something fundamentally new from a design system. It&#8217;s because it exposes a requirement we&#8217;ve been fudging all too often.</p><p>A system that doesn&#8217;t know how it wants to be used isn&#8217;t incomplete because AI arrived. It was already incomplete.</p><p>Its humans were just better at workarounds.</p><p>It&#8217;s why I&#8217;m skeptical of the idea that making design systems &#8220;AI-ready&#8221; is a new category of work.</p><p>AI <em>consumes</em> information differently than a person browsing a docs site. Markdown, frontmatter, metadata, and retrieval-friendly matter in a way they might not for a human reader.</p><p>That&#8217;s an important representation layer.</p><p>It doesn&#8217;t change the concept of a design system.</p><p>Design systems have always needed to explain more than just what exists. They need to explain how to use what exists, why, what it replaced, where it can flex, where it breaks, and what to do when something&#8217;s missing.</p><p>That&#8217;s design system maturity.</p><p>It just happens to parallel AI readiness.</p><p>The real shift is this:</p><div class="callout-block" data-callout="true"><p><em>AI won&#8217;t let you get away with building a mediocre component library, calling it a design system, and relying on human ingenuity to patch over the holes.</em></p></div><h3>Take away your components, and what have you got?</h3><p>Salesforce Lightning, Google Material, IBM Carbon. These systems didn&#8217;t arrive as immaculate frameworks, independent of existing product realities. They emerged out of large organizations that were already shipping at scale.</p><p>They codified familiar solutions, refined them, integrated them with existing ways of working.</p><p>The best systems grow out of prior knowledge.</p><p>At IBM, Carbon has components, tokens, and documentation. But what makes it strong is that it has a stance.</p><p>It explains how the system wants to be used. It makes decisions visible. It gives teams accessible components, but also explains how IBM thinks about accessibility.</p><p>What to avoid. Where to extend. How to think when the answer isn&#8217;t obvious.</p><p>I pointed an AI at Carbon when I was vibe coding a personal project. It got a solid result.</p><p>Carbon hasn&#8217;t been magically reimagined for AI. Carbon is powerfully explicit about how the system works.</p><p>The operating model is more important than the components.</p><p>Broader product and brand intelligence still matters too. A design system alone can&#8217;t tell AI what your company should build or what your product strategy is.</p><p>The design system has a narrower responsibility. How to make its own operating model legible.</p><p>That means decisions on what the system encourages. What makes a good extension? Which are the patterns we like and which do we tolerate? Why did we change that component? What&#8217;s the justification for that exception?</p><p>Humans usually want to ask those questions.</p><p>AI powers past all the missing answers.</p><h3>&#8220;Context-based&#8221; design systems, also known as &#8220;good&#8221; design systems</h3><p>So much of this &#8220;AI-ready design system&#8221; conversation misdiagnoses the problem. It&#8217;s as if &#8220;context&#8221; is some mysterious new discovery that we need to add for the machines to use.</p><p>Not the thing that separated a design system from a component library in the first place.</p><p>The system is not just its visible parts. It&#8217;s the reasoning that connects them.</p><p>TJ Pitre at Southleft describes part of this problem well in his writing on context-based design systems. His diagnosis is right: components and tokens alone are not enough. AI needs the context around the system, not just the artifacts inside it.</p><p>But this isn&#8217;t a new model for design systems.</p><p>It&#8217;s the old model.</p><p>Except now it&#8217;s being tested by a consumer that can&#8217;t quietly compensate for everything we failed to document.</p><h3>Plausible is not the same as good</h3><p>AI design system demos can produce something plausible when the room&#8217;s furnished. The framework. The components. The tokens. Codebase conventions it can imitate.</p><p>Plausible is not the same as correct.</p><p>Plausible means it holds together in that moment. Consistent spacing, the right component names. It passes the first glance test.</p><p>Good means the decision makes sense.</p><p>Good means the pattern fits the use case. That we understand the accessibility tradeoffs. Know that the exception in the implementation is intentional and justified. That it respects the system.</p><p>You can&#8217;t get that kind of quality from just your artifacts.</p><p>It depends on the receipts behind the artifacts.</p><p>People carry a lot of that context informally. In the heads of people who&#8217;ve been around long enough, who take the time to talk about it. In the old email threads that live long past their expiration date.</p><p>That&#8217;s worked better that it might have.</p><p>It&#8217;s also failed regularly.</p><p>This is how design systems drift. It is also how products drift away from design systems. It&#8217;s why you can build an inaccessible experience from 100% accessible components. The pieces are correct. The composition is not.</p><p>Someone might have the artifact, but not the context, the nuance or judgment to make the artifact useful.</p><p>AI accelerates that. A lot.</p><p>AI reaches for what&#8217;s visible. Uses the semantically close component. Follows a statistically likely pattern. It will make something that looks kinda aligned and completely miss the reason the alignment mattered.</p><p>That is a design system problem exposed by AI.</p><h3>Different format, same responsibility</h3><p>We don&#8217;t need to invent a separate discipline or domain around &#8220;AI design systems&#8221;.</p><p>We need to be more rigorous in doing the work that good systems always required. Then make that work consumable by machines as well as people.</p><p>Write guidance that explains usage, not just availability.</p><p>Document the reasons why, not just the end outcome.</p><p>Use examples as evidence of judgment.</p><p>Be explicit about exceptions.</p><p>Explain what to do when the system doesn&#8217;t have an easy answer.</p><p>Structure the content well. So that humans can read it and machines can retrieve it.</p><p>The design systems that handle AI well are ones that already understand what they&#8217;re for.</p><p>They will have a point of view. Patterns grounded in use. Teams that have captured not only what they shipped, but why it was worth sharing.</p><p>AI doesn&#8217;t make that a new requirement.</p><p>But it probably removes our ability to pretend the artifacts were ever enough.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Pitre, TJ. <em><a href="https://southleft.substack.com/p/context-based-design-systems-revisited">Context-Based Design Systems Revisited</a></em>. Slot Machine Substack, May 2026.</p></li><li><p>Whitehead, R. <em><a href="https://ioaglobal.org/blog/does-it-matter-ai-doesnt-understand-context/">Does It Matter If AI Doesn&#8217;t Understand Context?</a> </em>Institute of Analytics, Apr 2025.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@tonny_huang?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">tonny huang</a> on <a href="https://unsplash.com/photos/a-pile-of-boxes-that-are-sitting-on-the-ground-CNJAUQFRPps?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Don't get crabs]]></title><description><![CDATA[AI makes fine effortless. That's a problem.]]></description><link>https://www.robin-cannon.com/p/dont-get-crabs</link><guid isPermaLink="false">https://www.robin-cannon.com/p/dont-get-crabs</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 26 May 2026 15:02:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5512441c-c5cf-43c8-8f17-866058a49b79_4912x3264.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I was at a Knapsack Patterns event in Minneapolis recently. Lou Manning from ADP made a throwaway comment about how all AI applications will eventually become shadcn UIs.</p><p>That might be true.</p><p>Evolutionary biology has a concept called carcinization. Crustaceans that aren&#8217;t crabs will - independently - evolve into crab-like forms. They do it across lineages, and it&#8217;s happened multiple times through history.</p><p>Different species. Different environments. Same solution.</p><p>Crabs, as it turns out, are what you get when you optimize hard enough for a similar set of pressures.</p><div><hr></div><p>You might make the same observation about the application landscape. The same component libraries. Same sidebar navigation. Chat interface with an input field at the bottom. Muted palette, rounded corners, and empty states with friendly illustrations.</p><p>Nobody is copying anybody.</p><p>It&#8217;s a reaction to the same constraints. Foundation models producing the same outputs. And now we&#8217;re optimizing for the same LLM coding workflows. That creates a consistent pressure to ship fast, test cheap, and iterate immediately.</p><p>Same environment. Same selection pressures.</p><p>Same crab.</p><p>When everyone&#8217;s solving for the same things - speed, cost, LLM-friendliness - convergence isn&#8217;t a failure of imagination. It&#8217;s the logical outcome. And AI can make that convergence happen at speed.</p><div><hr></div><p>Carcinization is inevitable...if the selection pressures stay the same.</p><p>Crabs aren&#8217;t destiny. Crabs are what evolution produces to answer a specific question. If the question changes, the answer will too.</p><p>Teams building things that actually look different aren&#8217;t ignoring the constraints. But they&#8217;re adding at least one more - the need for something to be <em>good</em>. Not just fast and functional. And not just a matter of taste.</p><p>Good that requires someone to have made a decision about what good means.</p><p>AI can&#8217;t optimize that. It&#8217;s a human call.</p><p>Speed is table stakes. Cost is...if not free then converging to the point of consistency. The scarcity is wanting something specific that pushes against the path of least resistance.</p><div><hr></div><p>I wrote a piece last year about how AI changes the creative equation. Fast is free. Cheap is everywhere. So the only real differentiator left is good.</p><p>This is what that looks like in practice.</p><p>Yes, early AI adoption will make teams stand out. But AI will make <em>fine</em> effortless. The true differentiator is something ingenious, or newly imagined.</p><p>Convergent evolution is real. It happens anyway. AI means it&#8217;s coming for every product category, every interface pattern, every design system. And coming fast.</p><p>Time to decide whether you want to be a crab.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Hamers, L. <em><a href="https://www.scientificamerican.com/article/why-do-animals-keep-evolving-into-crabs/">Why Do Animals Keep Evolving into Crabs?</a> </em>Scientific American, Jun 2023.</p></li><li><p>Rizal, K. <em><a href="https://ai.gopubby.com/tyranny-of-smoothness-in-the-age-of-generative-ai-d60193df4902">Tyranny of Smoothness in the Age of Generative AI</a>. </em>AI Advances, Jan 2026.</p></li><li><p>Martignetti, T. <em><a href="https://www.fastcompany.com/91530169/ai-is-replacing-creativity-with-average">AI is replacing creativity with &#8216;average&#8217;</a>. </em>Fast Company, Apr 2026.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@mackenziejcruz?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Mackenzie Cruz</a> on <a href="https://unsplash.com/photos/red-and-black-crab-on-sand-V9ounv39B7k?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Your brand is an archaeological artifact]]></title><description><![CDATA[Why brand and design systems evolved apart - and why AI makes a fix even more urgent.]]></description><link>https://www.robin-cannon.com/p/your-brand-is-an-archaeological-artifact</link><guid isPermaLink="false">https://www.robin-cannon.com/p/your-brand-is-an-archaeological-artifact</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 19 May 2026 15:01:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/08cb8b0f-7f74-4dba-b4d2-1a58f76e65e2_5280x2970.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Where does a brand team&#8217;s work live?</p><p>Brand guidelines, brand standards. Decks sent to agencies. A Confluence page. Almost certainly a PDF or two.</p><p>Where does a design system team&#8217;s work live?</p><p>The component library. Tokens. The design system website. A Storybook instance. An npm package.</p><p>Two expressions of what a company is. Consistent and coherent in themselves. Often almost zero shared conversation.</p><p>It&#8217;s a strange structural failure in modern product organizations. I&#8217;m surprised it&#8217;s not talked about more often.</p><p>Brand sits in the marketing organization. Design systems sit somewhere in product, design, engineering organizations. They share a fundamental subject matter, but operate in organizational separation.</p><p>The results are predictable.</p><p>Design systems encode visual consistency without brand meaning. Brand guidelines describe how a company should feel...and nobody in product has read them.</p><p>The digital expression of the brand gets determined by whoever&#8217;s in the room.</p><p>Whose job actually is it to ensure that a product expresses the brand?</p><p>In most enterprises I&#8217;ve seen, the honest answer is nobody. Brand says what a company should feel like. Design systems say what the products should look like. And there&#8217;s a gap between the two where digital brand value can quietly disappear.</p><p>It&#8217;s an archaeological artifact.</p><p>Brand as an organizational function predates digital product as a discipline, by decades. And those structures have calcified. Brand in marketing, product in...product. And that&#8217;s held even after the product has become the primary brand experience for most companies.</p><p>Your brand is not the ad. Not the packaging. It&#8217;s the thing you use every day. But the org charts are already set.</p><p>I&#8217;ve seen the alternative. I built one of the few examples I see of how it can exist at scale. IBM Carbon doesn&#8217;t derive from component logic. It derives from a broader IBM design philosophy - that predates the system, runs deeper than any product services. But which defined a design language with awareness of, and collaboration with, the digital design system team.</p><p>The result is that I see Carbon in an IBM television ad or on a billboard. Or, more accurately, IBM products that use Carbon represent a coherent design language. Something that exists across mediums because it&#8217;s grounded in something universal.</p><p>That&#8217;s not the norm.</p><p>Brand guidelines evolved from print and campaign logic. Design systems evolved from digital product delivery logic. They haven&#8217;t evolved toward each other. Customers are the ones who absorb the incoherence.</p><p>This might have been manageable when product moved relatively slowly. Misalignment can be addressed (...or ignored!). Brand drift was visible enough for someone to notice before it compounds too much.</p><p>The faster delivery accelerates, the more difficult it is to manage. And we&#8217;re at the point of AI driving acceleration so that delivery might become almost incomprehensibly fast.</p><p>The models that let us ship in days instead of months will be generating interfaces, copy, and variations at a scale no brand team has ever accounted for. AI will drift because it has no context for what that brand is.</p><p>Design systems that carry genuine brand meaning - not just coherent visual rules, but the reasoning behind them - will compound in the right direction. Design systems that are sophisticated token libraries and nothing more will produce brand-incoherent experiences. Faster, and at greater volume, than before.</p><p>The AI case makes the structural fix urgent. But it was already necessary.</p><p>The product has been the brand for years.</p><p>The org chart just hasn&#8217;t caught up.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Weidemann, V (PhD). <em><a href="https://medium.com/@v_weidemann/the-unspoken-tension-product-vs-marketing-why-we-still-dont-speak-the-same-language-93b9f7e81b48">The Unspoken Tension: Product vs. Marketing &#8212; Why We Still Don&#8217;t Speak the Same Language</a></em>. Medium, May 2025.</p></li><li><p><em><a href="https://greygekko.com/integrated-product-and-brand-development/">Why do product and brand need to be developed together?</a></em> GreyGekko, Apr 2026.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@naeimj?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">naeim jafari</a> on <a href="https://unsplash.com/photos/an-aerial-view-of-a-city-with-a-lot-of-dirt-aDfL9xdyW8w?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Are friends electric?]]></title><description><![CDATA[Video killed the radio star. AI killed the GUI.]]></description><link>https://www.robin-cannon.com/p/are-friends-electric</link><guid isPermaLink="false">https://www.robin-cannon.com/p/are-friends-electric</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 05 May 2026 15:01:53 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/feb8e0e4-370b-49b9-ab20-14dfc4776743_7680x4320.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We&#8217;ve tried chatting with our computers. There was Clippy. My Amazon Alexa became the most over-engineered (and yet incredibly useful) voice activated cooking timer possible. But the dream of natural language as a primary interface kept arriving...and not quite working. Even the impressive demos were underwhelming in daily reality.</p><p>That&#8217;s shifted.</p><p>I interact with my computer almost entirely through conversation now. Not novelty. Not occasionally. It&#8217;s the primary mode.</p><p>I describe what I want in plain language. Despite the discomfort I mentioned in an essay a few weeks ago, I increasingly do so through voice as well as keyboard. Things get built. Files get created. Plans are written. Code is built. Systems get modified.</p><p>And then I look at what was made and decide whether it was right, or how to modify it. And we chat some more.</p><p>I don&#8217;t think I&#8217;m alone. In my circles, at least, feels like more people are operating this way. Directing tools through text and conversation, not clicking through interfaces that hold your hand.</p><p>GUIs didn&#8217;t disappear. But for a growing number of people they&#8217;re something you produce for others. They&#8217;re not where I operate from.</p><div><hr></div><h3>One of these things is not like the other</h3><p>Conversational interfaces are the mechanic.</p><p>What&#8217;s underneath that interface is not all the same.</p><p>I use Claude in a chat window. I use Claude Code in the terminal. In both cases I type, something responds, I evaluate the output. We have a conversation. The interaction pattern is recognizable across both of them.</p><p>Claude Code has system access. It writes files. It runs commands. It asks my permission to modify things. And if I say &#8220;yes, do that&#8221;, I&#8217;m not just approving the generation of a document. I&#8217;m approving an operation on my own system.</p><p>The &#8220;yes&#8221; can feel similar in Claude.ai as it does in Claude Code. The blast radius is not.</p><p>Approving an operation in a chat window isn&#8217;t the same as approving a system permission in a dialog box. And on one level we know that.</p><p>But there&#8217;s an older, visceral aspect to this. Conversation triggers social trust.</p><p>We trust things that communicate like people.</p><p>When Claude asks me ever-so-nicely whether it can have bash access, a part of my brain processes that like a colleague asking me for a favor. Not just a root access prompt.</p><p>When it&#8217;s just a dialog, I can dismiss it. When it&#8217;s a colleague, I can&#8217;t dismiss them.</p><p>Claude Code is genuinely the most capable thing I&#8217;ve used for building. That capability is the point. It&#8217;s exciting. But the same quality that makes it feel trustworthy - it&#8217;s fluent, reasonable, amenable - is exactly what makes the trust worth examining.</p><div><hr></div><h3>What&#8217;s going on?</h3><p>OK, let&#8217;s think about the first part first. Before we get cautious.</p><p>It&#8217;s really fucking cool that you can chat with your computer. Like actually chat to it, in ordinary language, and have it do useful stuff.</p><p>That&#8217;s new. That&#8217;s sci-fi made real for a lot of us. A genuine shift in how humans relate to machines. Maybe the biggest shift in interaction since the invention of the mouse and the window.</p><p>For all my life, computers needed me to learn their language to communicate. Commands, syntax, interfaces. So that the machine could parse it. We adapted to that tool.</p><p>Now...the tool adapts to me. It&#8217;s not perfect. It&#8217;s inconsistent. But so are people. At a very fundamental level, I think the direction of translation has reversed.</p><p>That&#8217;s not trivial. It&#8217;s not Clippy. It&#8217;s not a better search box.</p><p>It&#8217;s different.</p><div><hr></div><p>But there&#8217;s a question worth asking. I don&#8217;t have a clean answer to it yet.</p><p>What does it mean to approve operations that you don&#8217;t fully see? What does it mean to let something that feels like a real conversation make changes to systems that actually matter? How much do we care how the thing was built, if we ask for something and our agent makes it work?</p><p>We haven&#8217;t built the intuition to deal with that yet.</p><p>I think the interface arrived well before our instincts have had time to adjust.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Gibbons, S, et al. <em><a href="https://www.nngroup.com/articles/anthropomorphism/">The 4 Degrees of Anthropomorphism of Generative AI</a>. </em>NN/Group, Oct 2023</p></li><li><p>Milano, B. <em><a href="https://hls.harvard.edu/today/your-chatbot-may-be-the-friend-that-isnt/">Your chatbot may be the friend that isn&#8217;t.</a> </em>Harvard Law Today, Oct 2025.</p></li><li><p>Numan, G. <em><a href="https://www.youtube.com/watch?v=quAQqXX-7m0&amp;list=RDquAQqXX-7m0&amp;start_radio=1">Are &#8216;Friends&#8217; Electric?</a></em> Youtube (official), Jan 2020.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[A million useless MCPs]]></title><description><![CDATA[The protocol is the easy part.]]></description><link>https://www.robin-cannon.com/p/a-million-useless-mcps</link><guid isPermaLink="false">https://www.robin-cannon.com/p/a-million-useless-mcps</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 21 Apr 2026 15:03:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fb8424e1-2d9c-4a05-aafe-c00d7653888c_5760x3840.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Storybook has an MCP. Figma has an MCP. Zeroheight has an MCP. It&#8217;s a constant drip of messaging. Announcing it, leading with it.</p><p><em>We have an MCP.</em></p><p><em>We have an MCP.</em></p><p><em>We have an MCP.</em></p><p>Procurement teams talk about it as a qualifying feature. The word itself is talked about as if it&#8217;s a feature.</p><p>It isn&#8217;t one.</p><div><hr></div><p>I have an MCP. It has a cloud-based instance. You can hit it right now. It&#8217;ll pass all your tests. It&#8217;s fully compliant with MCP Inspector. Integrate it with any client.</p><p>It returns an empty object on every call. It has a zero toolset. </p><p>It&#8217;s useless. I built it as a thought experiment.</p><p>Because an MCP server is not a product. It&#8217;s a protocol.</p><p>A really useful protocol. One that this industry has converged on with exceptional speed. </p><p>But a protocol is a contract about how things plug into other things. It says nothing about what&#8217;s being plugged in. The protocol is an invitation, but if you accept it then you should have something to say.</p><div><hr></div><p>The MCP protocol doesn&#8217;t ask where the data comes from. It doesn&#8217;t say anything about what to do if sources contradict. It doesn&#8217;t say if its payload is authoritative or based on junk.</p><p>None of that is in the MCP spec. None of it has to be true to ship an MCP. But those are the only things that matter about whether that MCP is worth connecting to.</p><p>Having an MCP isn&#8217;t that much different from saying you have a GitHub repo. It can be empty, it can be filled with junk, it can be filled with excellence.</p><div><hr></div><p>I&#8217;ve argued at Knapsack to de-emphasize our MCP in our conversations.</p><p>The moment the term gets into the conversation, you get stuck in a comparative argument. A buyer can say &#8220;So does Storybook. So does Figma. What&#8217;s special about you having an MCP?&#8221; </p><p>At Knapsack I&#8217;ve been watching our IPE reconcile parallel implementations of a component, and surfacing where the disagreements are. Figuring out how sources coexist when they contradict. Normalizing those intakes into a consistent design system schema. </p><p>That&#8217;s the kind of work that&#8217;s valuable. We can serve that work through a protocol. But the protocol doesn&#8217;t have any opinions.</p><p>Products do.</p><div><hr></div><h4>Further reading:</h4><ul><li><p><em><a href="https://github.com/shinytoyrobots/compliant-empty-mcp">compliant-empty-mcp</a></em> on GitHub.</p></li><li><p><a href="https://modelcontextprotocol.io/docs/getting-started/intro">What is the Model Context Protocol?</a></p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@craftsmanconcrete_official?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Craftsman Concrete Floors</a> on <a href="https://unsplash.com/photos/empty-modern-warehouse-interior-with-polished-concrete-floor-3lkaszxWfGc?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA["You're lost, unless you have a rutter."]]></title><description><![CDATA[The Dutch East India company built an empire on knowledge. So can you.]]></description><link>https://www.robin-cannon.com/p/youre-lost-unless-you-have-a-rutter</link><guid isPermaLink="false">https://www.robin-cannon.com/p/youre-lost-unless-you-have-a-rutter</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 14 Apr 2026 15:01:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ccbd9261-2621-4261-a9b2-2b15f7c15eeb_6000x3791.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In <em>Shogun</em>, the Dutch ship pilots are the best in the world. Not just for being better sailors. Because they have better rutters. As James Clavell explains in the book:</p><blockquote><p><em>A rutter was a small black book containing the detailed observation of a pilot who had been there before. It recorded magnetic compass courses between ports and capes, headlands and channels. It noted the sounding and depths and color of the water and the nature of the seabed. It set down the how we got there and how we got back.</em></p></blockquote><p>But he adds something important to that description.</p><blockquote><p><em>But a rutter was only as good as the pilot who wrote it, the scribe who hand-copied it, the very rare printer who printed it, or the scholar who translated it. A rutter could therefore contain errors. Even deliberate ones. A pilot never knew for certain until he had been there himself. At least once.</em></p></blockquote><p>The Dutch understood this. They built better rutters. And they kept them secret. A library of the best routes to India, China, Japan, and the East Indies.</p><p>Maps they denied existed.</p><p>In 1611, Hendrik Brouwer charted a route between the Cape of Good Hope and Java that cut the voyage from a year to six months.</p><p>Their rivals sailed the same oceans. The Dutch got there in half the time, carried twice the cargo. And nobody outside the Dutch East India company knew exactly how.</p><p>That&#8217;s more than navigation. That&#8217;s a competitive moat.</p><div><hr></div><p>People are starting to realize that AI compounds what you feed it. Design systems, patterns, decisions - it produces at scale and speed that would have been inconceivable three years ago.</p><p>AI doesn&#8217;t know what context is worth compounding and what isn&#8217;t. Whatever you give it, it will amplify. Documented decisions and undocumented drift. Authoritative tokens, and a Figma file out of date with production. Enforced constraints, and those in one engineer&#8217;s institutional memory. All of it goes in.</p><p>Good context will compound into coherence at scale. Bad context into confident, fast-moving inconsistency. The AI isn&#8217;t lost, it&#8217;s just following charts that are wrong.</p><p><em>A rutter was only as good as the pilot who wrote it...</em></p><p>Most organizations haven&#8217;t verified theirs even once.</p><div><hr></div><p>A rutter wasn&#8217;t a single source. Historical rutters were compiled from multiple voyages, multiple pilots. Sometimes conflicting observations.</p><p>The value wasn&#8217;t just having the information. It was knowing which accounts to trust when they disagreed. The 1589 observation or the 1612 one. The pilot who navigated this strait in summer, versus the one in winter. The hint that was a lie.</p><p>Your product context is the same problem. Design system, yes. But brand guidelines, Figma libraries, components, content strategy, regulatory environment. Institutional memory about why one pattern exists and what happens if you overlook it. And the sources don&#8217;t always agree, so how does the AI navigate?</p><p>Organizations that integrate, orchestrate, and prioritize those sources - they&#8217;re building rutters. Every decision made is another voyage in the record. Every conflict resolution is a note to avoid a reef. Generation cycles that feed back into the system to make the next one more reliable.</p><p>It&#8217;s compounding advantage. Your competitors all have access to the same models. For all their genius, they&#8217;re commodities. But accumulated product knowledge, integrated, adjudicated, tells AI what <em>right</em> looks like for you. It&#8217;s the rutter your organization built, deliberately or not, for as long as you&#8217;ve been making decisions and writing them down. Its compounding nature means the earlier you build it, the harder it is to catch up.</p><p>The Dutch didn&#8217;t win the spice trade because their ships were faster. They won by taking the navigational system they built into proprietary, compounding knowledge. Then they protected that fiercely.</p><p>Same ocean. Same models.</p><p>You know what the difference was.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Clavell, J. <em>Sh&#333;gun</em>. Hodder &amp; Stoughton, 1975. </p></li><li><p>Bruijn, J. R. <em><a href="https://www.jstor.org/stable/20070358">Between Batavia and the Cape: Shipping Patterns of the Dutch East India Company</a></em>. Journal of South East Asian Studies, Sept 1980.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@genebrut?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Gene Brutty</a> on <a href="https://unsplash.com/photos/a-large-boat-floating-on-top-of-a-large-body-of-water-KfL1kTTrIKg?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Claude to Claude Code bridge]]></title><description><![CDATA[The key that gets you in the room]]></description><link>https://www.robin-cannon.com/p/the-claude-to-claude-code-bridge</link><guid isPermaLink="false">https://www.robin-cannon.com/p/the-claude-to-claude-code-bridge</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 31 Mar 2026 15:01:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ea420ca3-f1a9-4e85-b536-82cb81436702_5593x3621.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It took me a while to get my head round the idea that Claude and Claude Code live in separate environments. They&#8217;re both called &#8220;Claude&#8221;, right? But I didn&#8217;t really understand the practical implications until it started getting in my way.</p><p>I was out. Had my phone in my hand. And wanted to run a competitive analysis to see how Knapsack was addressing the market issue raised in an article I&#8217;d just read. Everything I needed was in my <code>cpo-skills</code> suite. Context, background data sources, routing tables to know which command to pull for which task.</p><p>All sitting in my vault. Perfectly organized. Inaccessible to me.</p><p>Because Claude Code lives in my CLI. My CLI lives on my laptop. My laptop was at home.</p><p>Claude.ai was right there. And it had far less that could help me.</p><p>It made me take a closer look at what claude.ai skills could do. My assumption - skills are prompts. A few hundred words providing context, available to trigger. Useful for tasks. But a long way from the depth I&#8217;d built in Claude Code.</p><p>I&#8217;ve built for depth. My <code>cpo-skills</code> suite has ten discrete command files, a delivery pipeline context document, and utilizes my specialized data-gathering agents wired to Linear, Slack, Github, etc. My <code>thought-leadership</code> suite carries its own context file - posting cadences, metrics, content pillars, conferences I&#8217;m tracking.</p><p>When Claude Code triggers these, that&#8217;s more than a prompt. It&#8217;s loading a working system.</p><p>But I realized that a claude.ai skill can be a gateway to that same system. On trigger, instead of containing all the intelligence itself, it instructs Claude to fetch it. It reads the README from my vault, loads the context file, consults the routing table, and pulls the specific command for the task at hand.</p><p>The files live in GitHub. Claude fetches them directly during the session. The skill in claude.ai is the key. The vault on GitHub is the room it opens.</p><p>My <code>cpo-skills</code> and <code>thought-leadership</code> skills in my Claude.ai project work in exactly this way. Each one is a few dozen lines. When I trigger the skill, it expands into the full suite I built in Claude Code. That&#8217;s context, routing logic, specialized commands, and in some cases an agent. All also available in that GitHub vault.</p><p>If you&#8217;re working in claude.ai you see a simple skill with a clear description. What runs is everything in the vault, accessed remotely.</p><p>This matters beyond just my own workflow. Claude Code solves for depth. That&#8217;s complex, stateful, multi-step work with persistent context. Claude.ai solves for accessibility. No terminal, no config files, far lower technical barrier.</p><p>The bridging pattern doesn&#8217;t collapse the distinction. The complexity stays in the vault, versioned and maintainable. The interface stays simple.</p><p>The person who builds the vault and the person who opens the door don&#8217;t have to be the same.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Salcan, Y. E. <em><a href="https://medium.com/@yunusemresalcan/claude-vs-claude-code-vs-cowork-which-one-do-you-actually-need-66d3952a2eb4">Claude vs Claude Code vs Cowork &#8212; Which One Do You Actually Need?</a> </em>Medium article, Feb 2026</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@alluntsyatko?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Alla Bila</a> on <a href="https://unsplash.com/photos/a-weathered-red-double-door-under-a-stone-archway-qys_X17KRRE?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for articles on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[We've always known the destination]]></title><description><![CDATA[A thirty-year detour to somewhere we knew we were meant to go]]></description><link>https://www.robin-cannon.com/p/weve-always-known-the-destination</link><guid isPermaLink="false">https://www.robin-cannon.com/p/weve-always-known-the-destination</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 24 Mar 2026 15:01:21 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b6658b15-a105-42ae-84ad-cdc930cb56d9_6000x3376.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There has always been one obvious destination for digital product delivery. A single interface where design intent and working reality are the same thing.</p><p>Not translated.</p><p>Not approximated.</p><p>Not handed off from one side and reconstructed on the other.</p><p>The same thing.</p><p>Our industry has been trying to build that for thirty years. We&#8217;ve gone from building bad versions of the right thing to building good versions of the wrong thing.</p><p>The jokes wrote themselves. Dreamweaver sites. FrontPage sites. If you were of a certain rebellious bent, HotDog sites. </p><p>These used to be mocked because they were the telltale sign of someone who didn&#8217;t really know what they were doing. Table-based layouts, inline styles, spaghetti markup so bloated that any developer would want to quietly rebuild the whole thing from scratch instead of fix it.</p><p>They were also empowering tools for a lot of people. They let you make stuff that was real. They weren&#8217;t being mocked because they were trying to unify the visual and the functional in a single interface. That instinct was right.  They were mocked because of how much they corrupted the code side of the equation.</p><p>The canvas was easy to navigate.</p><p>The code output was garbage.</p><p>So industry corrected. Realistically, given the technical limitations. Built an organizational culture around the separation of disciplines.</p><p>Serious designers used serious design tools.</p><p>Serious developers wrote serious code.</p><p>And between them a handoff ritual grew - redlines, specs, prototypes, tickets.</p><p>We created a workaround dressed up as a workflow.</p><p>The separation is artificial. We&#8217;ve always known this to some extent. Design systems were an obvious admission. Design intent encoded as structured, reusable truth rather than redrawn from scratch on every new screen. Tokens, components, semantic definitions; shared language that both sides could read. I&#8217;ve often joked that the irony of the name &#8220;design system&#8221; is that its primary consumers are usually developers.</p><p>AI closes the remaining distance. When structured design context can be interpreted directly into working interfaces, the translation layer becomes unnecessary. The middle dissolves, and the destination comes into view.</p><p>It&#8217;s why I find these code-to-canvas offerings so strange. Code-to-canvas takes a working interface - real interactions, data, behavior - and converts it back into static frames. </p><p>It argues that collaboration is only possible on drawings of the real thing, not the real thing itself.</p><p>Dreamweaver and FrontPage, for all their failures, at least understood where they needed to go. The visual and the functional needed to live together. They just didn&#8217;t have the technology to make their ambition real. The code they generated was the limitation, not the vision. </p><p>You can forgive a tool for being ahead of its time. But the technology exists now to make the canvas genuinely real - connected, live, executable. And it&#8217;s harder to forgive a deliberate turn away than a premature attempt at the right destination.</p><p>Our destination hasn&#8217;t changed. A canvas as a live interface into the system. The real thing made navigable, editable, collaborative. We&#8217;ve known that&#8217;s where we were going for a long time.</p><p>Surely this time.</p><div><hr></div><h4>Further reading</h4><ul><li><p><em><a href="https://www.robin-cannon.com/p/the-digital-workflow-is-obsolete">The digital workflow is obsolete</a></em>, on the collapse of the handoff model.</p></li><li><p><em><a href="https://www.robin-cannon.com/p/code-to-canvas-is-bonkers">Code to canvas is bonkers</a></em>, on Figma&#8217;s specific wrong turn.</p></li><li><p><em><a href="https://www.webmasterworld.com/html_editors/347.htm">FrontPage vs DreamWeaver</a></em>. Webmaster World.com discussion thread, July 2003.</p></li><li><p>Smith, E. <em><a href="https://tedium.co/2017/03/02/microsoft-frontpage-history-web-design-wysiwyg/">Your Code is Junky</a>.</em> Tedium, March 2017.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@d_mccullough?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Daniel McCullough</a> on <a href="https://unsplash.com/photos/an-architect-working-on-a-draft-with-a-pencil-and-ruler-HtBlQdxfG9k?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The design token cargo cult]]></title><description><![CDATA[How a useful tool can become dogma]]></description><link>https://www.robin-cannon.com/p/the-design-token-cargo-cult</link><guid isPermaLink="false">https://www.robin-cannon.com/p/the-design-token-cargo-cult</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 17 Mar 2026 15:01:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e64fab7a-7d14-466f-bc1e-f8472749686f_4032x3024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I was at IBM, Anna Gonzales was the thought leader on the token architecture for the Carbon Design System. It was tight, disciplined, and opinionated. And it worked. The structure encoded decisions the team had already fought for and made: what the system constrained, what it left open, where a component&#8217;s responsibility ended and a product team&#8217;s began. </p><p>You can read the system&#8217;s philosophy in how the tokens are organized.</p><p>The tokens aren&#8217;t what made the system work. The convictions are what made it work. The design tokens were an artifact of that conviction.</p><p>That distinction is the argument.</p><div><hr></div><p>The W3C Design Tokens Community Group published its first stable spec at the end of 2025. It&#8217;s genuinely good work. Years of collaboration to solve a hard coordination problem. How to share design decisions without everything fracturing every time someone changes a color.</p><p>The spec is deliberately minimal. It defines tokens, their types, and a reference system that lets one token point to another &#8212; so <code>color.text.primary</code> can reference <code>color.palette.black</code>, and changing the palette propagates everywhere. It adds <code>$extends</code> for group inheritance, a Color module with modern color space support, and a Resolver module for theming and context. </p><p>It&#8217;s a sensible, focused infrastructure that solves real problems. It&#8217;s the kind of foundation that would underpin exactly the discipline Anna built at IBM.</p><p>What practitioners have built <em>around</em> that foundation is something else. </p><p>The spec is agnostic about tiers. It defines how to express token relationships, not how many layers you should have. </p><p>The community is far less agnostic. Three-tier has become doctrine: primitive tokens at the base, semantic tokens that make purposeful claims about use, component tokens scoped to specific components. This is &#8220;mature token architecture&#8221; in most design systems discourse.</p><p>The gap between doctrine and practice is instructive. </p><p>Three-tier is what gets taught and recommended. Two tiers is what the major systems actually implement. </p><p>Carbon&#8217;s architecture doesn&#8217;t map neatly onto the primitive/semantic/component model - it has its own layering logic built around UI depth. Polaris has moved away from its component token layer. Material Design 3 publishes reference tokens and system tokens, and stops there.</p><p>Three-tier is the aspiration. Two-tier is what survives contact with a real system.</p><p>That gap should be a signal. The canonical systems couldn&#8217;t fully sustain the doctrine. And yet the doctrine keeps getting taught as the definition of maturity.</p><p>The problem isn&#8217;t the spec.</p><div><hr></div><p>The spec - unavoidably - creates <em>a thing to make</em>. &#8220;We&#8217;re implementing the W3C spec&#8221; can start to feel like a north star, when a real north star is missing.</p><p>At J.P. Morgan there was always a tension between a debate on token naming strategy and architecture coming before a simpler question was answered: <em>what is this design system for?</em></p><p>Naming debates aren&#8217;t a path to that clarity. They can be a replacement for it.</p><p>For Anna at IBM, the tokens were downstream. Philosophy came first. Tokens encoded the philosophy.</p><p>Where I&#8217;ve seen more struggle - at JPM, at some of the design systems I&#8217;ve worked with at Knapsack - is when that order is inverted. Taxonomy comes first, the thinking is supposed to emerge from it. Sometimes it does. Often the taxonomy becomes the only explicit structure of the system, and so it becomes load-bearing.</p><p>Which leads to a design system whose deepest held opinion is how to name its hover state.</p><p>Run naming conventions workshops because mature systems have naming conventions. Produce token JSON because good systems produce token JSON. That&#8217;s a cargo cult pattern. The mechanism becomes the mission.</p><div><hr></div><p>There is a failure mode you can identify: token counts scaling linearly with component complexity.</p><p>The Tetrisly design system acknowledged this problem. Their button component reached over 500 tokens. It enumerated every property of every state of every variant: background, border, text, icon, default, hover, focus, active, disabled, primary, secondary, danger, ghost, large, medium, small, dark mode, high contrast.</p><p>Before long you end up with <code>button-background-color-primary-large-hover-dark</code>, and hundreds of siblings.</p><p>The spec supports this. But at this point the abstraction provides no value over well-organized CSS. The overhead is real: tooling dependency, Figma sync, governance process. But there&#8217;s no additional leverage when your variable names map one-to-one to CSS properties you have to write anyway.</p><p>The promise of tokens is leverage - fewer, powerful constructs that express more than just flat specifics. 500 tokens is a clear failure of that promise - you may as well be writing CSS. Tetrisly acknowledged this, and they&#8217;ve very deliberately thinned their &#8220;component tier&#8221; so their model is much closer to a two-tier than three-tier model.</p><div><hr></div><p>Phillip Lovelace recently argued that tokens are even more important in an AI-driven workflow - token taxonomy can be an API the AI agent consumes. Semantic naming lets AI stop guessing your brand.</p><p>This is a worthwhile floor argument. AI generating UI from a token file produces more consistent output than generating from nothing. The W3C spec makes that even more reliable.</p><p>A floor isn&#8217;t a ceiling.</p><p>AI can traverse a token graph and resolve a name to a hex value. It can&#8217;t tell you why that color is right for a primary hover state, or if a destructive action should use the same token. Should a payment confirmation defer to stricter contrast constraints?</p><p>Those aren&#8217;t AI limitations. </p><p>Those are limitations of the information that tokens are supposed to carry.</p><p>Tokens encode what things look like. Not why. Not when. Not the conditions that change the answer.</p><p>The convictions in the best systems come from the decisions that precede them. An AI with access to those decisions - rules, intent, context - can do more interesting things than resolve color aliases.</p><div><hr></div><p>Tight token architecture delivers real value. Tokens are a powerful artifact of thinking, but not a substitute for it. The W3C spec describes something of genuine worth, when it&#8217;s built in the right order.</p><p>The design systems that work treat tokens as output. Philosophy first, constraints second, governance third. Tokens encode the decisions that have been made. But only <em>if </em>those decisions have been made. </p><p>Systems that struggle have the sequence backwards. And the quality of the spec actually makes the inversion easier. It&#8217;s a rigorous blueprint for the mechanism, and the foundation is left implicit.</p><p>Tokens with system conviction are infrastructure. Tokens that substitute for it are dogma. The difference is everything.</p><div><hr></div><h4>Further reading:</h4><ul><li><p>Frost, B. <em><a href="https://bradfrost.com/blog/post/the-many-faces-of-themeable-design-systems/">The Many Faces of Themeable Design Systems</a></em>. bradfrost.com</p></li><li><p>Gonzales, A. <em><a href="https://medium.com/carbondesign/introducing-figma-variables-and-a-consolidated-all-themes-library-d4893d1b8920">Introducing Figma variables and a consolidated &#8220;All themes&#8221; library!</a> </em>Carbon Design Blog, Aug 2023.</p></li><li><p><em><a href="https://www.designtokens.org/tr/2025.10/">Design Tokens Technical Reports</a></em>. W3C Community Group, Oct 2025.</p></li><li><p>Lovelace, P. <em><a href="https://www.designsystemscollective.com/design-systems-are-having-their-moment-70674f8ab197">Design Systems Are Having Their Moment</a></em>. Design Systems Collective, Feb 2026.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@jcanty123?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Jan Canty</a> on <a href="https://unsplash.com/photos/a-wooden-structure-sitting-on-top-of-a-rocky-beach-bz-FrwVCLDc?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[A design system isn't an aggregator. It's a contract.]]></title><description><![CDATA[Tools can aggregate assets. They can't make a system real.]]></description><link>https://www.robin-cannon.com/p/a-design-system-isnt-an-aggregator</link><guid isPermaLink="false">https://www.robin-cannon.com/p/a-design-system-isnt-an-aggregator</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Tue, 24 Feb 2026 16:03:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c60dcfc2-d84a-4e4b-b97d-0139f899a323_5472x3468.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week, Mariana Rita published &#8216;<em>Stop paying for design system documentation you can build yourself&#8217;</em>.</p><p>Her argument is that paid design system documentation platforms like Zeroheight are obsolete. Figma&#8217;s API is accessible. LLMs can write docs. Open-source tooling is good. Build the whole thing yourself, in two weeks, for free.</p><p>It&#8217;s a practical guide. It has sound tooling advice.</p><p>The piece treats design system documentation like it&#8217;s an asset display problem. How do you surface your Figma components and code repos in one place, more cheaply than current products offer? That&#8217;s a reasonable question.</p><p>But it&#8217;s based on a definition of design systems that I think is fundamentally too shallow. That&#8217;s a starting point that leads to something that looks like a solution, but isn&#8217;t.</p><div><hr></div><p>The article describes documentation as an &#8220;aggregator - the one place that ties everything together.&#8221; </p><p>Sit with that word.</p><p>An aggregator is passive. It collects. It displays. It points at sources and renders them side by side.</p><p>A design system isn&#8217;t an aggregator.</p><p>It&#8217;s a contract.</p><p>A design system is the authoritative agreement between disciplines - design, engineering, product - about what is true, what is intentional, and why. Documentation doesn&#8217;t just display agreement, it&#8217;s where that agreement becomes canonical. Intent becomes instruction. &#8220;We discussed this on Slack&#8221; becomes &#8220;this is how we build things.&#8221;</p><p>If your docs platform is a viewer on top of your Figma files and your code repos, that&#8217;s not a system of record. It&#8217;s a window. It doesn&#8217;t resolve disagreements.</p><div><hr></div><p>A DIY doesn&#8217;t solve one, vital problem - and the article doesn&#8217;t name it.</p><p>Figma is not a design system.</p><p>Storybook is not a design system.</p><p>Figma is a tool for designers. Their working files, experiments, intentions, abandoned interations. A designer&#8217;s environment - rich and full of things in-progress, deprecated, aspirational.</p><p>Storybook is a tool primarily for developers. It documents what has been implemented. It&#8217;s the engineering team&#8217;s environment - authoritative about code, somewhat indifferent to design rationale.</p><p>If you point an AI agent at your Figma library you&#8217;ll train it on your designers&#8217; hypotheses. Point it at your Storybook and you train it on engineering implementation that may lag design intent, exceed it, or quietly diverge.</p><p>Neither tool can answer the question on its own: <em>what is true?</em></p><p>When the Button component in Figma has rounded corners and the one in Storybook doesn&#8217;t, what does your documentation site say? If it just faithfully renders both sources, it&#8217;s documented your misalignment.</p><p>That&#8217;s not nothing. It&#8217;s useful to know. But it&#8217;s not a source of truth. It&#8217;s published disagreement.</p><p>The design system has to adjudicate. It has to carry a philosophical underpinning. A stance. Not just visual and technical inventory. It needs the decisions and reasoning that make the inventory coherent. Why did we make these choices? What are the governing principles? Which source wins when there&#8217;s a discrepancy? Why? </p><p>And how do we strive for excellence when there isn&#8217;t a source at all? A new pattern, an edge case, a platform you haven&#8217;t built for yet. The system of record doesn&#8217;t just arbitrate what exists. It also guides what should.</p><p>That&#8217;s not something that can be generated. It requires human judgment, authority, and a platform to enforce it.</p><div><hr></div><p>The article&#8217;s AI-readiness argument is sharp. And it&#8217;s where the initial error becomes most consequential.</p><p>The article is right to identify that design system documentation is increasingly an instruction layer for AI agents. &#8220;Robot food&#8221; as my colleague Chris Bloom would describe it. The context that makes generated UI consistent and correct. And it&#8217;s also right that static platforms unable to expose data in a structured, machine-readable way, fail at this job.</p><p>So it proposes replacing them with Docusaurus and MDX files maintained by a team. Which is also static. And manually curated. But it&#8217;s free and you own it.</p><p>But the answer to AI-readiness isn&#8217;t just a better documentation site. It&#8217;s a genuine system of record. Where documentation is generated from structured, authoritative, interconnected sources. Where the connection between intent and implementation is dynamic, not periodically reconciled.</p><p>Documentation that auto-updates when code changes isn&#8217;t a feature. It&#8217;s the entire point. </p><p>But it needs to have broader context to update in an intelligent, guided, way.</p><p>Otherwise you&#8217;re just feeding AI a snapshot. A snapshot of what was true when someone last updated a file. Or a snapshot of the AI&#8217;s guess at what a conflict resolution was. And the gap between &#8220;what the docs say&#8221;, &#8220;what the design file says&#8221; and &#8220;what is in production&#8221; is exactly the kind of ambiguity that makes AI-generated interfaces drift.</p><div><hr></div><p>The cost argument also dissolves. </p><p>The article compares platform licensing fees to zero. </p><p>But the real cost of a design system isn&#8217;t tooling. It&#8217;s the misalignment it prevents - or fails to prevent.</p><p>That&#8217;s the denominator.</p><p>The cost of rework because design and engineering interpreted a component differently. Cost of multiple QA cycles because the implementation didn&#8217;t match the spec. Cost of onboarding time because the documentation was out of date. Cost of inconsistent experiences because there wasn&#8217;t an authoritative answer to &#8220;how does this pattern work in iOS?&#8221;</p><p>The total cost of a DIY aggregator includes engineering time to build it, maintain it, update it when APIs change, wrangle the AI writing pipeline, manually curate the output. And, if it&#8217;s being seen as an aggregator, it includes the organizational cost of having a documentation site that&#8217;s a collection of assets rather than a system of authority.</p><p>That cost might seem invisible. Then it accumulates.</p><div><hr></div><p>None of this is an argument against open-source tooling. Or AI-assisted documentation. Or against genuine improvements in what&#8217;s accessible and buildable. Those are real changes.</p><p>Infrastructure isn&#8217;t neutral.</p><p>The choice of what you build - aggregation layer or system of record - has downstream consequences for every discipline that depends on it. It shapes what designers trust. What engineers implement. What AI agents consume. What your product becomes.</p><p>If self-building (not merely aggregating) your design system platform is the right approach, and the costs and benefits are fully considered, great. That&#8217;s IBM Carbon, and it&#8217;s one of the best design system websites out there.</p><p>And if a documentation platform can make the building of the site easier, so you can focus on the creation of the actual design system, all the better.</p><p>Design systems are not a collection of Figma components and Storybook stories, with a documentation site sitting on top. The documentation is the surface expression of something much deeper and more considered. </p><p>Decisions made. Rationale captured. Authority established.</p><p>You can build an aggregator in two weeks. Building a system of record takes longer. Because you have to decide what&#8217;s actually true.</p><p>That&#8217;s the work.</p><div><hr></div><p><em>I&#8217;m not a neutral observer.</em></p><p><em>I&#8217;m VP of Product at <a href="https://www.knapsack.cloud/">Knapsack</a>. We build infrastructure that makes design systems a live system of record - connecting design, code, and documentation as a unified source of truth.</em></p><p><em>I have a direct interest in this question, and you should read with that in mind. But the argument stands regardless.</em></p><div><hr></div><h4>Further reading:</h4><ul><li><p>Rita, M. <em><a href="https://medium.com/all-about-design-systems/stop-paying-for-design-system-documentation-you-can-build-yourself-a10f1390987f">Stop paying for design system documentation you can build yourself.</a></em> All about design systems, Feb 2026.</p></li><li><p>Aizlewood, J. <em><a href="https://clearleft.com/thinking/design-">Design systems don&#8217;t start with components.</a></em> Clearleft, July 2017.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@skillscouter?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Lewis Keegan</a> on <a href="https://unsplash.com/photos/text-XQaqV5qYcXg?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Code to canvas is bonkers]]></title><description><![CDATA[Figma's latest feature solves Figma's problem, not yours]]></description><link>https://www.robin-cannon.com/p/code-to-canvas-is-bonkers</link><guid isPermaLink="false">https://www.robin-cannon.com/p/code-to-canvas-is-bonkers</guid><dc:creator><![CDATA[Robin Cannon]]></dc:creator><pubDate>Wed, 18 Feb 2026 16:02:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/01c34f06-ba20-4823-9447-8cfc15c0b64c_6240x4160.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Dylan Field posted about a new Figma feature: the ability to bring work from Claude Code directly into Figma&#8217;s canvas. Capture a working UI - something that already works in production, staging, or localhost - and convert it into editable Figma frames.</p><p>Let&#8217;s break that down. You use an AI coding tool to generate a working interface. Real code, real interactions, and real data. Then take the functioning reality and convert it <em>back</em> into a static abstraction, so that people can look at it together on that canvas.</p><p>You had a building. Now you have an architect&#8217;s drawing of that building.</p><p>Figma is telling you this is progress. It isn&#8217;t. It&#8217;s a solution to Figma&#8217;s business problem - the growing irrelevance of the static canvas - presented as though it&#8217;s a solution to yours.</p><div><hr></div><p>In November I wrote about the collapse of the traditional digital workflow. That comfortable fiction where design happens here, code happens over there, and a handoff ritual connects the two. </p><p>AI is dissolving that middle layer. Structured design systems, and wider product context, can be interpreted into working interfaces. At that point, the abstraction layer between design intent and delivery reality becomes unnecessary. Design isn&#8217;t a stage before delivery. Design <em>is</em> delivery.</p><p>It&#8217;s a defining shift in how products get built. </p><p>Jonny Burch made a complementary argument last week, in his piece &#8220;Life After Figma is Coming.&#8221; His framing encompasses the tooling ecosystem: as code becomes the source of truth, design tools become interfaces on code, not the other way around. Code is the only correct source of truth - open, shared, with common standards. The canvas doesn&#8217;t disappear, but it has to become a view into reality and not a substitute for it.</p><p>These arguments point in the same direction. The canvas is no longer a staging ground for ideas that exist outside the system. It needs to be a live interface for ideas that exist within it.</p><p>Claude Code to Figma points in the opposite direction entirely.</p><div><hr></div><p>Field&#8217;s own framing is revealing. He describes working code as &#8220;tunnel vision&#8221; and that this feature will help you &#8220;escape&#8221; it. </p><blockquote><p>&#8230;the design canvas is better at navigating lots of possibilities than prompting in an IDE.</p></blockquote><p>The Figma blog announcing the Claude Code to Figma feature elaborates further. Solo code exploration is a &#8220;single-player environment&#8221;.</p><blockquote><p>&#8230;that speed of solo exploration can become a constraint.</p></blockquote><p>And as a contrast, they describe the canvas as a &#8220;shared space&#8221; where &#8220;the conversation changes and new possibilities open up.&#8221;</p><p>That&#8217;s a neat rhetorical move.</p><p>It takes a genuine limitation of current code-first workflows - collaboration and visual comparison are harder in a terminal - but reframes it so that the <em>working artifact</em> is the problem, and the <em>abstraction</em> is the solution. </p><p>Those are not the same thing. Needing better collaboration on working artifacts isn&#8217;t the same as needing to convert the artifacts into a different, lower-fidelity format in order to collaborate at all.</p><p>The answer to &#8220;how do we collaborate on code?&#8221; is not &#8220;convert it to not-code.&#8221;</p><p>The better solution is better collaboration tooling for code. Live preview sharing. Annotation layers on running applications. Structured feedback on deployed states. As Burch points out, a tooling ecosystem is already exploding in size - design interfaces that sit on top of production code, development environments that integrate design thinking. The pieces are coming together.</p><p>When you capture something from Claude Code and bring it into Figma you don&#8217;t add information. You remove it. You remove interactions, real data, actual behavior. You replace it with a picture of what it looked like at the moment of capture. </p><p>That&#8217;s trading truth for convenience, and claiming it&#8217;s a workflow improvement.</p><p>It&#8217;s an absurd proposition.</p><p>You have a working thing. You convert it into a non-working representation. You collaborate on the representation. At some point, presumably, someone has to make into a working thing again. That&#8217;s not a workflow, that&#8217;s a detour.</p><div><hr></div><p>This is not just a feature decision. It&#8217;s a strategic posture that serves Figma&#8217;s interests while actively working against the interests of people using it.</p><p>Figma&#8217;s business depends on it being the place where product decisions happen. Their value proposition is the collaborative canvas being the hub of product development. The more decisions happen in code-first environments - engineers and designers collaborating directly on running applications - the more Figma risks becoming peripheral.</p><p>Figma is trying to maintain its gravitational pull. Every feature needs to bring work <em>into</em> Figma, not enable work to happen <em>outside</em> it. </p><p>Field says this directly:</p><blockquote><p>Whether product building begins in a terminal, a prompt box, a visual UI or a hand-drawn sketch, we want Figma to be the place where it all comes together.</p></blockquote><p>That&#8217;s not a workflow insight. That&#8217;s a business objective.</p><p>Claude Code to Figma is not about improving your workflow. It&#8217;s about preventing your workflow from leaving Figma behind.</p><p>And this is the part of the framing that I find genuinely dishonest. Figma&#8217;s blog post presents this as a way to unlock creativity and open up collaboration. Something to liberate teams from the constraints of solo code exploration. But the people who are using Claude Code to build working interfaces <strong>aren&#8217;t constrained</strong>. They are <strong>ahead</strong>. They have the real thing. The feature asks them to go backwards. Sacrifice fidelity, leave something functional and return to an abstraction layer. Because Figma needs them to.</p><p>That&#8217;s a solution for Figma. It&#8217;s not a solution for the people building products.</p><p>The timing might reinforce the point. Figma went public, and the stock has dropped a lot since IPO. It&#8217;s a legitimate question whether AI-native tools make traditional design canvas less central to digital product development. So product announcements also need to be a message to investors: <em>the canvas is still essential</em>. </p><p>Code-to-canvas is a defensive move dressed up as innovation.</p><div><hr></div><p>The honest version of this feature announcement might say: &#8220;More product work is starting in code. We need to pull that back into our ecosystem so we can remain relevant.&#8221; That&#8217;s a legitimate business challenge. And I have sympathy for the difficulty of Figma&#8217;s position. They built something genuinely great, and the ground is shifting beneath it.</p><p>But don&#8217;t tell me it&#8217;s for my benefit. The loss of fidelity isn&#8217;t &#8220;opening up new possibilities.&#8221; A functioning prototype being flattened into a picture isn&#8217;t &#8220;changing the conversation&#8221;. The conversation was already happening - and in a richer, more honest medium. This feature interrupts the conversation to bring it back to Figma.</p><p>It&#8217;s not about fewer canvases. It&#8217;s about more honest ones. More real. Connected to live systems, reflecting real state. Enabling collaboration on the actual artifact rather than a simulation. The future of the canvas is as an interface to code, not a destination to convert code back to.</p><p>Code to canvas is the wrong direction. The future is canvas as code. Features like this are going to look increasingly strange as the rest of industry figures that out.</p><div><hr></div><p><em>I&#8217;m not a neutral observer.</em></p><p><em>I&#8217;m VP of Product at Knapsack. We&#8217;re building in the place where structured design systems and product context meet AI-driven delivery. </em></p><p><em>But the proximity is also why this absurd code-to-canvas direction is so acute for me. When you&#8217;re working on systems that make design directly executable, watching someone propose converting execution back into abstraction feels like someone printing out a Google Doc so that they can fax it.</em></p><p><em>I'm presenting an expanded version of these ideas at the <a href="https://developersummit.com/session/the-digital-workflow-is-obsolete-how-to-survive-the-end-of-the-canvas">Great International Developer Summit</a> in April 2026.</em></p><div><hr></div><h4>Further reading:</h4><ul><li><p><em><a href="https://www.robin-cannon.com/p/the-digital-workflow-is-obsolete">The digital workflow is obsolete</a></em> - the end of abstraction, and the start of design as delivery.</p></li><li><p>Burch, J. <em><a href="https://jonnyburch.com/life-after-figma/">Life after Figma is coming (and it will be glorious)</a></em>. jonnyburch.com, Feb 2026.</p></li><li><p>Seiz, G. &amp; Kern, A. <em><a href="https://www.figma.com/blog/introducing-claude-code-to-figma/">From Claude Code to Figma: Turning production code into editable Figma designs</a></em>. Shortcut - Figma&#8217;s editorial newsletter, Feb 2026.</p></li><li><p>Field, D. <em><a href="https://www.linkedin.com/pulse/claude-code-figma-design-dylan-field-e5ilc/?trackingId=aca2hHS1Q2OmpfNYVAIeWA%3D%3D">Claude Code to Figma Design</a></em>. LinkedIn, Feb 2026.</p></li><li><p>Flowers, E. <em><a href="https://eflowers.substack.com/p/if-you-ask-a-designer-what-they-want">If you ask a designer what they want, they will say faster horses</a></em>. Zero Vector, Feb 2026.</p></li></ul><p><em>Article photo by <a href="https://unsplash.com/@version2beta?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Rob Martin</a> on <a href="https://unsplash.com/photos/red-and-white-stop-sign-tte1gbfGEeY?utm_source=unsplash&amp;utm_medium=referral&amp;utm_content=creditCopyText">Unsplash</a>.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.robin-cannon.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe for essays on design, technology, and culture - plus original fiction.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item></channel></rss>