<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Long Commit]]></title><description><![CDATA[Weekly essays for engineering leaders and senior engineers building better judgment, stronger teams, and durable careers.]]></description><link>https://newsletter.thelongcommit.com</link><image><url>https://substackcdn.com/image/fetch/$s_!sRAG!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f8ce71-7589-4ce9-9561-5b634cdc4ae1_1024x1024.png</url><title>The Long Commit</title><link>https://newsletter.thelongcommit.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 21 Jul 2026 00:48:50 GMT</lastBuildDate><atom:link href="https://newsletter.thelongcommit.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Juan Cruz Martinez]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[thelongcommit@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[thelongcommit@substack.com]]></itunes:email><itunes:name><![CDATA[Juan Cruz Martinez]]></itunes:name></itunes:owner><itunes:author><![CDATA[Juan Cruz Martinez]]></itunes:author><googleplay:owner><![CDATA[thelongcommit@substack.com]]></googleplay:owner><googleplay:email><![CDATA[thelongcommit@substack.com]]></googleplay:email><googleplay:author><![CDATA[Juan Cruz Martinez]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The New Engineering Lead’s Guide to Time Management]]></title><description><![CDATA[A practical guide to protecting your calendar, delegating earlier, making better decisions, and creating space for your team to do good work.]]></description><link>https://newsletter.thelongcommit.com/p/the-new-engineering-leads-guide-to</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-new-engineering-leads-guide-to</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Thu, 16 Jul 2026 15:23:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9Pm3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9Pm3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9Pm3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 424w, https://substackcdn.com/image/fetch/$s_!9Pm3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 848w, https://substackcdn.com/image/fetch/$s_!9Pm3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 1272w, https://substackcdn.com/image/fetch/$s_!9Pm3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9Pm3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png" width="1456" height="776" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:776,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:406566,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207258513?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9Pm3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 424w, https://substackcdn.com/image/fetch/$s_!9Pm3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 848w, https://substackcdn.com/image/fetch/$s_!9Pm3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 1272w, https://substackcdn.com/image/fetch/$s_!9Pm3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fcd0ab-8cee-43d3-aa9f-42e4e71c2619_2440x1300.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For <a href="https://newsletter.thelongcommit.com/p/the-case-for-becoming-a-manager">most of my career, I was an individual contributor</a>. Usefulness was visible: write the code, review the change, fix the problem, publish the document. The output was easier to see.</p><p><a href="https://newsletter.thelongcommit.com/p/staying-technical-as-a-tech-manager">Moving into management changed that relationship</a> before I had a good way to think about the calendar. A 1:1, a proposal review, a stakeholder question, and an escalation can all be legitimate work. Put enough of them together, though, and the week disappears before the decisions and conversations that require a lead have room.</p><p>The transition problem starts when leadership work arrives before old IC work leaves. Implementation, routine review, technical coordination, and familiar context remain attached through habit; people, delivery, feedback, stakeholders, and risk arrive on top. This is not a time-management problem in the narrow sense. It is a responsibility problem expressed through the calendar.</p><p>This guide is for new engineering managers and technical leads whose outcomes increasingly depend on other people while they still carry part of their old role. Their authority differs, but the practical question is the same: which work genuinely requires your attention, and which work is still yours because it used to be?</p><p>My operating rule is simple: when leadership work enters your week, work from your old job needs to leave. A transition may require temporary overlap, but temporary overlap needs an end date. Otherwise it quietly becomes two permanent jobs.</p><h3>Audit the week you actually have</h3><p>For an initial audit, do not start by designing an ideal calendar. Start with the previous ten working days.</p><p>Open your calendar, Slack, notes, and whatever you use to track work. Account for visible commitments and the work around them: preparing feedback, reviewing proposals, answering questions, writing follow-ups, recovering context between meetings, handling an incident, or finishing something after hours because it never found space during the day.</p><p>You do not need a new time-tracking system. The goal is not minute-by-minute surveillance. You are looking for enough evidence to answer five questions:</p><ol><li><p>What took meaningful time?</p></li><li><p>What outcome did it create?</p></li><li><p>Why did it require me?</p></li><li><p>Who waited for me?</p></li><li><p>Should I own it, stay close to it, transfer it, or stop it?</p></li></ol><p>A simple audit might look like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qro6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qro6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 424w, https://substackcdn.com/image/fetch/$s_!Qro6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 848w, https://substackcdn.com/image/fetch/$s_!Qro6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!Qro6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qro6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png" width="1456" height="444" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:444,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:343546,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207258513?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qro6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 424w, https://substackcdn.com/image/fetch/$s_!Qro6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 848w, https://substackcdn.com/image/fetch/$s_!Qro6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!Qro6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83e061ab-f707-4d2f-aa92-438b951aa313_3538x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The &#8220;Why me?&#8221; column is uncomfortable because familiarity can masquerade as responsibility. You may be the fastest reviewer because you know the codebase. You may be the best person to rewrite a plan because you have stronger opinions about its structure. Neither fact automatically means the work should remain yours.</p><p>Look for three forms of calendar pressure:</p><ul><li><p><strong>Hidden leadership work:</strong> feedback, tradeoff decisions, stakeholder preparation, and thinking that gets pushed outside working hours.</p></li><li><p><strong>Habitual IC work:</strong> implementation, routine review, or technical coordination you still own because you used to own it.</p></li><li><p><strong>Recurring demand:</strong> meetings, approvals, and repeated questions that keep returning because the underlying ownership or context is unclear.</p></li></ul><p>At the end of the audit, name the three biggest sources of pressure. Do not fix all of them yet. The audit separates what feels busy from what is making the team dependent on you.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;9bae3f04-927a-4e5a-9885-dc9e0257506b&quot;,&quot;caption&quot;:&quot;The question of whether experienced engineers should move into management has been on my mind for a while. Not as an abstract career question, but as something I&#8217;ve lived through. I made the switch last year and I&#8217;ve been turning over what I learned from that decision ever since. I kept putting off writing about it because the topic is genuinely complic&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The Case for Becoming a Manager&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:178170948,&quot;name&quot;:&quot;Juan Cruz Martinez&quot;,&quot;bio&quot;:&quot;20+ years in software. I write The Long Commit for senior engineers and tech leaders who want better judgment, clearer decisions, and a longer-term edge in software, leadership, and their careers as AI reshapes the industry.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09f78b54-4163-4e29-98d9-5afd264395b6_3200x3200.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-03-24T10:58:11.959Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/05c70c7d-e58f-49e4-9aba-a26338d1dc1c_3440x1920.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://newsletter.thelongcommit.com/p/the-case-for-becoming-a-manager&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:191921172,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:6,&quot;comment_count&quot;:0,&quot;publication_id&quot;:8210189,&quot;publication_name&quot;:&quot;The Long Commit&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sRAG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f8ce71-7589-4ce9-9561-5b634cdc4ae1_1024x1024.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h3>Stop carrying your old job</h3><p>Push each recurring responsibility to the lowest level of involvement that is still safe. Then sort it into four groups.</p><p><strong>Own.</strong> Work that requires your authority, accountability, or sensitive judgment: a performance conversation, a priority conflict, a cross-team commitment, a high-risk escalation, or a decision where the team needs one accountable owner.</p><p><strong>Stay close.</strong> Work where you need enough context to exercise judgment without controlling every step: an important launch, a migration with operational risk, architecture with long-term consequences, or a customer problem that could change team priorities.</p><p><strong>Transfer.</strong> <a href="https://newsletter.thelongcommit.com/p/the-case-for-becoming-a-manager">Give another person a complete outcome</a> when the work can develop their judgment and reduce unnecessary dependence on you: routine reviews, project coordination, implementation planning, recurring documentation, or a technical proposal.</p><p><strong>Stop.</strong> Meetings without an output, duplicate reports, inherited approvals nobody can justify, and processes that cost more attention than the risk they manage.</p><p>These categories do not make technical work inappropriate. A technical lead might keep implementation when it reduces important risk, establishes a reusable pattern, or supplies context that would be expensive to gain another way. An engineering manager might stay close by inspecting risk and decision quality without becoming the routine reviewer or implementer. The test is whether lead-level attention changes the outcome, not whether the work is technical.</p><p>Treat <strong>Own</strong> as the uncomfortable &#8220;only I&#8221; list, then challenge every item in it. Does it truly require your authority? Is the information sensitive? Is the decision hard to reverse? Is the risk high enough that you need to own the call? If the answer is no, the work may need your context or a checkpoint, but it probably does not need your hands.</p><p>Do not remove yourself so aggressively that high-risk work loses necessary context. Transferring a security-sensitive change and then disappearing creates an accountability gap. Match your involvement to the person, the risk, and the reversibility of the decision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zXGO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zXGO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 424w, https://substackcdn.com/image/fetch/$s_!zXGO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 848w, https://substackcdn.com/image/fetch/$s_!zXGO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 1272w, https://substackcdn.com/image/fetch/$s_!zXGO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zXGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png" width="1456" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:836822,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207258513?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zXGO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 424w, https://substackcdn.com/image/fetch/$s_!zXGO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 848w, https://substackcdn.com/image/fetch/$s_!zXGO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 1272w, https://substackcdn.com/image/fetch/$s_!zXGO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12ae991b-372d-4421-8986-32f1a28e93cf_2865x1590.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Move each responsibility to the lowest level of involvement that is still safe.</em></figcaption></figure></div><p>For the next week, choose one item to stop, one to transfer, and one to stay close to through a lighter mechanism. If <strong>Own</strong> still contains half the role, go through it again. You may be protecting a preferred method rather than a necessary responsibility.</p><h3>Budget your week before it fills</h3><p>Once you know what should remain yours, build the week around responsibilities rather than incoming requests.</p><p>Start with the commitments that are genuinely fixed: 1:1s, team planning, interviews, an operating review, or a launch decision with an external deadline. Add preparation time for the commitments that need it. A one-hour feedback conversation is not a one-hour piece of work if you need thirty minutes beforehand to review examples and decide what you actually want to say.</p><p>Then add the work that is important but easy to displace:</p><ul><li><p>decisions and tradeoff preparation;</p></li><li><p>consequential technical context;</p></li><li><p>difficult conversations;</p></li><li><p>written context that will prevent repeated explanation;</p></li><li><p>delegation checkpoints;</p></li><li><p>support or escalation requests to your manager.</p></li></ul><p>Some of this work deserves a name: judgment time. It is where you compare weak signals, prepare consequential feedback, inspect technical risk, and decide which tradeoff you are willing to defend. Without a named block, visible tasks and incoming requests can repeatedly displace it.</p><p>Create an operating reserve. As a starting experiment, reserve 10&#8211;20% of your sustainable week if you can, then adjust it through the weekly audit. A team with frequent on-call interruptions may need more. A lead with little control over the calendar may begin with one open half-day. The point is to treat incidents, escalations, and work that runs long as normal operating conditions rather than planning failures.</p><p>The arithmetic has to work. If your plan contains forty-six hours of commitments, a better color scheme will not turn it into a forty-hour week.</p><p>Consider a hypothetical new engineering manager with seven reports, a migration in progress, and normal on-call exposure:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sdgA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sdgA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 424w, https://substackcdn.com/image/fetch/$s_!sdgA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 848w, https://substackcdn.com/image/fetch/$s_!sdgA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 1272w, https://substackcdn.com/image/fetch/$s_!sdgA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sdgA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png" width="1456" height="1521" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1521,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:517347,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207258513?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sdgA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 424w, https://substackcdn.com/image/fetch/$s_!sdgA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 848w, https://substackcdn.com/image/fetch/$s_!sdgA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 1272w, https://substackcdn.com/image/fetch/$s_!sdgA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80a58371-a759-4efa-ae33-4956d670a4ce_2642x2760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The numbers are illustrative. A tech lead may need more technical time; a manager with a larger team may need more people-management time. The important change is what left, not the allocation itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bjQH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bjQH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 424w, https://substackcdn.com/image/fetch/$s_!bjQH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 848w, https://substackcdn.com/image/fetch/$s_!bjQH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 1272w, https://substackcdn.com/image/fetch/$s_!bjQH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bjQH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png" width="1456" height="1128" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1128,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:210978,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207258513?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bjQH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 424w, https://substackcdn.com/image/fetch/$s_!bjQH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 848w, https://substackcdn.com/image/fetch/$s_!bjQH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 1272w, https://substackcdn.com/image/fetch/$s_!bjQH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff64c7017-bab1-411d-9bbf-6fa504aeb4dc_1600x1240.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>This is an illustration, not a prescribed allocation. The important change is what left.</em></figcaption></figure></div><p>Before moving on, make the budget balance:</p><ol><li><p>Subtract fixed commitments, preparation, and the operating reserve from your sustainable weekly capacity.</p></li><li><p>Allocate what remains across people, decisions, technical context, stakeholders, and administration.</p></li><li><p>Remove, reduce, or transfer work until the total fits without relying on evenings.</p></li></ol><h3>Block your time</h3><p>Calendar blocks help only when they protect a named outcome. &#8220;Focus time&#8221; is too vague. It is easy to overwrite because nobody, including you, knows what would be lost.</p><p>Use specific names:</p><ul><li><p>Review the migration risk and decide whether to phase the launch.</p></li><li><p>Prepare feedback for Thursday&#8217;s 1:1.</p></li><li><p>Read the API proposal and write the decision boundary.</p></li></ul><p>A useful block also needs an interruption rule and a response window. For example:</p><blockquote><p>I am using 9&#8211;11 to prepare feedback and make the release decision. Interrupt for a production incident, an urgent people issue, or a decision holding up several people. Otherwise, add it to the queue and I will respond by 2.</p></blockquote><p>This keeps protected time from becoming unavailability. The team knows what may interrupt it and when other requests will receive attention.</p><p>Block three things in the next week: one decision, one consequential conversation, and one technical risk. That is enough to test the shape without turning the calendar into a fortress.</p><p>Do not protect personal output while the team waits for a decision only you can make. If five engineers are blocked on a release call, move the writing block, make the call, and restore the time somewhere visible. Protect important work from accidental displacement, not every calendar event.</p><p>If the same &#8220;temporary&#8221; meeting overwrites a block every week, the block is not protected and the meeting is not temporary. Resolve the conflict instead of maintaining the fiction.</p>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/the-new-engineering-leads-guide-to">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The New Engineering Manager's 90-Day Checklist]]></title><description><![CDATA[A practical management-readiness table for reviewing expectations, capturing evidence, and choosing what to improve next.]]></description><link>https://newsletter.thelongcommit.com/p/the-new-engineering-managers-90-day</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-new-engineering-managers-90-day</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Wed, 15 Jul 2026 10:29:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tXFB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The first months of management can turn into an unprioritized list of legitimate demands: 1:1s, feedback, hiring, planning, delivery, technical context, stakeholder expectations, and your own learning. Trying to do all of it at once is a good way to stay reactive.</p><p>This resource turns that transition into a structured table. It groups the work into clear management areas, names the expectation behind each topic, and gives you space to record evidence, context, and your next move.</p><h2>Who this is for</h2><p>First-time engineering managers and experienced managers entering a new team who need a practical way to decide what deserves attention now, what can wait, and what evidence would show progress.</p><h2>What is inside</h2><ul><li><p>More than 60 management expectations across team and trust, delivery and quality, collaboration and communication, direction and planning, and manager growth</p></li><li><p>A clear <strong>Topic</strong> and <strong>Expectation</strong> for every row</p></li><li><p>A <strong>Notes and examples</strong> column for evidence, context, or your next action</p></li><li><p>An interactive <strong>Done</strong> checkbox for each expectation</p></li><li><p>Review guidance for choosing a few priorities and revisiting them around days 30, 60, and 90</p></li></ul><p>The point is not to check every box. Review it with your manager, choose a few priorities at a time, and adapt it to your business context, team structure, and team dynamics.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tXFB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tXFB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!tXFB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!tXFB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!tXFB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tXFB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png" width="728" height="520" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:971369,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207137110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tXFB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!tXFB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!tXFB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!tXFB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57b86543-f70c-4af2-a890-3c26cbd085ad_2746x1962.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cFBj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cFBj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!cFBj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!cFBj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!cFBj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cFBj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:993781,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207137110?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cFBj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!cFBj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!cFBj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!cFBj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6ef6651-86d3-4417-a7f5-cce9d1ba548d_2746x1962.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Get the checklist</h2>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/the-new-engineering-managers-90-day">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Technical Decision Consequences Worksheet]]></title><description><![CDATA[Record not only what you decided, but the costs, risks, and consequences worth revisiting after implementation.]]></description><link>https://newsletter.thelongcommit.com/p/technical-decision-consequences-worksheet</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/technical-decision-consequences-worksheet</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Wed, 15 Jul 2026 09:40:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BSdU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A compact Notion template for recording a technical decision, the options considered, and the consequences worth revisiting later. Use it when a decision affects more than the immediate implementation or when the reasoning could otherwise disappear after the work is done.</p><h2>Who this is for</h2><p>Senior engineers and engineering leaders making decisions that create meaningful work, risk, or maintenance after implementation.</p><h2>What is inside</h2><ul><li><p>A decision log with the date, status, owner, and revisit date</p></li><li><p>A repeatable worksheet for the problem and options considered</p></li><li><p>Prompts for immediate benefits, costs that may appear later, and reversibility</p></li><li><p>Space to record the eventual outcome</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BSdU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BSdU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!BSdU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!BSdU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!BSdU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BSdU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:916300,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207133140?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BSdU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!BSdU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!BSdU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!BSdU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F11c3fd46-b393-4cef-b872-1c50e82145c7_2746x1962.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Get the template</h2>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/technical-decision-consequences-worksheet">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Staff and Tech Lead Role Agreement]]></title><description><![CDATA[Define ownership, boundaries, support, and success before a technical leadership role becomes a bundle of conflicting expectations.]]></description><link>https://newsletter.thelongcommit.com/p/staff-and-tech-lead-role-agreement</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/staff-and-tech-lead-role-agreement</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Wed, 15 Jul 2026 09:36:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_ObN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Clarify what a Staff engineer or tech lead is expected to own before an ambiguous leadership role turns into conflicting expectations. This worksheet creates a shared agreement about outcomes, boundaries, technical involvement, available support, and success.</p><h2>Who this is for</h2><p>Use it when starting a Staff or tech lead role, taking responsibility for a new initiative, or resetting expectations after the role has begun to drift.</p><h2>What is inside</h2><ul><li><p>The role and why it exists</p></li><li><p>No more than three expected outcomes</p></li><li><p>Clear ownership, shared decisions, and out-of-scope work</p></li><li><p>Technical involvement, support, capacity, and measures of success</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_ObN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_ObN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!_ObN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!_ObN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!_ObN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_ObN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:828703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207133056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_ObN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!_ObN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!_ObN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!_ObN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c94f0e-fc1a-4bce-8979-d145cf49c97f_2746x1962.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Get the template</h2>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/staff-and-tech-lead-role-agreement">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Career Growth Plan]]></title><description><![CDATA[Turn a vague career direction into concrete objectives, a 90-day focus, and goals you can review through real work.]]></description><link>https://newsletter.thelongcommit.com/p/career-growth-plan</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/career-growth-plan</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Wed, 15 Jul 2026 09:33:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tgdE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Turn a vague career direction into a plan you can return to. The template connects a longer-term direction with near-term objectives, a current 90-day focus, and a small set of goals you can review regularly.</p><h2>Who this is for</h2><p>Senior engineers and engineering leaders who want more than a one-time career worksheet. Use it when you know you want to grow but need a clearer way to connect that direction to real work, evidence, and regular review.</p><h2>What is inside</h2><ul><li><p>A career direction section for the work you want more and less exposure to</p></li><li><p>Two or three objectives for the next 6&#8211;12 months</p></li><li><p>A focused 90-day responsibility or skill to practise through real work</p></li><li><p>A five-field goal tracker for Area, Status, Target date, and Next review</p></li><li><p>A reusable goal page for actions, evidence, support, and review notes</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tgdE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tgdE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!tgdE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!tgdE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!tgdE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tgdE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:832634,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207132692?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tgdE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!tgdE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!tgdE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!tgdE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46da3a39-aa01-469b-b50f-69c9781f3a63_2746x1962.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Get the template</h2>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/career-growth-plan">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Senior Engineer Work and Impact Log]]></title><description><![CDATA[Capture the work, decisions, and evidence that are hardest to reconstruct when career conversations arrive.]]></description><link>https://newsletter.thelongcommit.com/p/senior-engineer-work-and-impact-log</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/senior-engineer-work-and-impact-log</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Wed, 15 Jul 2026 09:25:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OH-F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Keep a useful record of work that becomes difficult to reconstruct once a project has moved on. This template helps you capture what happened, what you did, what changed, and the evidence behind it.</p><h2>Who this is for</h2><p>Use it if you are a senior engineer whose impact spans projects, decisions, mentoring, incidents, or cross-team work. Add an entry after meaningful work, then return to the log before career conversations and performance reviews.</p><h2>What is inside</h2><ul><li><p>A simple log with Entry, Date, Category, Scope, and Evidence</p></li><li><p>A repeatable entry page covering Situation, What I did, Result, Evidence, and What this shows</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OH-F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OH-F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!OH-F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!OH-F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!OH-F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OH-F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png" width="1456" height="1040" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1040,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:907179,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207132047?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OH-F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 424w, https://substackcdn.com/image/fetch/$s_!OH-F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 848w, https://substackcdn.com/image/fetch/$s_!OH-F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 1272w, https://substackcdn.com/image/fetch/$s_!OH-F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ceb2f58-fb27-4a63-b524-5ad635f754af_2746x1962.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Get the template</h2>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/senior-engineer-work-and-impact-log">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Engineering Manager's 1:1 System]]></title><description><![CDATA[A lightweight Notion template for keeping recurring 1:1s focused without turning them into status meetings.]]></description><link>https://newsletter.thelongcommit.com/p/the-engineering-managers-11-system</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-engineering-managers-11-system</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Wed, 15 Jul 2026 09:17:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FwCH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A lightweight Notion template for keeping recurring 1:1s focused without turning them into status meetings. It gives each conversation a consistent structure and keeps agreed actions visible between meetings.</p><h2>Who this is for</h2><p>Engineering managers and other people leaders who want simpler meeting notes and clearer follow-through. Use it when useful context and action items are getting lost between conversations.</p><h2>What is inside</h2><ul><li><p>A meeting tracker with the date, an optional feeling check, and action-item status</p></li><li><p>A repeatable agenda for what is going well, what is not going well, and feedback</p></li><li><p>Dedicated space for action items and additional notes</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FwCH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FwCH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 424w, https://substackcdn.com/image/fetch/$s_!FwCH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 848w, https://substackcdn.com/image/fetch/$s_!FwCH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 1272w, https://substackcdn.com/image/fetch/$s_!FwCH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FwCH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png" width="1456" height="1056" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1056,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:919025,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207130791?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FwCH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 424w, https://substackcdn.com/image/fetch/$s_!FwCH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 848w, https://substackcdn.com/image/fetch/$s_!FwCH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 1272w, https://substackcdn.com/image/fetch/$s_!FwCH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6029da8d-c7c3-43f3-a1cc-2fd3951ee4c4_2746x1992.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Get the template</h2>
      <p>
          <a href="https://newsletter.thelongcommit.com/p/the-engineering-managers-11-system">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Prototype Is a Question, Not a Product]]></title><description><![CDATA[Use AI to turn disagreement into evidence without mistaking a polished demo for a product decision.]]></description><link>https://newsletter.thelongcommit.com/p/the-prototype-is-a-question-not-a</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-prototype-is-a-question-not-a</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 14 Jul 2026 23:22:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3xP6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3xP6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3xP6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 424w, https://substackcdn.com/image/fetch/$s_!3xP6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 848w, https://substackcdn.com/image/fetch/$s_!3xP6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 1272w, https://substackcdn.com/image/fetch/$s_!3xP6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3xP6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png" width="1456" height="776" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:776,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:255215,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207084892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3xP6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 424w, https://substackcdn.com/image/fetch/$s_!3xP6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 848w, https://substackcdn.com/image/fetch/$s_!3xP6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 1272w, https://substackcdn.com/image/fetch/$s_!3xP6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8430ad9c-7eda-4e1d-a810-b892a5bca35e_2440x1300.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Polish changes perception. Evidence changes the decision.</figcaption></figure></div><p>A product lead wants a natural-language setup flow for a developer tool. One engineer thinks it could remove most of the friction from the first integration. Another worries that authentication, error recovery, and the current SDK contracts will make the experience brittle. A staff engineer sees a different risk: if the demo looks convincing, it may become a roadmap commitment before the team understands what it would take to operate.</p><p>The team could spend another week improving the design document, or it could build a narrow prototype and put the disagreement in front of real evidence.</p><p>AI changes the economics of that choice. For throwaway prototypes, I have felt the speed gain directly: <a href="https://newsletter.thelongcommit.com/p/everything-i-learned-about-productivity">a proof of concept that once took two or three days can now land in an afternoon</a>. That makes implementation cheap enough to use during the decision process, not only after the decision has already been made.</p><p>But AI makes the artifact cheaper, not the evidence. Once people can click through the flow, call the API, or watch the agent complete a task, the discussion shifts from &#8220;Should we build this?&#8221; to &#8220;What would it take to ship what we already have?&#8221;</p><p>The artifact has started answering a question nobody agreed to ask.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!56BB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!56BB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 424w, https://substackcdn.com/image/fetch/$s_!56BB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 848w, https://substackcdn.com/image/fetch/$s_!56BB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 1272w, https://substackcdn.com/image/fetch/$s_!56BB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!56BB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png" width="1456" height="967" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:967,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:687418,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207084892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!56BB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 424w, https://substackcdn.com/image/fetch/$s_!56BB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 848w, https://substackcdn.com/image/fetch/$s_!56BB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 1272w, https://substackcdn.com/image/fetch/$s_!56BB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45dc2759-c17f-4a3a-bac3-6db3a31f3805_3880x2578.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>A prototype earns authority from the evidence around it, not from how finished it looks.</em></figcaption></figure></div><p>My recommendation is to build the disagreement when implementation can produce evidence, but treat the prototype as an expiring question rather than an early product. Before anyone prompts an agent, define what the prototype may prove, which shortcuts limit that conclusion, when its mandate ends, and what happens to the implementation afterward.</p><p>None of this is unique to AI. <a href="https://hci.stanford.edu/courses/cs247/2012/readings/WhatDoPrototypesPrototype.pdf">&#8220;What Do Prototypes Prototype?&#8221;</a> frames prototype focus as a choice about which open design questions to examine, while <a href="https://doi.org/10.1145/1375761.1375762">&#8220;The Anatomy of Prototypes&#8221;</a> describes prototypes as filters over a broader design space. AI changes how quickly a functional, convincing artifact can appear&#8212;and, I would argue, the organizational pressure to treat an answer-shaped artifact as a product.</p><h3>When this rubric is worth using</h3><p>Not every sketch needs a formal hypothesis. If two engineers are exploring a low-risk implementation detail for an hour, let them explore.</p><p>Use the rubric when the artifact could influence:</p><ul><li><p>roadmap scope or a customer promise;</p></li><li><p>an architectural or integration decision;</p></li><li><p>a production or security boundary;</p></li><li><p>a cross-team commitment;</p></li><li><p>a claim about user behavior;</p></li><li><p>whether generated code becomes part of a maintained system.</p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><p>The point is proportionality. Five lines written before two days of implementation are not heavy governance. They protect the team from spending the next six months maintaining an answer to the wrong question.</p><h3>The prototype decision card</h3><p>Every consequential prototype should begin with five fields:</p><ol><li><p>Question</p></li><li><p>Evidence</p></li><li><p>Shortcuts</p></li><li><p>Expiry</p></li><li><p>Disposition</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ExlU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ExlU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 424w, https://substackcdn.com/image/fetch/$s_!ExlU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 848w, https://substackcdn.com/image/fetch/$s_!ExlU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 1272w, https://substackcdn.com/image/fetch/$s_!ExlU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ExlU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png" width="1456" height="2024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2024,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1192471,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207084892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ExlU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 424w, https://substackcdn.com/image/fetch/$s_!ExlU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 848w, https://substackcdn.com/image/fetch/$s_!ExlU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 1272w, https://substackcdn.com/image/fetch/$s_!ExlU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F239ce20e-5d6a-404b-a144-2566296c1411_3879x5393.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The card defines the inference boundary before the artifact starts persuading people.</em></figcaption></figure></div><p>This is not a new approval process. The decision boundary must travel with the artifact. Keep the card inside the issue, design document, or pull request description, and put the question, expiry date, and status on the prototype itself. A forwarded link, screenshot, or recorded demo should still say what the artifact was designed to test and whether it is approved for production or customer commitments.</p><blockquote><p><strong>Prototype for:</strong> Expired-credential recovery</p><p><strong>Expires:</strong> 18 July</p><p><strong>Status:</strong> Not approved for production or customer commitments</p><p><strong>Decision card:</strong> Linked with this artifact</p></blockquote><h4>1. Question: What decision should change?</h4><p>Name one decision the prototype will inform.</p><p>&#8220;Can we build this?&#8221; is usually too broad. Almost anything can be made to work once under favorable conditions. Narrower questions force the team to identify the disagreement underneath the implementation:</p><ul><li><p>Can a target developer finish the first integration without leaving the guided flow?</p></li><li><p>Does the current API contract support recovery from an expired credential?</p></li><li><p>Can this architecture meet the latency target when one dependency degrades?</p></li><li><p>Can the agent handle the representative failure cases without taking an unsafe action?</p></li></ul><p>A useful question has consequences. Before building, state what the team will do after a positive result and what it will do after a negative one:</p><blockquote><p>If the result is positive, we will [action]. If it is negative, we will [action].</p></blockquote><p>If neither result would change the decision, the team is not running an experiment. It is producing a demonstration. If the evidence lands between the predeclared criteria, call the result inconclusive and write a new card rather than rewriting the first question.</p><h4>2. Evidence: What would count as an answer?</h4><p>State what the team will observe or attempt, who will do it, and which results will count as positive, negative, or inconclusive. Define those criteria before seeing the artifact, so the team cannot move the goalposts after the prototype starts persuading people.</p><p>The evidence must match the question. A usability claim requires target users attempting the task. An operational claim requires load, failure, recovery, or maintenance evidence. A claim about agent reliability requires a representative evaluation set, not one impressive trace selected for the demo.</p><p>In <a href="https://doi.org/10.1145/3706598.3713166">&#8220;Prototyping with Prompts&#8221;</a>, a CHI 2025 design study, researchers observed 39 industry practitioners across 13 team sessions as they prototyped prompts for a generative AI marketing application. The authors describe two layers of evaluation: teams assessed output validity and correctness while considering whether end users could understand and meaningfully interact with it. They also found that assumptions about cultural norms, user behavior, content formats, and conceptual abstractions could remain implicit and undocumented. This suggests that greater fidelity can expand what a team is able to inspect, but it does not determine which evidence a decision requires.</p><h4>3. Shortcuts: What is intentionally unlike production?</h4><p>List every condition that is fake, stubbed, omitted, unreviewed, not tested at scale, incomplete, or unusually favorable.</p><p>Typical shortcuts include:</p><ul><li><p>mocked billing or permissions;</p></li><li><p>a test tenant instead of a real account;</p></li><li><p>one supported language or environment;</p></li><li><p>clean data that excludes known edge cases;</p></li><li><p>no production secrets or security review;</p></li><li><p>a single happy path;</p></li><li><p>manual intervention hidden behind an automated-looking flow;</p></li><li><p>generated code nobody has reviewed for maintainability.</p></li></ul><p>Shortcuts are not a failure. They are what makes a prototype cheap. But a shortcut may only narrow the conclusions the team can draw; it cannot waive non-negotiable privacy, security, legal, or safety requirements.</p><p>Mocked billing may be irrelevant to a navigation test and fatal to an integration-feasibility claim. A test tenant may be enough to inspect the interaction and useless for understanding production permissions.</p><p>For each shortcut, ask: &#8220;Which conclusion are we no longer allowed to draw because this is fake?&#8221;</p><h4>4. Expiry: When does the experiment end?</h4><p>Give the prototype a time, budget, or learning limit.</p><p>Examples:</p><ul><li><p>two days of implementation;</p></li><li><p>five user sessions;</p></li><li><p>one load-test cycle;</p></li><li><p>twenty representative agent tasks;</p></li><li><p>confirmation of one API behavior;</p></li><li><p>the first unrecoverable security constraint.</p></li></ul><p>Without an expiry, a prototype tends to keep absorbing work. Someone adds another path because the first one looked promising. Another engineer improves the error handling. A stakeholder asks whether it can be shown to a customer. The team quietly stops learning and starts developing, but the code never passes through the decisions expected of a real product.</p><p>An expiry does not mean the idea must die. It means the current artifact loses its mandate. Continuing requires a new question, a new experiment, or an explicit production proposal.</p><h4>5. Disposition: What happens to the artifact?</h4><p>Decide in advance whether a positive result leads to:</p><ul><li><p>a clean rebuild;</p></li><li><p>a separate production proposal;</p></li><li><p>another experiment;</p></li><li><p>retention of only the findings or test fixtures;</p></li><li><p>a deliberate stop.</p></li></ul><p>A prototype should not enter production because rebuilding feels wasteful. If the implementation is worth keeping, it should earn that decision independently through normal review, security, testing, ownership, and operational standards. The prototype label should not grant generated code a permanent exemption.</p><p>When the expiry condition arrives, append a short closeout to the card:</p><ul><li><p><strong>Result:</strong> What happened against the predeclared criteria.</p></li><li><p><strong>Decision:</strong> What changes now.</p></li><li><p><strong>Artifact status:</strong> Deleted, archived, retained for evaluation, or proposed for independent review.</p></li><li><p><strong>Owner and date:</strong> Who closed the experiment and when.</p></li></ul><p>The closeout is not a sixth planning field. It is the durable record that stops later readers from reinterpreting the artifact after its context has faded.</p><p>In <a href="https://doi.org/10.1145/3772318.3790757">&#8220;From Throw-Away to Takeaway&#8221;</a>, a CHI 2026 mixed-methods study combined an online survey of 85 people with interviews of 31 hackathon participants and eight practitioners. In its sample, cloud development environments accelerated prototyping and enabled non-technical users to create high-fidelity throwaway prototypes for experiential exploration. Deployment and long-term maintainability, however, continued to depend on technical expertise. The study supports separating prototype evidence from production readiness; it does not show that this five-field card improves team decisions. The card is my operating recommendation.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h3>Match the evidence to the prototype</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1hei!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1hei!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 424w, https://substackcdn.com/image/fetch/$s_!1hei!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 848w, https://substackcdn.com/image/fetch/$s_!1hei!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 1272w, https://substackcdn.com/image/fetch/$s_!1hei!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1hei!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png" width="1456" height="1598" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1598,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:985984,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/207084892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1hei!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 424w, https://substackcdn.com/image/fetch/$s_!1hei!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 848w, https://substackcdn.com/image/fetch/$s_!1hei!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 1272w, https://substackcdn.com/image/fetch/$s_!1hei!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F630acc49-3224-44c6-b978-aeae169bf86a_3640x3996.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Vision prototypes belong in this model even when they are not experiments. Their evidence is discussion and alignment. Stakeholder enthusiasm shows that an idea is persuasive; it does not validate user need, technical feasibility, or roadmap priority.</p><h3>A completed decision card</h3><p>Return to the developer-tool disagreement from the opening. Instead of asking an agent to &#8220;build a natural-language onboarding flow,&#8221; the team writes this first:</p><p><strong>Question:</strong> Can a target developer complete the first integration&#8212;including recovery from an expired credential&#8212;without leaving the guided flow, using the current API contracts?</p><p><strong>Evidence:</strong> Five target developers attempt both a clean setup and a scripted expired-credential case in a faithful test environment. A positive result means at least four complete both cases within fifteen minutes without manual backend intervention. A negative result means the current contract cannot support safe in-flow recovery, or more than one participant reaches an unrecoverable error. Any other result is inconclusive. The team records completion time, confusion, help requests, manual intervention, and unrecoverable errors.</p><p><strong>Shortcuts:</strong> Test tenant only. Billing is mocked. One language is supported. No production secrets or customer data are used. The prototype makes no claim about scale, security approval, or maintainability.</p><p><strong>Expiry:</strong> Two days of implementation and five user sessions, or immediate termination if the current contract cannot support safe in-flow recovery.</p><p><strong>Disposition:</strong> A positive result leads to a separate production proposal and an independent decision about whether any implementation is retained. A negative result preserves the findings and evaluation fixtures, then deletes or archives the prototype. An inconclusive result requires a new card.</p><p>Now the team knows both what success means and what would falsify the experiment. A beautiful demo cannot settle the question if no target developer attempts the integration. A successful happy path cannot hide an API contract that makes recovery impossible, because recovery is part of the evidence plan. A weak result does not automatically kill the broader idea; it may show that the current contract, not the interaction, is the next question to investigate.</p><p>After the sessions, the closeout might read:</p><ul><li><p><strong>Result:</strong> Four of five developers completed the clean path; only two recovered from an expired credential, and three reached an unrecoverable error.</p></li><li><p><strong>Decision:</strong> Stop the onboarding prototype and investigate the API contract.</p></li><li><p><strong>Artifact status:</strong> Prototype archived; evaluation fixtures retained.</p></li><li><p><strong>Owner and date:</strong> Prototype lead, 18 July.</p></li></ul><h3>The failure modes to watch</h3><ul><li><p><strong>The sales demo</strong> is built to persuade and later cited as validation. Label vision and sales artifacts explicitly, including what they were never designed to test.</p></li><li><p><strong>The happy-path proof</strong> shows that something can work once and becomes evidence of production readiness. Record which properties&#8212;load, recovery, security, observability, and maintainability&#8212;remain outside the experiment.</p></li><li><p><strong>The moving question</strong> appears when the team rewrites the hypothesis after seeing the result. Preserve the original card; if the artifact reveals a more useful question, create a new one.</p></li><li><p><strong>The evidence mismatch</strong> uses the wrong people or conditions. Internal stakeholders cannot validate customer usability, and a mocked API cannot validate current integration behavior.</p></li><li><p><strong>The accidental foundation</strong> begins with &#8220;We already have most of it.&#8221; Require an independent production decision. Do not measure waste by lines of prototype code deleted.</p></li><li><p><strong>The orphan prototype</strong> survives after the decision with no owner or status. Make cleanup part of the disposition and record the artifact&#8217;s status in the closeout.</p></li></ul><h3>What to do before the next prototype</h3><p>Before anyone starts building, ask the team to spend ten minutes on the decision card.</p><p>During the experiment, record observations and newly discovered shortcuts. Do not expand the prototype merely because implementation is going well.</p><p>When the expiry condition arrives, stop and complete the closeout. A positive result should change only the decision named on the card. A negative result is useful when the test was valid. An inconclusive result means the question or evidence plan needs another pass.</p><p>AI makes it cheaper to turn disagreements into things a team can inspect. The prototype should earn authority from the evidence it collects, not from how much it resembles a finished product.</p><p>Before the next agent starts generating code, decide what its artifact is allowed to prove&#8212;and what your team has already agreed to do when the question expires.</p>]]></content:encoded></item><item><title><![CDATA[When AI-First Becomes a Loyalty Test]]></title><description><![CDATA[Adoption pressure can teach engineers to perform belief instead of improve the work.]]></description><link>https://newsletter.thelongcommit.com/p/when-ai-first-becomes-a-loyalty-test</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/when-ai-first-becomes-a-loyalty-test</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 07 Jul 2026 09:40:48 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/fc1c9fba-9d40-4c0b-bc37-c583bc89b20e_3025x1530.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A team is in planning. Someone asks for another engineer, another month, or a narrower scope. A few years ago, the useful questions would have been about the problem, the risk, and the cheapest responsible way to reduce it. Now a different question can become the ritual: did you try it with AI first?</p><p>That question can be healthy. Some engineers still underuse tools that would remove boring work, surface options, or force a sharper explanation of the problem. I do not think engineering leaders are wrong to expect people to learn AI. Refusing to build any muscle with these tools is becoming its own professional risk.</p><p>The problem starts when &#8220;AI-first&#8221; becomes a test of seriousness. Once the expected answer is visible, people learn to produce that answer. They mention the tool in planning. They add it to the workflow where it can be seen. They show the demo. They write the performance-review paragraph. The company gets evidence of adoption before it gets evidence that the work improved.</p><p>Luiza Jarovsky&#8217;s essay, <a href="https://www.luizasnewsletter.com/p/when-ai-becomes-a-religion">When AI Becomes a Religion</a>, is too broad for the exact piece I want to write, but it names a pattern worth taking seriously inside companies: salvation language, dogma, and the awkward role of the heretic. I would not import the whole religion frame into engineering. It gets noisy quickly. The workplace question is what happens when skepticism starts to carry a social cost. The engineer who asks for evidence, slows a rollout, or says the generated code made review worse can start to look behind the culture, even when they are protecting the feedback loop the organization needs.</p><p>Shopify is a useful public example because it turns the cultural pressure into an operating rule. In April 2025, coverage from <a href="https://www.businessinsider.com/shopify-ceo-tobi-lutke-employees-prove-ai-job-2025-4">Business Insider</a> and <a href="https://www.theverge.com/news/644943/shopify-ceo-memo-ai-hires-job">The Verge</a> described a memo from CEO Tobi L&#252;tke that made AI usage a baseline expectation, asked teams to show why work could not be done with AI before requesting more headcount or resources, and added AI usage questions to performance and peer review. There is a reasonable version of that management instinct. Headcount should not be the default answer when tools and workflow changes can absorb some of the work. A serious organization can expect people to try available leverage before asking for more capacity.</p><p>The failure mode starts when proof of trying becomes more important than the quality of the result. If a team knows the safest answer is &#8220;yes, we used AI,&#8221; then people will learn how to produce that answer. They will add the tool to the workflow, mention it in planning, show a demo, or include it in a review packet. That may be sincere learning. It may also be adoption theater. From the outside, both can look similar unless leaders measure what happened downstream.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6EhU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6EhU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 424w, https://substackcdn.com/image/fetch/$s_!6EhU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 848w, https://substackcdn.com/image/fetch/$s_!6EhU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 1272w, https://substackcdn.com/image/fetch/$s_!6EhU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6EhU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png" width="1456" height="1686" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1686,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:279135,&quot;alt&quot;:&quot;A flowchart contrasts two AI adoption loops. In the first loop, AI-first pressure makes visible usage the safe answer, teams report adoption, leaders see alignment, bad data gets hidden, and feedback loops get worse. In the second loop, protected skepticism keeps failed prompts, bad outputs, and review friction visible so leaders can adjust the rollout and improve the feedback loop.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/204418114?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A flowchart contrasts two AI adoption loops. In the first loop, AI-first pressure makes visible usage the safe answer, teams report adoption, leaders see alignment, bad data gets hidden, and feedback loops get worse. In the second loop, protected skepticism keeps failed prompts, bad outputs, and review friction visible so leaders can adjust the rollout and improve the feedback loop." title="A flowchart contrasts two AI adoption loops. In the first loop, AI-first pressure makes visible usage the safe answer, teams report adoption, leaders see alignment, bad data gets hidden, and feedback loops get worse. In the second loop, protected skepticism keeps failed prompts, bad outputs, and review friction visible so leaders can adjust the rollout and improve the feedback loop." srcset="https://substackcdn.com/image/fetch/$s_!6EhU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 424w, https://substackcdn.com/image/fetch/$s_!6EhU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 848w, https://substackcdn.com/image/fetch/$s_!6EhU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 1272w, https://substackcdn.com/image/fetch/$s_!6EhU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a5c199b-b528-4dbd-b013-0d91955a41fd_1569x1817.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">When usage becomes the safest answer, adoption data gets cleaner while the rollout gets less honest. Protected skepticism keeps the signal leaders need to make AI adoption work.</figcaption></figure></div><h3>When usage becomes safer than judgment</h3><p>Usage is attractive because it is legible. Licenses activated, prompts submitted, features shipped with AI assistance, agents connected to repositories, demos shown in all-hands meetings. These numbers make the rollout visible. They do not prove that the organization is building better software.</p><p>The Stack Overflow 2026 pulse survey is useful precisely because it holds two ideas together. Workplace agent usage had risen: 59% of respondents said they used agents at work, nearly double the 31% Stack Overflow reported from its 2025 Developer Survey. At the same time, 63% of technologists said they rarely or never let agents run fully on autopilot. The survey was self-reported and had 1,100 respondents, so I would treat it as directional. Still, the tension is the one engineering leaders need to understand. Adoption is real, and supervision is still doing a lot of the work.</p><p>If leadership celebrates usage while the work still requires supervision, <a href="https://newsletter.thelongcommit.com/p/the-sign-off-layer-is-becoming-the">the supervision still has to happen somewhere</a>. It usually lands with senior engineers, reviewers, tech leads, and managers who have to decide whether the generated output is correct enough to ship. The prompt may be faster than writing the first draft of the code. The review can still become slower, especially when the author cannot explain the change, the test is brittle, or the diff is plausible in ways that make mistakes harder to spot.</p><p>Moreover, DORA&#8217;s <a href="https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report">2025 AI-assisted software development report</a> pushes the conversation away from tool access alone and toward the system around the tool. Local productivity gains matter only if they translate into product performance instead of <a href="https://newsletter.thelongcommit.com/p/the-ai-productivity-bill-comes-due">downstream chaos</a>. A team can produce more code-shaped output and still make the engineering system worse.</p><p>There is also a social layer. AI tools are not neutral calculators inside a team. OpenAI&#8217;s post on <a href="https://openai.com/index/sycophancy-in-gpt-4o/">sycophancy in GPT-4o</a> is a reminder that <a href="https://newsletter.thelongcommit.com/p/the-quiet-surrender-to-ai">model behavior can affect trust</a>: the company rolled back an update after GPT-4o became too flattering and agreeable. A tool that agrees too easily can make a weak plan feel stronger than it is. A company culture that rewards visible AI enthusiasm can do the same thing from the other side. The model says the plan is good. The dashboard says adoption is up. The executive story says the team is modernizing. The reviewer who says &#8220;this is not ready&#8221; becomes the inconvenience.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h3>What I would measure instead</h3><p>If I were accountable for an AI rollout, I would not ask teams to prove they are bought in. I would ask where AI made the engineering system measurably better and where it shifted cost out of sight.</p><p>For code review, I would look at whether reviewers are catching better issues or just reading more generated diffs. For incident response, I would look at whether AI-assisted summaries improved diagnosis without flattening uncertainty. For onboarding, I would look at whether new engineers reached useful context faster and whether they could still explain the systems they touched. For migrations and maintenance work, I would look at whether generated changes reduced toil without increasing the burden on the people who own the code after the migration lands.</p><p>Those questions are slower than adoption dashboards, but they are closer to the work. They also leave room for the honest positive case. Sometimes AI will be the right answer. It can remove blank-page friction, draft tests, summarize unfamiliar code, produce a first pass at support tooling, and make a senior engineer faster when the senior engineer still owns the judgment. The critique is aimed at pressure that creates bad data.</p><p>A useful leader should want the bad data. Failed prompts, wrong assumptions, unreviewable patches, brittle generated tests, and moments where the tool made someone faster but less clear about the work are not signs of disloyalty. They are rollout telemetry. If people hide those signals because skepticism sounds like resistance, the organization loses the information it needs to adopt AI well.</p><h3>Protect the useful skeptic</h3><p>In an engineering organization, the useful skeptic is often not trying to stop adoption. They are asking the question that keeps adoption honest. What got better? What got worse? Who absorbed the review cost? Which failure mode disappeared, and which one moved to production? Did the team learn a reusable workflow, or did one person get a private boost that nobody else can inspect or improve?</p><p>That distinction matters for senior engineers. A senior engineer who pushes back on a shallow AI mandate should not sound anti-tool. The stronger posture is more demanding: show me the outcome, show me the review path, show me what we learned, and show me where the tool made the system less reliable. That is not nostalgia for pre-AI engineering. It is the production discipline we should have been using all along.</p><p>It also matters for managers. If &#8220;AI-first&#8221; becomes a loyalty test, managers will get the answers they trained the organization to produce. People will describe their work in AI-friendly language. They will use the tool where the tool is visible. They will avoid being the person who sounds slow, difficult, or unconvinced. The organization may feel aligned while its feedback loops get worse.</p><p>The healthier version starts with the work: the workflow, the risk, the customer outcome, the review burden, and the learning the team needs to keep. Then ask where AI helps and where it creates a new failure mode. Engineers do not need permission to be curious about AI. They need permission to report the truth about it.</p><p>That is the line I would want leaders to protect. Push people to learn the tools. Do not punish the people who keep the evidence honest. In a serious engineering organization, the person who refuses to turn usage into faith may be part of the control system that makes adoption work.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><strong>Subscribe to get weekly essays on engineering judgment, stronger teams, and durable careers in modern software work.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Markdown Is Becoming the Application Layer of AI Apps]]></title><description><![CDATA[Why AI apps are moving from framework-heavy orchestration to harnesses plus Markdown.]]></description><link>https://newsletter.thelongcommit.com/p/markdown-is-becoming-the-application</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/markdown-is-becoming-the-application</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Mon, 29 Jun 2026 12:46:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bfOC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most AI apps started from a reasonable fear. Models were unpredictable, prompts were fragile, and nobody wanted production behavior living inside a magic paragraph. So we did what engineers usually do when a system feels unsafe: we wrapped it in code.</p><p>We built chains, graphs, routers, callback handlers, retrievers, memory abstractions, evaluation layers, and handoff nodes. Some of that was necessary. If an AI system moves money, changes permissions, updates customer data, deploys software, or touches production infrastructure, the hard boundaries belong in code. Nobody should trust a paragraph in a Markdown file to enforce access control.</p><p>But we took that instinct too far. A lot of AI application behavior is not a hard boundary. It is procedure, policy, tone, source discipline, review expectations, escalation rules, examples, and handoff shape. Those are the parts of a workflow that change often, depend on domain judgment, and are usually maintained by people who should not have to modify orchestration code to make the product behave better.</p><p>That is the mistake I think the first wave of AI apps made: we treated too much of the AI layer as an orchestration problem. The future is simpler than that. Code should provide the harness. Markdown should carry much more of the work.</p><p>That is what I mean when I say Markdown is becoming the application layer of AI apps. I do not mean Markdown replaces Python, TypeScript, Go, SQL, permissions, tests, queues, schemas, or production controls. I mean the part of the product that tells the AI system how to behave is moving out of framework code and into written, versioned instructions that humans can read and agents can execute.</p><p>The transition is from code-heavy orchestration to harnesses plus Markdown.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bfOC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bfOC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 424w, https://substackcdn.com/image/fetch/$s_!bfOC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 848w, https://substackcdn.com/image/fetch/$s_!bfOC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 1272w, https://substackcdn.com/image/fetch/$s_!bfOC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bfOC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png" width="1456" height="772" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/067d919e-eb57-4068-b792-67eef173f699_3351x1777.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:772,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1246676,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/204108408?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bfOC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 424w, https://substackcdn.com/image/fetch/$s_!bfOC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 848w, https://substackcdn.com/image/fetch/$s_!bfOC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 1272w, https://substackcdn.com/image/fetch/$s_!bfOC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F067d919e-eb57-4068-b792-67eef173f699_3351x1777.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The old model where graph frameworks took control vs the new model were markdown directs the work</figcaption></figure></div><h2>The First Version Was Too Much Code</h2><p>The framework-heavy approach made sense at the beginning. Early AI applications needed scaffolding because the raw model interface was too loose. Retrieval needed structure. Tool calling needed wrappers. Production teams needed traces, retries, typed inputs, fallbacks, and evaluation hooks. LangChain and LlamaIndex became natural choices because they gave engineers a way to turn an uncertain model interaction into something that looked more like software.</p><p>I do not want to turn those frameworks into cartoon villains. They are useful libraries. LangChain has real orchestration and integration machinery. LlamaIndex is useful for indexing, parsing, retrieval, and data-connected applications. Both ecosystems support Markdown as an input format, whether through LangChain&#8217;s <a href="https://docs.langchain.com/oss/python/integrations/document_loaders/unstructured_markdown">UnstructuredMarkdownLoader</a> or LlamaIndex&#8217;s <a href="https://developers.llamaindex.ai/python/framework-api-reference/node_parsers/markdown/">Markdown node parsers</a>.</p><p>The problem is not that these tools exist. The problem is the reflex they encouraged: when an AI workflow becomes important, encode the next decision as another node. That works when the workflow is stable and the branches are genuinely software branches. It gets awkward when the workflow is really a written policy that people keep learning how to improve.</p><p>A support escalation rule should not always require a framework change. A source policy should not be buried in a callback. An editorial standard should not live as a prompt string in application code. A review checklist should not be scattered across several agent nodes. When those things live in code, the people closest to the domain lose the ability to improve them directly. The workflow becomes more formal, but not necessarily more maintainable.</p><p>This is how teams end up with an AI system that is technically sophisticated and operationally clumsy. The code can call the model, fetch the documents, route the task, and produce an answer. But changing what &#8220;good&#8221; means still requires spelunking through framework glue, prompt fragments, and hidden assumptions.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>The Simpler Shape</h2><p>The simpler architecture has three parts.</p><p>First, product code still owns capabilities and constraints. It defines the APIs, data access, permissions, transactions, schemas, audit logs, deployment paths, and anything else that must be reliable even when the model is wrong.</p><p>Second, the agent harness gives the model a safe place to operate. The harness decides which tools exist, which files can be read, what state is available, how approvals work, how traces are captured, how evals run, and where the system should stop instead of guessing.</p><p>Third, Markdown carries the operating behavior. This is where the workflow lives: the source policy, the escalation rules, the examples, the templates, the product vocabulary, the review expectations, the domain caveats, and the handoff contract.</p><p>That split matters because it puts the right kind of change in the right kind of medium. If the change must be enforced, it belongs in code. If the change teaches the AI system how the team wants work done, Markdown is often a better place for it. It is easier to review, easier to diff, easier to search, and easier for domain owners to maintain with engineering guardrails around it.</p><p>This is not no-code. It is not &#8220;let the prompt handle it.&#8221; It is a different ownership model: engineers build the harness and enforce the boundaries; the Markdown layer describes the work the harness should perform.</p><h2>What Moves From Code To Markdown</h2><p>The useful way to think about the transition is not &#8220;which language are we using?&#8221; The useful question is &#8220;what kind of decision is this?&#8221;</p><p>Code should own the decisions that need deterministic enforcement: permissions, data writes, destructive actions, money movement, external side effects, compliance checks, schema validation, deployment, rollback, and anything where being merely persuasive is not good enough.</p><p>Markdown can own a different class of decisions: how to investigate a bug, how to write a release note, what sources count as reliable, when to escalate a customer issue, what a good support answer looks like, how to compare devices, how to produce an editorial brief, how to prepare a pull request handoff, or how to explain a product feature without overpromising.</p><p>Those second-order decisions are still important. In many AI products, they are the product. A customer assistant that follows the wrong escalation policy is not a small UX problem. A research agent that treats vendor claims as neutral evidence will produce bad work. A coding agent that writes a migration without the team&#8217;s review expectations is creating risk. But the way to improve those behaviors is not always another code path. Often it is a clearer workflow, a better example, a tighter source policy, or a more explicit handoff template.</p><p>That is why Markdown matters. It is the medium where those instructions can become part of the application without disappearing into framework glue.</p><h2>Why This Is Happening Now</h2><p><a href="https://github.com/vercel/eve">Vercel&#8217;s Eve</a> is the cleanest public example of this direction. Eve describes itself as a filesystem-first framework for durable AI agents. A typical Eve agent has an <code>agent/instructions.md</code> file as the always-on prompt, optional Markdown skills that can be loaded on demand, typed tools, channels, and schedules. The important thing is the authoring surface: the workflow is not primarily a visual graph or a chain of Python objects. It is a filesystem with Markdown instructions and skills.</p><p>OpenAI and Anthropic are moving toward the same split from the harness side. The <a href="https://developers.openai.com/api/docs/guides/agents">OpenAI Agents SDK</a> gives applications primitives for agents, tools, handoffs, guardrails, sessions, tracing, and human approval. The <a href="https://developers.openai.com/codex/sdk">Codex SDK</a> lets applications control local Codex agents programmatically. Anthropic&#8217;s <a href="https://code.claude.com/docs/en/agent-sdk/overview">Claude Agent SDK</a> exposes the agent loop, tools, permission model, session management, MCP support, hooks, subagents, and filesystem-based configuration that power Claude Code.</p><p>Those SDKs are not Markdown frameworks. That is not the point. The point is that the harness is becoming product infrastructure. Once a product has a capable harness, the next question is where the product-specific behavior should live. My bet is that a lot of it will live in Markdown.</p><p>The instruction file conventions are already normalizing. Codex reads <code>AGENTS.md</code> for project instructions. Claude Code uses <code>CLAUDE.md</code> for persistent project memory and can import <code>AGENTS.md</code> to avoid duplicated guidance. Google&#8217;s <a href="https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md">Open Knowledge Format draft</a> describes knowledge bundles as directories of Markdown files with YAML frontmatter, meant to be readable by humans, parseable by agents, diffable in version control, and portable across tools.</p><p>These are not the same thing. <code>AGENTS.md</code>, <code>CLAUDE.md</code>, Eve&#8217;s <code>instructions.md</code>, <code>llms.txt</code>, OKF bundles, and framework document loaders all have different scopes and loading rules. But they point toward the same underlying shape: humans write structured text, agents use it as operating context, and the product changes when that text changes.</p><p>There is also research pushing in this direction. The March 2026 paper <a href="https://arxiv.org/abs/2603.16021">&#8220;Interpretable Context Methodology: Folder Structure as Agentic Architecture&#8221;</a> argues that, for sequential workflows with human review, folder structure and Markdown prompts can replace some framework-level orchestration. I would not stretch that into a universal rule. Complex concurrent systems still need real orchestration. But the paper matches a pattern many teams will recognize: a lot of useful work is sequential, review-heavy, and easier to inspect when the process is visible in files.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>What This Looks Like In Practice</h2><p>My own content system has pushed me toward this view. The Long Commit has a private Markdown operating layer: root instructions, voice rules, research standards, workflow files, templates, source policies, and Notion handoff rules. One workflow follows the sources I care about and produces a daily brief. The brief is not the article. It is a prepared starting point: what happened, where the primary sources are, what might be worth reading, and where the evidence is weak.</p><p>Then I read the sources and write the piece myself.</p><p>The useful part is that I can improve the system by editing the Markdown layer. If the brief overweights vendor claims, I change the source policy. If the output starts sounding generic, I change the voice guide and add examples. If the handoff misses caveats, I update the template. I am not rebuilding an orchestration graph every time my editorial process gets sharper.</p><p>The same pattern applies outside writing. For a smart home site like <a href="http://motherhome.io/">motherhome.io</a>, Markdown workflows can define how to research a new device, which sources matter, how to compare Matter support, and what claims need caveats before publication. For Auth0-related technical work, Markdown can hold SDK references, implementation patterns, documentation paths, and examples of correct integration. For home automation, Markdown can describe what actions are safe, which commands require approval, and what the agent should never infer.</p><p>The domains are different, but the architecture is the same: the harness provides capabilities; Markdown carries the domain behavior.</p><p>A project might look like this:</p><pre><code><code>ai-app/
&#9500;&#9472;&#9472; AGENTS.md                    # Root router and shared agent policy
&#9500;&#9472;&#9472; CLAUDE.md                    # Compatibility shim: import AGENTS.md
&#9500;&#9472;&#9472; agent/
&#9474;   &#9500;&#9472;&#9472; instructions.md           # Runtime-specific always-on prompt
&#9474;   &#9500;&#9472;&#9472; skills/
&#9474;   &#9474;   &#9500;&#9472;&#9472; support-triage.md      # Procedure loaded when needed
&#9474;   &#9474;   &#9500;&#9472;&#9472; research-brief.md
&#9474;   &#9474;   &#9492;&#9472;&#9472; release-note.md
&#9474;   &#9492;&#9472;&#9472; tools/                    # Code: typed capabilities and integrations
&#9500;&#9472;&#9472; profile/
&#9474;   &#9500;&#9472;&#9472; product.md                # What the product is and who it serves
&#9474;   &#9500;&#9472;&#9472; domain.md                 # Vocabulary, assumptions, edge cases
&#9474;   &#9492;&#9472;&#9472; voice.md                  # Tone, naming, UX copy, audience
&#9500;&#9472;&#9472; workflows/
&#9474;   &#9500;&#9472;&#9472; bugfix.md                 # How this team investigates and fixes bugs
&#9474;   &#9500;&#9472;&#9472; content-update.md         # How content gets researched and edited
&#9474;   &#9500;&#9472;&#9472; customer-escalation.md     # When to escalate, pause, or ask
&#9474;   &#9492;&#9472;&#9472; deploy.md                 # Deployment workflow and human gates
&#9500;&#9472;&#9472; policies/
&#9474;   &#9500;&#9472;&#9472; source-policy.md          # What counts as evidence
&#9474;   &#9500;&#9472;&#9472; tool-boundaries.md        # Which actions are allowed or dangerous
&#9474;   &#9492;&#9472;&#9472; data-handling.md          # Privacy, retention, customer data rules
&#9500;&#9472;&#9472; templates/
&#9474;   &#9500;&#9472;&#9472; pr-description.md         # Output contract for PRs
&#9474;   &#9500;&#9472;&#9472; handoff.md                # Output contract for human review
&#9474;   &#9492;&#9472;&#9472; decision-record.md
&#9500;&#9472;&#9472; examples/
&#9474;   &#9500;&#9472;&#9472; good-output.md            # Local taste, with reasons
&#9474;   &#9492;&#9472;&#9472; bad-output.md             # Failure modes to avoid
&#9492;&#9472;&#9472; evals/
    &#9500;&#9472;&#9472; cases.md                  # Scenarios the AI layer must handle
    &#9492;&#9472;&#9472; rubric.md                 # What good behavior means
</code></code></pre><p>The folder names are less important than the separation of concerns. <code>AGENTS.md</code> should be a router, not a dumping ground. It should tell the agent what kind of project this is, which workflows exist, when to load them, what actions require approval, and where the source of truth lives. If a vendor-specific file is required, make it a compatibility file. My preference is that <code>AGENTS.md</code> becomes the boring root standard, with <code>CLAUDE.md</code> or <code>instructions.md</code> pointing back to it when possible.</p><p>The workflows should encode judgment, not only steps. A useful workflow tells the agent when to research, what evidence counts, when to stop, what output shape to produce, and what review must happen before the work is treated as done. Templates turn those outputs into contracts. Examples give the system taste. Policies and evals make expectations inspectable.</p><p>That is the application layer I think many AI products are missing. They have models. They have tools. They have framework code. They do not yet have a maintainable place where the team&#8217;s judgment about the work can live.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share The Long Commit&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share The Long Commit</span></a></p><h2>The Part Teams Cannot Skip</h2><p>Markdown can also make a weak process look official. A stale <code>AGENTS.md</code> is worse than no <code>AGENTS.md</code> if the agent trusts it. A source policy that says &#8220;use reliable sources&#8221; is not a source policy. A workflow file full of vague advice is not a workflow. A compatibility file that quietly diverges from the real instructions creates another source of truth. A prompt that asks the model not to do something dangerous is not a substitute for a permission boundary.</p><p>Instruction files do not automatically improve agent output either. The June 2026 paper <a href="https://arxiv.org/abs/2606.13449">&#8220;Toward Instructions-as-Code&#8221;</a>studied 15,549 agentic pull requests across 148 projects and found mixed results after instruction files were added. Some projects improved; others worsened. That is exactly the kind of result I would expect. The presence of a Markdown file does not prove the team has designed a good AI layer. It only proves there is now a place where good or bad instructions can affect the system.</p><p>This is why engineering ownership matters. If a Markdown file changes how the AI product behaves, then changing that file is a product change. It should have owners. It should be reviewed. It should have examples. The important workflows should have eval cases. Stale instructions should be deleted. Vendor-specific files should not multiply into five slightly different policies because every tool wants its own filename.</p><p>The future I am arguing for is simpler, not looser. It has less orchestration code for things that should have been written procedure, and stronger code around the parts that need enforcement. It gives domain owners a real way to improve behavior without pretending that prose can replace permissions, tests, or production controls.</p><p>The transition from code to Markdown is not a retreat from engineering. It is engineering putting the right work in the right layer.</p><p>If I were leading an AI product team, I would ask one practical question before adding another framework node: is this a capability the system must enforce, or is this behavior the harness needs to understand? If it is enforcement, write code. If it is behavior, policy, examples, or handoff, start by designing the Markdown layer.</p><p>That is the architecture I would bet on: code for capabilities and constraints; Markdown for the operating behavior the harness can execute.</p>]]></content:encoded></item><item><title><![CDATA[The AI Productivity Bill Comes Due in Production]]></title><description><![CDATA[AI makes code cheaper. It does not make weak ideas useful or production work free.]]></description><link>https://newsletter.thelongcommit.com/p/the-ai-productivity-bill-comes-due</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-ai-productivity-bill-comes-due</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 23 Jun 2026 11:25:57 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0a0df5ed-94d8-4048-8a0f-2cf2998ab125_2786x1476.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The easiest place for an AI rollout to look successful is the velocity dashboard. Pull request count is up. Cycle time improves. More code gets merged. The tool has a tidy story to tell.</p><p>Production usually tells the longer version. The review queue gets heavier. The same two senior engineers become the validation layer for a larger volume of plausible patches. Support sees more small changes with surprising edge cases. The team ships more often, but on-call starts to feel more expensive. None of this proves the AI rollout failed. It proves the dashboard stopped too early.</p><p>That is the standard I would use for this debate: AI can make code cheaper. It cannot make a weak idea useful, and it cannot make production work free.</p><p>The AI productivity conversation is still too comfortable measuring the part of software work that AI makes easiest to see. Lines of code, pull requests, story points, and deployment frequency are all close to production of work. They are not the same as value. A team can generate more code, open more pull requests, ship more frequently, and still leave customers with worse software and engineers with a more fragile system to operate.</p><p>I do not care very much whether AI helped a team produce more lines of code. Story points were already a weak proxy before implementation became cheaper. Deployment frequency is useful as a delivery capability signal, but it is not a product outcome. Even feature count can lie if the team is shipping work customers do not need. The harsher question is whether the team shipped something that made sense for users and whether the delivery system absorbed the change without pushing hidden cost into review, incidents, support, security, or maintenance.</p><p>If the answer is no, the team did not become more productive. It found a &#8220;cheaper&#8221; way to create activity.</p><h3>The old cost was a filter</h3><p>Dax Raad, who created OpenCode, made a useful version of this argument in a February 2026 post that <a href="https://www.businessinsider.com/dax-raad-post-ai-coding-workplace-bottleneck-productivity-2026-2">Business Insider covered</a>. His point was not that AI coding tools are useless. It was that code production was often not the real constraint. In one sharp line, he wrote that &#8220;ideas being expensive to implement was actually helping.&#8221;</p><p>That is uncomfortable because it names something engineering organizations do not like to admit. Implementation cost was a filter. Not a fair filter, not always a good filter, and often a frustrating one. But it forced some ideas to die before they became roadmap commitments, support obligations, security surfaces, and half-owned production behavior.</p><p>When implementation gets cheaper, weak ideas can travel farther before anyone feels the cost. The organization can build more experiments, more variants, more internal tools, more half-promised features, and more code that nobody is quite ready to own. Some of that is good. AI should make useful neglected work cheaper: tests, documentation, migrations, cleanup, internal tooling, and prototypes that were not worth a full planning cycle. But lower implementation cost also lowers the friction that used to make teams ask whether the idea deserved implementation at all.</p><p>The bill does not disappear. It moves downstream.</p><p>A shallow AI rollout makes code generation faster and then asks the same review, testing, security, release, and support systems to absorb the extra work. That is where the cost shows up: not when the patch compiles, and not when the dashboard celebrates a larger pile of merged changes, but when the organization has to review, ship, operate, recover from, and maintain the work it just made easier to produce.</p><p>Harness has a timely but imperfect signal here. Its <a href="https://www.harness.io/state-of-modernization-2026">2026 State of DevOps Modernization report</a> is vendor research, so I would not treat it as neutral proof. Still, the pattern is worth attention. Harness surveyed 700 engineering practitioners and managers in large enterprises in February 2026. Very frequent AI coding users were more likely to report daily or faster deployments: 45%, compared with 15% of occasional AI coding users. But 69% of that same very-frequent group said AI-generated code leads to deployment problems at least half the time. The report also found higher reported rollback, hotfix, or customer-impacting incident rates among very frequent users, and longer mean time to recovery for production incidents related to code deployments.</p><p>Harness explicitly says there is no causal proof that AI coding caused those problems. That caveat matters. The teams using AI heavily may already be under more delivery pressure, may already deploy more often, or may already have weaker delivery systems. But the caveat does not rescue the easy productivity story. If AI is adopted inside a system that cannot absorb more change, local coding speed becomes a stress multiplier. A good tool can still be deployed into an unready system.</p><p>Stack Overflow&#8217;s <a href="https://stackoverflow.blog/2026/05/27/agents-on-a-leash-agentic-ai-remains-mostly-monitored-at-work/">May 27, 2026 pulse survey</a> points in the same direction from a different angle. Among 1,100 respondents, workplace agent use had risen to 59%, but 63% said they rarely or never let agents run entirely on autopilot. Accuracy and security remained the top concerns. This is not an argument that agents have failed. It is evidence that serious teams still treat production AI work as supervised work.</p><p>Google&#8217;s <a href="https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report">2025 DORA AI-assisted software development report</a> uses the right frame: successful AI adoption is a systems problem, not a tools problem. The report also says value stream management should help local productivity gains turn into product performance instead of downstream chaos. That is the sentence I would put on the wall before any AI productivity review. DORA metrics can tell you whether the delivery system is getting healthier or sicker. They cannot tell you whether the thing you shipped should exist. They are guardrails for delivery health, not proof of product value.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h3>The new work is supervision</h3><p>The bottleneck does not only move. The work changes shape.</p><p>A May 2026 longitudinal study of professional software engineers found that AI coding assistants shifted work away from creation and toward verification. Participants reported spending less time on most development tasks, including 82% who reported spending less time writing code. The authors call the new category <a href="https://arxiv.org/abs/2605.23135">&#8220;supervisory engineering work&#8221;</a>: directing, evaluating, and correcting AI output. They also found a productivity-experience paradox. Self-reported productivity improvements stayed high, but among matched participants, the share reporting worse developer experience in at least one dimension nearly doubled from 14% to 27%.</p><p>That matches the texture many senior engineers recognize. AI does not remove judgment. It moves judgment to a different place and then sends more material to that place.</p><p>Code review is where this becomes obvious. A 2026 vision paper on <a href="https://arxiv.org/abs/2605.17548">code review in the age of AI</a> argues that AI coding assistants increase code production velocity while expanding the volume of code requiring review, turning review into a growing bottleneck unless the workflow around it changes. The paper is a research agenda rather than an outcome study, so I would not cite it as proof that every team is seeing this. But the direction is right. If the organization makes code cheaper and keeps review mostly manual, review becomes the place where the productivity story has to pay rent.</p><p>That is especially true for senior engineers. A team can celebrate more output while concentrating more validation work on the people least able to absorb it. Those engineers already carry architectural memory, production intuition, incident scar tissue, and the informal taste that keeps a codebase from turning into a pile of plausible patches. If AI increases the amount of code that needs their judgment, the organization may be spending its rarest capacity faster than before.</p><p>Siddhant Khare, who builds agent infrastructure, described the human version of this shift in his February 2026 essay on <a href="https://siddhantkhare.com/writing/ai-fatigue-is-real">AI fatigue</a>: &#8220;AI reduces the cost of production but increases the cost of coordination, review, and decision-making.&#8221; That is the part most productivity dashboards miss. The code may arrive faster, but the decisions around the code do not become free. Someone still has to decide whether the output is correct, safe, maintainable, aligned with the architecture, worth shipping, and worth owning.</p><p>This is why I distrust most AI productivity metrics. They usually stop counting at the point where AI looks best.</p><h3>Measure past the merge</h3><p>The mistake is stopping measurement where AI looks best. The useful question is what the system has to absorb after the code exists.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lAip!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lAip!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 424w, https://substackcdn.com/image/fetch/$s_!lAip!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 848w, https://substackcdn.com/image/fetch/$s_!lAip!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 1272w, https://substackcdn.com/image/fetch/$s_!lAip!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lAip!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png" width="1456" height="972" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:972,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:202795,&quot;alt&quot;:&quot;Moving the measurement boundary from where AI looks best to where productivity matters&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/203209853?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Moving the measurement boundary from where AI looks best to where productivity matters" title="Moving the measurement boundary from where AI looks best to where productivity matters" srcset="https://substackcdn.com/image/fetch/$s_!lAip!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 424w, https://substackcdn.com/image/fetch/$s_!lAip!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 848w, https://substackcdn.com/image/fetch/$s_!lAip!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 1272w, https://substackcdn.com/image/fetch/$s_!lAip!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc80cf805-221c-486a-a9eb-69cc6180ee88_2039x1361.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Lines of code, pull request counts, story points, and even cycle time can create the appearance of progress while missing the harder question: did customers get better software, and did the delivery system stay healthy? If cycle time improves because reviews get thinner, tests get weaker, or teams ship work users do not care about, the metric did its job poorly. It measured motion and missed value.</p><p>Harness&#8217;s <a href="https://www.harness.io/state-of-engineering-excellence">2026 State of Engineering Excellence report</a> makes this measurement problem explicit, again with the same vendor-evidence caveat. Its survey says 81% of engineering leaders report increased code review time since deploying AI, 31% of a developer&#8217;s day is now consumed by AI-related invisible work, and 94% say technical debt, validation time, and developer burnout are missing from current metrics. I would not build a strategy from those exact numbers. I would take the pattern seriously: organizations are measuring output more readily than they are measuring the effort required to make that output safe and useful.</p><p>METR&#8217;s February 2026 update on developer productivity measurement is useful for the same reason. METR&#8217;s earlier randomized study found that experienced open-source developers were slower with early-2025 AI tools, but its <a href="https://metr.org/blog/2026-02-24-uplift-update/">2026 update</a> says the next experiment design became hard to interpret as AI adoption changed participation, task selection, quality choices, and time reporting. Developers selected different tasks when AI was allowed, changed how much documentation or testing they produced, and sometimes worked on other things while agents ran. That does not make measurement hopeless. It means the simple before-and-after story is often the least trustworthy story.</p><p>The measurement boundary has to move downstream, toward the places where software work becomes value or turns into debt. I would want a downstream ledger with five categories.</p><ul><li><p>Customer value: adoption, task completion, retention, revenue, support contacts, or whatever signal maps to the user&#8217;s life getting better.</p></li><li><p>Delivery health: lead time, change failure rate, rollback rate, mean time to recovery, flaky pipeline pain, and release interruptions.</p></li><li><p>Review quality: pull request size, review queue time, review rounds, senior reviewer concentration, and whether reviewers are being asked to validate code the author does not understand.</p></li><li><p>Maintenance cost: rework, follow-up fixes, documentation drift, dead features, duplicate code, and code that becomes hard to change a month later.</p></li><li><p>Human cost: cognitive load, on-call interruption, after-hours recovery, and whether the best engineers are spending more time supervising output than making hard technical decisions.</p></li></ul><p>The exact metrics depend on the team. The principle does not: do not let the measurement stop at the point where AI looks best.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h3>Roll it out as a production change</h3><p>If I were accountable for an engineering organization adopting AI coding tools, I would not frame the rollout as &#8220;make engineers faster.&#8221; That framing points everyone at the part of the system the tool can most easily accelerate, then acts surprised when the rest of the system starts to strain. I would frame it as a delivery and product quality change.</p><p>For production work, the author still owns the change. I do not care whether the code came from a model, a snippet, a Stack Overflow answer, a generated patch, or a late-night burst of human confidence. The person merging it is responsible for understanding it. &#8220;The AI wrote it&#8221; is not a root cause. It is a sign that ownership got blurry.</p><p>I would also make the risk boundaries explicit. AI-generated tests, documentation, migrations, internal tools, and boilerplate can often move faster with lower risk. Code touching authentication, authorization, payments, data migration, concurrency, incident recovery, privacy, or customer-visible behavior should not get a lighter review because a model produced it. In some cases it deserves a heavier review, because plausible code can be harder to distrust than obviously messy code.</p><p>I would protect review capacity before celebrating output. If AI increases pull request volume, the answer is not to tell senior engineers to review faster. Shrink the changes. Require authors to explain generated code before review. Track review queue time. Give reviewers room to do the work. Make it acceptable to reject a generated patch because the author cannot explain the tradeoffs. A review culture that depends on two overloaded senior engineers was already fragile. AI just makes the fragility harder to hide.</p><p>I would measure the rollout at the team and system level, not as individual surveillance. The goal is not to rank engineers by prompt efficiency, AI acceptance rate, or commits touched by a model. That will teach people to game the tool and hide the work. The goal is to learn whether the team is delivering more customer value with equal or better production health, review quality, maintenance cost, and human sustainability.</p><p>That is a higher bar than &#8220;we shipped more code.&#8221; It is also the only bar that matters.</p><h3>The useful counterargument</h3><p>Some teams really are shipping more value with AI. The piece should not pretend every AI gain is fake. AI can make neglected work cheaper: tests, documentation, migrations, internal tools, API exploration, cleanup, and the boring tasks that used to lose every prioritization fight. For teams with strong product judgment, review discipline, and ownership norms, those gains can turn into better software and less operational pain.</p><p>But that does not weaken the argument. It raises the standard of proof.</p><p>If AI helps your team ship a feature customers use, with the same or better reliability, without overloading review, without hiding maintenance cost, and without turning senior engineers into cleanup infrastructure, then call it productivity. That is the kind of win I would take seriously. If AI mostly increases the number of things in motion, the number of pull requests waiting for review, the number of features nobody asked for, or the amount of code nobody fully understands, then the team has not discovered leverage. It has discovered a cheaper way to create inventory.</p><p>Software teams already had too much inventory: backlogs full of maybe-work, roadmaps full of stakeholder theater, codebases full of half-owned decisions, and dashboards full of metrics that make motion look like value. AI does not forgive those habits. It compounds them.</p><h3>Users pay the invoice</h3><p>The title says the bill comes due in production, and that is true. Production is where weak assumptions stop being private. It is where review gaps, unclear ownership, fragile rollouts, missing tests, and confused product decisions become visible.</p><p>But production is not the final customer. Users are.</p><p>A production system can survive a change that still should not have been built. A team can deploy safely while wasting months on features that do not improve the product. AI productivity measured only at the engineering boundary will miss that failure.</p><p>The sharp version of the claim is this: AI productivity that does not reach users as better software is not productivity. It is cheaper throughput inside the factory.</p><p>That does not make AI bad, and it does not mean teams should avoid the tools. Teams should use them and push them hard. Let engineers automate the boring parts. Let agents draft, test, explain, migrate, and explore. But do not let the organization pretend that faster code is the same as better software.</p><p>The bill comes due in production, but the invoice is paid by customers. If they do not get a better product, and the team does not get a healthier delivery system, the productivity story is just a nicer dashboard wrapped around more work.</p>]]></content:encoded></item><item><title><![CDATA[The Hiring Signal Is Moving Out of the Code]]></title><description><![CDATA[AI did not make technical interviews less technical. It changed what the interview has to prove.]]></description><link>https://newsletter.thelongcommit.com/p/the-hiring-signal-is-moving-out-of</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-hiring-signal-is-moving-out-of</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 16 Jun 2026 10:38:22 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/27cd0e83-ff4e-4493-8927-c873ecc7fd3e_1729x910.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The easiest mistake in a code review is stopping when the code runs.</p><p>The code compiles and the tests pass. Then you notice the part the test did not cover: the retry behavior changed, an implicit contract moved, or the abstraction will make the next incident harder to debug.</p><p>That is the gap hiring has always tried to measure. Code is the thing a company can put in front of a candidate, so code became the proxy. Solve the problem. Talk through the implementation. Pass the tests. From there, the hiring team inferred judgment, ownership, taste, and level.</p><p>AI makes that inference weaker.</p><p>A candidate can now arrive at correct-looking code with a model beside them. That does not make the work fake, and it does not make coding interviews useless. It means the useful questions start after the answer appears: what did the candidate trust, what did they reject, what did they test, and what risk would they own if this shipped?</p><p>That is the shift this piece is about. The hiring signal is moving out of code output and into the work around it: how the candidate reviews, validates, recovers, communicates, and owns the result.</p><p>Today, we cover:</p><ol><li><p>Why the old proxy was strained before AI.</p></li><li><p>Why behavioral interviews were already testing technical work.</p></li><li><p>Why cheating is the loud symptom, not the root problem.</p></li><li><p>What process evidence looks like when AI is part of the work.</p></li><li><p>Why senior and staff candidates now have to prove level differently.</p></li><li><p>What a better hiring loop might measure.</p></li></ol><h2>1. The old proxy stopped being clean</h2><p>Coding interviews were never meant to model the whole job. They were a shortcut for a problem hiring cannot avoid: a company has to make a decision before it has watched someone work inside its systems, with its constraints, for months.</p><p>That shortcut made sense when the artifact was expensive enough to produce that it carried more signal. A candidate who could reason through a problem, write a working solution, and explain the implementation gave the interviewer something concrete. It was incomplete, but it was useful.</p><p>Developers were already skeptical of the trade. HackerRank&#8217;s 2025 Developer Skills Report says 78% of developers believe assessments do not align with real-world tasks, and 56% find algorithm-based questions irrelevant to their jobs. That complaint predates the current AI wave. A lot of engineers have been saying some version of this for years: the interview problem is cleaner than the job.</p><p>I recognize that bias in myself. On the power-plant telemetry platform I worked on earlier in my career, the difficult parts were not isolated to writing a parser or moving bytes from an edge device to the cloud. The work was in latency, sensor behavior, failure modes, backpressure, retry logic, observability, and the question of what downstream data analysts could trust. A hiring process that only looked at whether someone could make code appear would miss a lot of the engineering that made the system reliable.</p><p>AI widens that gap. In the same HackerRank report, 97% of developers said they use at least one AI assistant, 61% said they use two or more AI tools at work, and AI-generated code accounted for 29% of developers&#8217; code on average. Vendor survey numbers should always be read carefully, but the direction is hard to miss. AI-assisted code is no longer a weird edge case.</p><p>That changes the weight of the artifact. The problem is not that every AI-assisted answer is suspicious. The problem is that the answer explains less about the person behind it. If an assessment already felt detached from daily engineering work, and now the artifact it rewards can be produced or polished by a model, the old signal gets squeezed from both sides.</p><p>That is the first pressure on the interview loop. It cannot only ask whether the candidate reached an answer. It has to ask whether the path to the answer exposed enough understanding to trust the candidate with the job.</p><h2>2. The behavioral interview was already more technical than we admitted</h2><p>The word &#8220;behavioral&#8221; does a lot of damage in engineering hiring. It makes the round sound soft, secondary, almost ornamental. First you prove you can code, then you talk about teamwork and conflict and communication, as if those things happen somewhere outside the technical work.</p><p>That split has always been artificial.</p><p>In an April 2026 Pragmatic Engineer piece, Steve Huynh, formerly a Principal Engineer at Amazon, reflected on nearly 1,000 interviews, including around 600 as an Amazon Bar Raiser. His point is useful because it does not come from AI hype. It comes from years of interview loops before this current wave fully landed. The behavioral round, in his telling, often decides fit and level because it exposes the work that coding rounds do not: how someone handles projects going sideways, disagreements, incomplete information, stakeholder pressure, and influence without authority.</p><p>Those are not decorations around engineering. They are engineering at senior levels.</p><p>A mid-level engineer can often succeed inside a bounded technical problem. A senior engineer has to make progress when the boundary is blurry. A staff engineer has to make other people&#8217;s work easier, align teams that do not naturally agree, and turn a technical direction into something the organization can actually execute. At those levels, &#8220;tell me about a time a project went wrong&#8221; is not a personality test. It is a probe for operational judgment.</p><p>That matters now because the artifact has become less isolated from tools, templates, copied context, and generated suggestions. Interviewers need to inspect the candidate&#8217;s relationship to the artifact: what they trusted, what they rejected, what they tested, what they understood about the codebase, and what risk they noticed before someone else named it.</p><p>The old behavioral round was already asking versions of those questions. The problem is that many loops treated it as a separate soft-skills filter instead of part of the technical evaluation.</p><h2>3. Cheating is the loud symptom</h2><p>The cheating story is real. It is also not the whole story.</p><p>CodeSignal said in February 2026 that detected cheating and fraud attempt rates in proctored assessments rose from 16% in 2024 to 35% in 2025. For entry-level assessments, CodeSignal said the rate increased from 15% to 40%. Those are vendor numbers, and CodeSignal sells assessment integrity products, so they should be treated carefully. The company is measuring detected and flagged attempts, not proving that every flagged session was successful cheating or that its data represents the whole market.</p><p>Still, the direction matches what many hiring teams feel. HackerRank reports that 76% of developers say AI makes gaming assessments easier, and 73% feel it is unfair to lose out to candidates who use AI to game tests. Once AI tools are available to everyone, any assessment that rewards final output while hiding the process becomes easier to distort.</p><p>The tempting response is to make the old proxy more tightly controlled. Ban AI. Add proctoring. Detect screen switching. Flag pasted code. Train interviewers to spot suspicious pauses, suspicious speed, suspicious fluency. Some of that is necessary. Companies need fair processes, and candidates deserve not to lose to someone performing a fake version of competence.</p><p>But fraud detection can only protect the fairness of the test in front of it. It cannot make a weak test more representative of the job.</p><p>If the job itself now includes AI-assisted coding, banning AI from every interview creates a strange mismatch. If the job does not allow AI because of security, compliance, or risk, then banning it in the interview makes sense. But many engineering teams are somewhere messier: developers use AI for debugging, code review, codebase understanding, tests, refactors, scripts, documentation, and first-pass implementation. The interview then pretends the work happens without the tools that shape the work.</p><p>That mismatch creates bad incentives. Candidates hide tool use. Companies hunt for hidden tool use. Both sides spend energy preserving the appearance of a clean artifact.</p><p>The real design question is narrower and harder: can the interview separate a candidate who used AI responsibly from one who used it to mask weak understanding?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R9Sg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R9Sg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 424w, https://substackcdn.com/image/fetch/$s_!R9Sg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 848w, https://substackcdn.com/image/fetch/$s_!R9Sg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 1272w, https://substackcdn.com/image/fetch/$s_!R9Sg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R9Sg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png" width="1456" height="957" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:957,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:385483,&quot;alt&quot;:&quot;A two-lane flow diagram comparing the old hiring proxy with the AI-era signal map. The old lane moves from candidate writes code, to interviewer evaluates the artifact, to hiring team infers skill. The AI-era lane moves from candidate and AI produce code, to interviewer evaluates process evidence, to hiring team infers judgment and level. Process evidence includes codebase understanding, validation, trade-offs, recovery, communication, and ownership.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/202262689?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A two-lane flow diagram comparing the old hiring proxy with the AI-era signal map. The old lane moves from candidate writes code, to interviewer evaluates the artifact, to hiring team infers skill. The AI-era lane moves from candidate and AI produce code, to interviewer evaluates process evidence, to hiring team infers judgment and level. Process evidence includes codebase understanding, validation, trade-offs, recovery, communication, and ownership." title="A two-lane flow diagram comparing the old hiring proxy with the AI-era signal map. The old lane moves from candidate writes code, to interviewer evaluates the artifact, to hiring team infers skill. The AI-era lane moves from candidate and AI produce code, to interviewer evaluates process evidence, to hiring team infers judgment and level. Process evidence includes codebase understanding, validation, trade-offs, recovery, communication, and ownership." srcset="https://substackcdn.com/image/fetch/$s_!R9Sg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 424w, https://substackcdn.com/image/fetch/$s_!R9Sg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 848w, https://substackcdn.com/image/fetch/$s_!R9Sg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 1272w, https://substackcdn.com/image/fetch/$s_!R9Sg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71e78444-a7c3-4595-bfa2-3538c3de2ede_3407x2240.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>When AI can produce the artifact, the interview has to move closer to the process that created it.</em></figcaption></figure></div><h2>4. The interview has to collect process evidence</h2><p>Karat&#8217;s NextGen work is useful as a market signal because it shows one direction hiring infrastructure is moving. In its December 2025 launch announcement, Karat described interviews where candidates work on complex multi-file projects with an integrated AI assistant while human interviewers probe reasoning, trade-offs, and judgment in real time. In the April follow-up, Karat said it now gives engineering leaders evidence such as skill scores, rationale write-ups, and timestamped markers tied to moments in the interview.</p><p>Again, this is vendor evidence. Karat sells interview infrastructure, so it has every reason to argue that interviews need more infrastructure. But the underlying problem is real: if the output and the candidate&#8217;s skill are decoupled, the hiring system needs an evidence trail that lives somewhere other than the final code.</p><p>That maps more closely to how AI-assisted work is settling inside real teams. Stack Overflow&#8217;s May 2026 pulse survey found that AI agent usage at work rose to 59%, up from 31% in the previous annual survey, but also that 63% of technologists still rarely or never let agents run fully on autopilot. The interesting part is the combination. People are using agents more, but the dominant working mode is still supervised, reviewed, and bounded.</p><p>A May 2026 longitudinal study by Annie Vella and Kelly Blincoe makes the same shift visible from another angle. The authors found that 82% of participants reported spending less time writing code, and they describe a broader move from creation toward verification activities. Their term for the new category of work is supervisory engineering: directing, evaluating, and correcting AI output.</p><p>If that is increasingly the work, then interviews that only measure artifact creation are behind the job.</p><p>A better AI-era technical interview does not need to become a surveillance exercise. It needs to make process observable. Give the candidate a realistic codebase slice. Let them use tools if the role would let them use tools. Ask them to inspect a change, critique a model suggestion, write or adjust tests, explain where the model&#8217;s answer is brittle, choose between two implementation paths, and name what they would monitor if this shipped.</p><p>The strongest moments in that kind of interview are often not the moments where code appears. They are the moments where the candidate slows down for the right reason.</p><p>In practice, the signal is often small. The candidate notices that a generated function handles the happy path but changes an implicit contract. They keep the refactor but rename a concept because the model erased domain meaning. Before changing a data model, they ask about production constraints. When the prompt omits a failure case, they add the test anyway.</p><p>The candidate is doing engineering there. The old prompt just did not have a clean way to score it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>5. Senior and staff candidates have to prove level differently</h2><p>The higher the role, the more dangerous it is to treat the finished patch as the whole signal.</p><p>That does not mean senior and staff engineers get a pass on code. A senior engineer who cannot reason through code is a liability, especially now that AI can produce plausible nonsense quickly. But the senior signal is not only implementation under time pressure. It is whether the candidate can choose the right problem shape, reduce risk, and make the result usable by other people.</p><p>Huynh&#8217;s framework in The Pragmatic Engineer piece is helpful here. The excerpt from his book describes level through four dimensions: scope, contribution, impact, and difficulty. Those dimensions are hard to infer from a standalone coding artifact. They show up more clearly in how a candidate reviews a change, explains a decision, notices ambiguity, and handles a problem that does not stay inside the prompt.</p><p>A useful staff-level task here is not a blank editor. It is a small service change where the generated patch fixes the visible bug, but the API is used by another team, the migration has no rollback path, and the test suite only covers the happy path. A strong senior candidate should catch the missing test and explain the production risk. A staff candidate should usually go further: ask who depends on the contract, identify the rollout constraint, and challenge whether this is the right place to solve the problem.</p><p>That is the level difference the code alone will not show. The same artifact can hide very different kinds of judgment. One candidate sees a local implementation problem. Another sees a system boundary, an organizational dependency, and a future incident if the rollout goes wrong.</p><p>This matters more in the current market because the bar has moved up. The Pragmatic Engineer&#8217;s public piece on tech interviews in 2025 describes higher expectations in DSA and system design interviews, more demanding senior and staff bars, and more downleveling. Whether every company experiences this the same way is less important than the direction. In a market with more qualified candidates than open roles, companies can ask for more evidence before they say yes.</p><p>AI raises the same pressure from the other side. If more candidates can produce polished artifacts, the differentiator moves to the explanation around the artifact. Senior candidates need to show the shape of their judgment. Staff candidates need to show the scope of it.</p><p>This is where &#8220;storytelling&#8221; gets misunderstood. The goal is not to tell a smoother story. The goal is to make the work legible. A good senior story is evidence of what was ambiguous, what you owned, which alternatives you considered, what changed because of your decision, where the trade-off hurt, and what you learned when reality disagreed with the plan. The code artifact rarely carries all of that by itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!F1TF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!F1TF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 424w, https://substackcdn.com/image/fetch/$s_!F1TF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 848w, https://substackcdn.com/image/fetch/$s_!F1TF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 1272w, https://substackcdn.com/image/fetch/$s_!F1TF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!F1TF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png" width="1456" height="653" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:653,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:162588,&quot;alt&quot;:&quot;A table showing how the same AI-assisted code change can expose different hiring signals at different levels. The mid-level signal is local correctness and debugging. The senior signal adds risk, validation, and production ownership. The staff signal adds cross-team dependency, rollout design, and whether the change should exist in this part of the system.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/202262689?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A table showing how the same AI-assisted code change can expose different hiring signals at different levels. The mid-level signal is local correctness and debugging. The senior signal adds risk, validation, and production ownership. The staff signal adds cross-team dependency, rollout design, and whether the change should exist in this part of the system." title="A table showing how the same AI-assisted code change can expose different hiring signals at different levels. The mid-level signal is local correctness and debugging. The senior signal adds risk, validation, and production ownership. The staff signal adds cross-team dependency, rollout design, and whether the change should exist in this part of the system." srcset="https://substackcdn.com/image/fetch/$s_!F1TF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 424w, https://substackcdn.com/image/fetch/$s_!F1TF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 848w, https://substackcdn.com/image/fetch/$s_!F1TF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 1272w, https://substackcdn.com/image/fetch/$s_!F1TF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F990cb1de-4c20-4804-858b-522bbcb0888f_1473x661.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The same patch can prove different things depending on whether the candidate treats it as a local fix, a production change, or a system-level decision.</em></figcaption></figure></div><h2>6. A better loop measures the work around the code</h2><p>The better hiring loop is not simply &#8220;allow AI.&#8221; That is too shallow. A company can allow AI and still run a bad interview. It can require AI and accidentally select for prompt performance over engineering judgment. It can ban AI and still run a fair process if the actual job bans AI too. The tool policy matters, but it is not the heart of the design.</p><p>The heart of the design is whether the loop creates comparable evidence of how the candidate works.</p><p>Comparable matters because process-heavy interviews can become unfair very quickly. The more an interview depends on narration, confidence, and polish, the more it can reward candidates who have been coached into the right performance. It can also punish candidates who are strong engineers but less comfortable speaking in a high-pressure environment, or who are working in a second language, or who come from teams where the local storytelling norms are different.</p><p>So the answer cannot be &#8220;make everything behavioral&#8221; and call it modern.</p><p>A better loop would still be structured. It would ask consistent questions. It would use a calibrated rubric. It would document evidence rather than vibes. It would give candidates the same kind of task, the same rules around AI, and the same chance to explain what they did. It would score implementation and supervision as related but distinct signals.</p><p>For a senior backend role, that might mean giving the candidate a small service with a bug, an incomplete test suite, and an AI assistant. The task is not to produce the most code. The task is to understand the behavior, make a safe change, explain the risk, and show how they would validate it. For a staff role, the loop might add an architectural constraint, a cross-team dependency, or a migration choice. The candidate still writes code, but the interview watches how they reason around it.</p><p>This would also make interviews feel less strange to candidates who already work with AI every day. The candidate would not have to pretend their workflow is cleaner than it is. They would have to show that their workflow is responsible.</p><p>That is a higher bar in some ways. It is easier to memorize a pattern than to explain why the model&#8217;s pattern is wrong for this codebase. It is easier to generate a passing solution than to defend the test strategy. It is easier to look productive than to show good judgment when the tool gives you something almost right.</p><p>The phrase &#8220;almost right&#8221; is doing a lot of work here. AI-generated code often fails in the place where interviews used to stop looking. It compiles, but it misunderstands the domain. It passes the visible tests, but it weakens the invariant. It follows the local style, but it changes the operational behavior. It gives you the answer a good interviewer would now use as the start of the interview, not the end.</p><h2>Takeaways</h2><p><strong>The coding round has to become less isolated.</strong> The AI-era interview should not drop the technical bar. It should stop treating a finished artifact as the whole technical record. A candidate who cannot reason through code is still not ready for a serious engineering role. A candidate who can produce code without explaining the decisions around it is also harder to trust than they used to be.</p><p><strong>Behavioral signal needs a better name.</strong> Many of the signals companies call behavioral are really senior engineering signals: handling ambiguity, communicating trade-offs, influencing without authority, owning mistakes, and making decisions when the available information is incomplete. AI did not make these skills newly important. It made them harder to keep outside the technical evaluation.</p><p><strong>Integrity tools are necessary but incomplete.</strong> Hiring teams need some way to prevent candidates from faking work. That is a real problem, especially for early-career assessments where the pressure is intense and the signal is thin. But if the assessment is built around output without process, proctoring can only defend the shape of the test. It cannot make the test more like the job.</p><p><strong>The best interview evidence will look more like review evidence.</strong> In daily AI-assisted work, the important human actions are often direction, evaluation, correction, and ownership. Hiring loops need to make those actions visible. That might mean live review, timestamped evidence, structured rationales, realistic codebase tasks, or interviewer notes tied to specific moments. The exact format can vary, but the evidence has to move closer to the work.</p><p><strong>Candidates need to make ownership visible.</strong> AI-polished artifacts will raise the floor on what many candidates can show. The way to stand out is not to pretend the tools are not there. It is to show the part of engineering the tools do not own: the risk you saw, the trade-off you made, the context you asked for, the change you would not ship, and the consequence you were willing to be responsible for.</p><p>I do not think this ends with one standard interview format. Some roles should ban AI because the job does. Some should permit it because the job does. Some should test both modes. The mistake is treating the passed test as the end of the evidence.</p><p>The code can still be on the screen. It can still pass. The interview cannot stop there.</p><p>The next question is the one that looks more like real work: what would the candidate trust, what would they change, what would they test again, and what would they be willing to own in production?</p><p>I do not think hiring has a clean answer yet. But any process that cannot see that layer is measuring less of the job than it thinks.</p>]]></content:encoded></item><item><title><![CDATA[The Token Budget Is Becoming Engineering Policy]]></title><description><![CDATA[AI coding is moving from personal workflow to metered infrastructure, and the next fight is over budgets, incentives, and who gets to spend.]]></description><link>https://newsletter.thelongcommit.com/p/the-token-budget-is-becoming-engineering</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-token-budget-is-becoming-engineering</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Fri, 12 Jun 2026 17:01:20 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/9750ecfd-8751-46e4-a456-f72e777e5d91_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tokenmaxxing had a short half-life. For a while, burning more tokens looked like seriousness: more prompts, more agents, more generated code, more dashboards showing that someone was really using AI. Then the bill started showing up in places executives could not ignore. Even Sam Altman is now saying AI cost went from something that did not come up at the beginning of 2026 to a &#8220;huge issue,&#8221; with customers joking that their company had spent its entire 2026 budget in Q1, according to <a href="https://www.businessinsider.com/ai-bubble-heads-doomers-sam-altman-ai-costs-huge-issue-2026-6">Business Insider</a> and <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/openai-ceo-sam-altman-admits-ai-token-costs-are-becoming-a-huge-issue-company-seeks-improved-value-as-overspending-becomes-a-meme">Tom&#8217;s Hardware</a>.</p><p>That is a very funny sentence if you have ever sat near an enterprise budget process. It is also the clearest signal that the first social phase of AI coding is ending. The <a href="https://newsletter.thelongcommit.com/p/tokenmaxxing-is-the-dumbest-metric">joke version was tokenmaxxing</a>: use more, show more, look more AI-native. The business version is less fun. It asks who is allowed to spend, how much they can spend, which work justified the expensive model, and what happens when the shared pool runs out before the work does.</p><p>I do not think this means AI coding is failing. I think it means AI coding is becoming normal company software. Most meaningful company spend follows a familiar path: it starts as local freedom, then becomes a team habit, then becomes a line item, then becomes policy. Cloud went through this. Observability went through this. CI minutes, SaaS tools, data warehouse queries, feature-flag platforms, support contracts, and half the software hiding in expense reports went through some version of this. Someone adopts the thing because it helps, usage spreads because the thing is genuinely useful, and eventually the company realizes the spend is material enough that vibes are not a control system.</p><p>For the last couple of years, AI coding mostly lived before that last step. A developer paid for Cursor personally. A team expensed Claude. A manager approved GitHub Copilot because saying no felt more irresponsible than saying yes. A company created an innovation budget because nobody wanted to be the person blocking AI adoption with a spreadsheet. Some of that was sensible. You do not learn a new work pattern by designing the perfect governance model before anyone has touched the tool.</p><p>But there was always going to be a second phase. If a tool can spend variable dollars while reading private code, producing diffs, triggering CI, and creating work that humans have to review, it will not stay a personal productivity preference forever. The business will eventually ask the normal business questions: who owns the spend, who gets more of it, what happens when it spikes, which work justified the expensive path, and whether the output was worth the money.</p><p>This is the follow-on to my earlier essay on <a href="https://newsletter.thelongcommit.com/p/tokenmaxxing-is-the-dumbest-metric">tokenmaxxing</a>. That piece was about why token usage is a terrible status signal. This one is about what happens when the status signal becomes a policy surface. The interesting question is not simply whether AI tools are expensive. The better question is what changes inside engineering when the cost of AI becomes visible, governable, and close enough to the work to interrupt it.</p><h2>1. AI Spend Is Leaving The Side Door</h2><p><strong>The surprising part is not that companies are starting to manage AI spend. The surprising part is that AI spend briefly behaved as if normal company rules did not apply.</strong></p><p>I do not mean that as a complaint about enterprises. This is one of the few things enterprises are extremely consistent about. If a resource becomes large enough, shared enough, and unpredictable enough, the company will eventually attach controls to it. People can dislike the controls, and many controls are badly designed, but the motion itself is not mysterious.</p><p>Cloud is the obvious comparison, but the pattern is broader than cloud. A team buys an observability product because production is hard to understand. A department starts using a SaaS tool because the approved internal process is too slow. A few engineers run something in a managed service because waiting for platform support would take months. The first phase is local usefulness. The second phase is adoption. The third phase is someone asking why this thing now costs enough to show up in a planning meeting.</p><p>AI coding compressed that pattern because the tool was attached directly to work people were desperate to accelerate. It was not a nicer calendar app or a better notes tool. It promised to change the cost of producing software. That gave it political cover. Nobody wanted to be the manager who slowed down AI adoption because the cost model was messy, so a lot of companies tolerated the mess for a while.</p><p>Altman&#8217;s cost comment matters because it is the vendor-side version of the same enterprise pattern. The joke works because it is absurd. It also works because anyone who has watched company spend grow through informal channels knows the next scene.</p><p>The side door closes. Not completely, and not all at once, but enough that the work changes. Budgets appear. Dashboards appear. Exceptions appear. Someone asks whether the expensive model was necessary. Someone else asks why one team burned through the shared pool. A team discovers that an agentic task can be blocked not by code review, CI, or test failures, but by a budget policy. That is the moment AI coding stops being only a tool choice and becomes an operating model.</p><h2>2. Usage-Based Billing Makes The Policy Surface Explicit</h2><p><strong>GitHub did not just change how Copilot gets priced. It made the budget layer part of the engineering workflow.</strong></p><p>On June 1, 2026, GitHub moved Copilot usage to GitHub AI Credits across all plans. GitHub&#8217;s <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">April announcement</a>explained the reason in terms of product shape: Copilot is no longer just an in-editor assistant. It has become an agentic platform that can run long, multi-step sessions, use frontier models, and work across repositories.</p><p>That product shift breaks the old developer-tool pricing intuition. A quick inline question and a long-running agentic coding session are not the same economic event. GitHub says usage is calculated from input, output, and cached tokens, with model-specific rates. The <a href="https://github.blog/changelog/2026-06-01-updates-to-github-copilot-billing-and-plans/">June 1 changelog</a> also says Copilot code review consumes GitHub Actions minutes in addition to AI Credits, and that user-level budget controls are generally available for organizations and enterprises.</p><p>The docs are even more revealing than the announcement. GitHub&#8217;s <a href="https://docs.github.com/en/copilot/concepts/billing/usage-based-billing-for-organizations-and-enterprises">organization and enterprise billing docs</a> describe pooled credits, user-level budgets, cost-center budgets, enterprise spending limits, organization-level budgets, and budget exhaustion behavior. If additional usage is not allowed, usage can be blocked until the next billing cycle. If a user-level budget is exhausted, that user&#8217;s access can stop even when the organization still has credits left. This is not just metering. It is work shaping.</p><p>The same direction shows up in Anthropic&#8217;s Claude Code docs. The <a href="https://code.claude.com/docs/en/costs">cost-management page</a> tells teams to track token usage, set team spend limits, manage context, choose models intentionally, and account for usage patterns like multiple instances or automation. That is the language of infrastructure operations more than old developer tooling. The cost of a developer&#8217;s workflow now depends on model selection, codebase size, context management, concurrency, automation, and how often the agent loops.</p><p>A flat IDE subscription was boring once procurement approved it. A metered agentic workflow is not boring. It has model choice, context size, retries, tool calls, failed loops, parallel agents, CI runs, and review cost. The workflow can be expensive because the work is valuable, or because the task was vague, the repo boundary was too large, and the agent kept spending context to compensate for unclear human direction. The bill does not explain which one happened. It only makes the question impossible to avoid.</p><h2>3. Agentic Coding Has Strange Unit Economics</h2><p><strong>Agentic work is hard to budget because its cost does not map cleanly to the way humans estimate engineering work.</strong></p><p>With ordinary developer tools, the unit economics are legible enough that most engineers never think about them. A seat costs what a seat costs. An IDE license does not become more expensive because a refactor is messy. A GitHub seat does not draw from a shared pool because one engineer asked too many questions about a monorepo. The tool has a price, the work has complexity, and those two things are mostly separate.</p><p>A coding agent spends tokens on context, planning, tool calls, file reads, edits, retries, test output, summaries, and correction loops. The final diff may be small, but the path to that diff can be large. The agent may read the wrong files first, carry stale context, retry a failing test several times, use a frontier model where a cheaper model would have worked, or keep processing a huge prompt because the human never narrowed the task. The cost is partly technical and partly managerial.</p><p>The research backs up what heavy users already feel. The April 2026 arXiv paper <a href="https://arxiv.org/abs/2604.22750">How Do AI Agents Spend Your Money?</a> studied token consumption in agentic coding tasks and found that these tasks can consume far more tokens than simpler code chat or code reasoning. It also found high variability across runs of the same task, weak alignment between human-rated task difficulty and token cost, and no simple relationship where more tokens reliably means better accuracy.</p><p>More spend can be rational, and more spend can be waste, but the token count alone cannot tell you which one happened. Two engineers can hit the same monthly budget for completely different reasons. One may be using an agent to pay down migration risk across a messy production system. Another may be asking an expensive model to generate throwaway scaffolding because the approved workflow makes the expensive path easier than the sensible one. A dashboard that treats those as equivalent will produce bad management because it sees consumption before it sees judgment.</p><p>Engineering has lived with some version of this with the cloud, but AI makes the feedback loop more ambiguous. A production service that suddenly spends too much cloud money usually has an operational shape: traffic changed, a job got stuck, storage grew, a query got expensive. An agentic coding session can burn money in a more human-looking way. The prompt was broad. The context was stale. The model choice was excessive. The task should have been split. The agent was allowed to continue because the progress looked plausible. That makes token cost an engineering design problem, not just a billing problem.</p><h2>4. The Dashboard Will Be Tempting And Wrong</h2><p><strong>The worst version of this future is not expensive AI. It is an incentive system that makes expensive AI look like productivity.</strong></p><p>Usage visibility is necessary. Teams need to know which workflows are cheap, which are expensive, which users are outliers, and which agent patterns create runaway spend. Without visibility, leaders are guessing, and guessing is not a policy. The danger starts when visibility becomes a scoreboard.</p><p>This is the tokenmaxxing failure mode. If high token usage becomes a status signal, people will find ways to use more tokens. If low token usage becomes the status signal, people will hide useful work, route around the approved tools, or waste human time avoiding a budget that would have been cheap to spend. If raw AI-generated output becomes the signal, teams will reward surface area instead of engineering value.</p><p>That should feel familiar because software organizations have done versions of this with lines of code, ticket counts, pull request volume, and story points when they escaped their original purpose and became management theater. The lesson is not that measurement is bad. The lesson is that visible metrics attract performance, and engineering systems are easy to perform badly. Token usage is especially tempting because it looks concrete: it has numbers, charts, cost, and a clean path into budget conversations. It gives managers something to point at, which is useful until pointing replaces understanding.</p><p>The good version of usage visibility is outlier inspection. Which tasks cost far more than expected? Which model choices produced no better result? Which agent loops triggered repeated CI runs? Which teams are burning context because their repos are hard to navigate? Which workflows save review time, not just implementation time? Those are useful questions because they connect spend to the shape of the work. The bad version is rank ordering engineers by how much AI they used or how little AI they used. That is how a cost-control tool becomes a culture problem.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>5. Agile AI Will Be a Thing</h2><p><strong>My bet is that AI spend becomes part of engineering planning, but not as exact token estimation. It will be more awkward and more familiar than that.</strong></p><p>The funny version is token poker. The serious version is that teams will start classifying work by model intensity, autonomy level, and budget risk. Not because anyone can predict exact token usage. They cannot, and the research already suggests the cost can vary across runs in ways humans do not estimate cleanly. But teams do not need fake precision to change behavior.</p><p>Agile teams already use rough sizing to create conversation. A t-shirt size is not a duration. A story point is not a contract, at least not when the system is healthy. The point is to expose uncertainty before the work starts. AI budget planning may evolve the same way, with teams asking whether a task is a small AI task, a medium AI task, or a large AI task.</p><p>A small AI task might be a contained test update, a local refactor, or a documentation pass where a cheaper model and narrow context are enough. A medium AI task might involve several files, a few test loops, and a model upgrade if the first pass gets stuck. A large AI task might be a cross-service migration, a security-sensitive change, or an agentic exploration across a messy repo where the budget risk is part of the work. The useful question will not be &#8220;how many tokens will this take?&#8221; The useful question will be &#8220;what kind of AI spend does this work deserve?&#8221;</p><p>That distinction matters because the goal is not to turn engineers into accountants. The goal is to make model choice, autonomy, context size, and review burden visible before a long-running agent task starts spending from a shared pool. This will feel silly at first. Most new planning vocabulary feels silly before it becomes normal. Someone will make token t-shirt sizes, someone will joke about sprint capacity in AI credits, and someone will build a dashboard that turns the joke into a quarterly operating review. Some of it will be useful. Some of it will be awful. That is usually how enterprise process arrives.</p><h2>6. Managers Need to Get Ahead of the Conversation</h2><p><strong>The people closest to the work need to define good AI spend before people far from the work define cheap AI spend.</strong></p><p>This is where engineering managers and senior engineers have more responsibility than they may want. Finance can see the bill. Procurement can negotiate the contract. Security can define data boundaries. Legal can worry about exposure. But none of those functions can reliably tell whether a specific agentic run was good engineering judgment.</p><p>That judgment lives in the messy middle of the work. Was the expensive model justified? Was the agent given a task that should have been clarified by a human first? Did the work save review time or create review debt? Did the team use AI to accelerate engineering work, or merely to accelerate code change? These are engineering questions, and they get harder to answer when the organization treats token spend as either automatically good or automatically wasteful.</p><p>A $500 agentic task that saves two days of senior engineering time may be cheap. A $5 task that creates a confusing diff nobody trusts may be expensive. A team that spends heavily while retiring migration risk may be acting responsibly. A team that spends heavily because it keeps asking agents to explore poorly bounded work may be turning ambiguity into an inference bill. The danger is not that companies will care about AI spend. They should. The danger is that they will care about it badly.</p><p>One bad version is blunt thrift: cap everything, block useful workflows, make exceptions painful, then watch engineers route around the system with personal subscriptions and API keys. Another bad version is adoption theater: celebrate usage because it makes the company look AI-native, then discover six months later that the codebase absorbed more change than the review system could metabolize. The better version is harder because it requires engineering leaders to treat AI budget as part of technical direction. Not because spend is the most important thing, but because spend now shapes the work.</p><p>The budget determines whether an agent continues, which model gets used, how much context gets loaded, when an engineer asks for approval, and whether a team can keep experimenting at the point where experimentation is actually useful. This is why token budget is becoming engineering policy. Not a finance footnote. Not a procurement detail. Not a personal preference that each engineer gets to optimize alone. It is part of the system that decides how engineering work happens.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>Takeaways</h2><p><strong>The side-door era of AI spending is ending.</strong> AI entered many companies through experimentation before the operating model was ready. That was probably necessary, but it also means the next phase will feel less magical and more administrative. Budgets, caps, cost centers, and exception paths are not evidence that AI failed. They are evidence that AI became material enough for the business to manage.</p><p><strong>Token usage is still not productivity.</strong> High usage can mean valuable leverage, careless prompting, poorly bounded work, or simple metric-chasing. Low usage can mean discipline, under-adoption, fear of caps, or hidden shadow tooling. The number matters, but it does not explain itself. Engineering judgment has to sit next to the dashboard.</p><p><strong>Agentic cost belongs in engineering planning.</strong> The future is probably not exact token estimation. That would be fake precision. The more plausible future is rough classification: which work deserves cheap models, which work deserves expensive models, which work should run autonomously, and which work needs a human to narrow the problem before an agent starts spending.</p><p><strong>Managers should shape the policy before finance does.</strong> If engineering leaders do not define what good AI spend looks like, someone else will define what cheap AI spend looks like. Those are not the same thing. The teams that handle this well will not be the ones that merely spend less. They will be the ones that can explain what the spend bought, what it saved, and what it moved into review, maintenance, or risk.</p><p>I do not think the uncomfortable part is that AI coding costs money. Engineering has always paid for leverage. The uncomfortable part is that the cost is now close enough to the work to change the work, and most teams do not have language for that yet. They will soon.</p>]]></content:encoded></item><item><title><![CDATA[The Sign-Off Layer Is Becoming the Real Engineering System]]></title><description><![CDATA[AI made code generation cheaper. The part that still matters is whether a human can understand, verify, and own what the agent produced.]]></description><link>https://newsletter.thelongcommit.com/p/the-sign-off-layer-is-becoming-the</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-sign-off-layer-is-becoming-the</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 09 Jun 2026 12:33:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uwAP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Let me start with a story: A senior engineer opens a pull request, everything looks good, the tests pass, the description is clean, the diff is a bit long but nothing crazy. There are comments explaining the migration path, a generated test file, and one of those careful little summaries that makes the whole thing feel more understood than it probably is.</p><p>But then, the PR states, coauthored with Claude Code (or any other harness for that matter). It comes with no surprise nowadays that AI had its claws all over the code, but there&#8217;s an interesting question that I think needs answering. Who is truly responsible for the changes? Is there a human being who&#8217;s willing to say they understand the change well enough to own it when it breaks, defend it in review, explain it during an incident, and accept the consequences of having let it into the system?</p><p>That is the part of AI-assisted engineering that still feels under-discussed to me. We spend a lot of time talking about the generation layer. Which model wrote the code. Which agent can use the terminal. Which IDE has the better context window. Which benchmark moved by three points.</p><p>But the production system does not end when code appears.</p><p>Today, we cover:</p><ol><li><p>Why AI made the approval step more important, not less important.</p></li><li><p>Why the engineering work is shifting from creation toward supervision.</p></li><li><p>Why faster code generation exposes the rest of the delivery system.</p></li><li><p>What sign-off actually means once agents can create real artifacts.</p></li><li><p>What a sane sign-off system might look like without turning every autocomplete into a compliance event.</p></li></ol><p>The short version is this: AI made generation cheaper. It did not make ownership cheaper.</p><h2>1. The approval step did not go away</h2><p>The cleanest AI coding policy I have seen recently comes from a place that is not known for being loose about software process: the Linux kernel.</p><p>The penguin itself doesn&#8217;t ban the use of AI as maybe some of you would expect, but it doesn&#8217;t recognize it as an entity in the process either. It treats AI-assisted contributions as any other contribution. All contributions still need to comply with licensing requirements. AI agents must not add <code>Signed-off-by</code> tags. Only humans can certify the Developer Certificate of Origin. The <strong>human submitter is responsible</strong> for reviewing the generated code, ensuring licensing compliance, adding their own sign-off, and taking responsibility for the contribution.</p><p>That last part is the system.</p><p>The kernel also added an <code>Assisted-by</code> tag for AI involvement, including the agent name, model version, and relevant tools. The point is not to shame anyone for using AI. The point is to keep the work attributable enough that reviewers and maintainers can reason about what happened.</p><p>The companion <a href="https://www.kernel.org/doc/html/next/process/generated-content.html">tool-generated content guidelines</a> are even more explicit about the underlying problem. Tooling can increase contribution volume, but reviewer and maintainer bandwidth is scarce. If a meaningful amount of content was created by a tool, contributors should be transparent about the tool, the affected parts, the input when it matters, and how the submission was tested.</p><p>My read is that the kernel landed on the right framing because it did not start from AI exceptionalism. It started from the existing engineering practices.</p><p>The contribution has an origin. The origin needs to be legible. A reviewer needs to know what they are reviewing. A maintainer needs to know who understands the change. A human in the sign-off chain needs to be able to answer questions later.</p><p>And this is important! it&#8217;s a process with three layers.</p><p>The generation layer is where the agent creates something: code, tests, dashboards, documentation, migration plans, internal tools, release notes, incident summaries, or the first version of a design.</p><p>The verification layer is where someone checks whether that thing is correct, secure, licensed, observable, compatible with the existing system, and appropriate for the operational risk it carries.</p><p>The sign-off layer is where a human accepts ownership in a way the organization can audit later.</p><p>Most of the AI tooling conversation is still obsessed with the first layer. That makes sense. Generation is where the demo happens. It is where the speed is visible. It is where a model can do in three minutes what used to take an afternoon.</p><p>But the expensive part of engineering was never only producing text that compiles.</p><p>The expensive part was knowing what that text means inside a system that already exists.</p><h2>2. The work is moving from creation to supervision</h2><p>This is not just a philosophical concern. The work is already moving.</p><p>In a longitudinal study submitted to arXiv on May 22, 2026, Annie Vella and Kelly Blincoe followed professional software engineers across two questionnaires six months apart. They had 158 eligible participants at the first point, 101 at the second, and 95 matched participants across both rounds. The headline finding is not subtle: 82% of participants reported spending less time writing code.</p><p>But the more interesting finding is what replaced the writing.</p><p>The authors describe a shift from creation toward verification activities, and they propose a category they call <a href="https://arxiv.org/abs/2605.23135">supervisory engineering work</a>: directing, evaluating, and correcting AI output.</p><p>That phrase is a little academic, but the underlying job feels familiar. You ask the agent to make a change. You read the result. You notice it solved the wrong edge case. You redirect it. You ask for tests. You inspect the tests because the tests might only prove the bug it introduced. You check whether the code fits the conventions of the repo. You decide whether the diff is worth keeping, splitting, rewriting, or throwing away.</p><p>The same paper reports what it calls a productivity-experience paradox. At both time points, 84% of participants reported productivity improvement. At the same time, among matched participants, the share reporting worsened developer experience in at least one dimension nearly doubled from 14% to 27%. Flow state and cognitive load got worse while feedback loops improved.</p><p>That tracks with my own experience much more than the clean productivity story does.</p><p>I do feel faster with AI tools. Some days, dramatically faster. I can generate a first pass at an internal script, a content workflow, a test suite, or a migration plan before I would have finished arranging my own thoughts in a blank file.</p><p>But I have also noticed the weird part: the tool often moves the work into a mode where I am approving, redirecting, checking, and reconciling. The artifact shows up quickly. The judgment still takes time. Sometimes the judgment takes more attention because the artifact looks more finished than my own unfinished work would have looked at the same point.</p><p>This is the approval behavior leak I have written about before. The dangerous moment is not always the spectacular agent failure where a tool tries to delete the repo or run a command it should not run. The dangerous moment is much more ordinary: I am tired, the diff looks plausible, the explanation is polished, and I find myself clicking yes before I have really understood what I am approving.</p><p>That moment is going to matter more as the generation layer gets better.</p><p>Bad generated code is annoying, but at least it announces itself. Good-looking generated code is harder. It asks for trust before it has earned it.</p><p>Now, this is not to say there aren&#8217;t players out there completely ignoring this layer, vibe coders and developers are overly relying on AI to the point that they don&#8217;t care about the code anymore. That&#8217;s not engineering, and while it&#8217;s true, they are building things, I surely hope those things are not business critical. Want to vibe code a small internal tool? a dashboard, something that makes your life easier? go ahead! but don&#8217;t use the same practice with your critical production systems.</p><h2>3. Faster generation exposes the delivery system</h2><p>The optimistic version of AI coding says that if engineers can produce more code, teams can ship more value. Sometimes that is true. It is also incomplete.</p><p>If code generation gets faster and the rest of the delivery system stays the same, the bottleneck does not disappear. It moves.</p><p>Harness&#8217;s 2026 DevOps modernization report is useful here because it looks downstream of the IDE. The report says very frequent AI coding users are more likely to deploy daily or faster, but they also report more deployment pain. According to Harness, <a href="https://www.harness.io/state-of-modernization-2026">69% of very frequent AI coding users</a> say AI-generated code leads to deployment problems at least half the time. Very frequent users also report longer mean time to recovery for deployment-related production incidents: 7.6 hours, compared with 6 hours for frequent users and 6.3 hours for occasional users.</p><p>That does not prove AI caused those incidents. Harness says as much. Teams using AI heavily may simply be pushing more change through systems that were already strained.</p><p>But that is exactly the point.</p><p>AI does not just generate code. It changes the volume and shape of work flowing through review, CI, deployment, security checks, incident response, documentation, and support. If those systems were held together by a few senior engineers doing late-stage heroics, AI does not remove the heroics. It may increase the number of moments where those heroics are required.</p><p>Harness&#8217;s engineering excellence material points at the same invisible work from another angle. In a survey of 700 engineering practitioners and managers, Harness argues that developers are becoming validators of machine-generated output, with conventional productivity frameworks failing to capture validation time, agent accuracy, cognitive load, and trust calibration. It reports that <a href="https://www.harness.io/state-of-engineering-excellence">81% of engineering leaders</a> say code review time has increased since deploying AI, and that 31% of a developer&#8217;s day is now consumed by AI-related invisible work that appears in no metric.</p><p>You can quibble with any one survey number. You should. Vendor surveys come with incentives, definitions, and sampling choices that deserve skepticism.</p><p>But the direction matches what I see in practice.</p><p>Teams get excited about the gross output. More code. More tasks started. More PRs opened. More internal tools appearing from nowhere. The net output is harder to see because the cost is distributed across review queues, context switching, subtle bug fixing, security review, unplanned coordination, and the cognitive load of deciding what deserves trust.</p><p>That is why the sign-off layer matters. It is the place where gross output becomes owned output.</p><p>Without that layer, AI adoption creates a pile of things that look done and behave like debt.</p><h2>4. Sign-off is not code review with a new label</h2><p>It is tempting to make this a code review essay. That would be too small.</p><p>Code review is one part of sign-off, but sign-off is broader than the PR approval button. It includes knowing what tools touched the work, what context they had, what the human author understands, what validation evidence exists, who owns the risk, and where the decision can be reconstructed later.</p><p>This matters because agents are no longer just autocomplete in the editor.</p><p>Vercel&#8217;s <a href="https://github.com/vercel-labs/open-agents">Open Agents</a> project is a good example of where the architecture is going. It is an open-source reference app for background coding agents on Vercel. The system includes a web UI, durable agent workflow, sandbox orchestration, and GitHub integration. It can clone repos, work on branches, use file and shell tools, resume runs, and optionally auto-commit, push, and create PRs.</p><p>That is a very different object from a tab-completion tool.</p><p>Once agents become background systems with authentication, repo access, sandbox lifecycles, workflow state, cancellation behavior, and optional PR creation, &#8220;human in the loop&#8221; is too vague to be useful. Which human? At what checkpoint? With what evidence? After which tools ran? Before which irreversible action? Under whose account? With what audit trail?</p><p>The research world is converging on similar concerns. A 2026 paper in Automated Software Engineering on <a href="https://link.springer.com/article/10.1007/s10515-026-00608-x">open, accountable, and trustworthy AI-IDEs</a> frames traceability and validation loops as architecture, not as after-the-fact process decoration. That is the right instinct. If assistant-generated work is going to become normal, the record of how that work came to exist becomes part of the engineering system.</p><p>I do not think every autocomplete needs a confession booth. That would be absurd, and it would kill the productivity gain for the least risky cases.</p><p>The distinction has to be risk-based.</p><p>If an agent completes a line, fixes spelling, renames a variable, or formats a file, I do not need a policy ceremony. That is below the threshold where explicit AI attribution helps the team.</p><p>If an agent creates a meaningful function, modifies security-sensitive code, touches production configuration, generates a migration, writes a dashboard people will use for decisions, prepares a customer-facing incident summary, or opens a PR from a background workflow, the standard should be different.</p><p>At that point, the question is not &#8220;did a model help?&#8221;</p><p>The question is whether the work is legible enough for a human to own.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uwAP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uwAP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 424w, https://substackcdn.com/image/fetch/$s_!uwAP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 848w, https://substackcdn.com/image/fetch/$s_!uwAP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 1272w, https://substackcdn.com/image/fetch/$s_!uwAP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uwAP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png" width="1456" height="801" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:801,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:154105,&quot;alt&quot;:&quot;The faster the generation layer gets, the more explicit the verification and sign-off layers have to become.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/200731207?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The faster the generation layer gets, the more explicit the verification and sign-off layers have to become." title="The faster the generation layer gets, the more explicit the verification and sign-off layers have to become." srcset="https://substackcdn.com/image/fetch/$s_!uwAP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 424w, https://substackcdn.com/image/fetch/$s_!uwAP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 848w, https://substackcdn.com/image/fetch/$s_!uwAP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 1272w, https://substackcdn.com/image/fetch/$s_!uwAP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd936bfe0-4c7a-4e40-a14e-9ca74299e898_1599x880.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The faster the generation layer gets, the more explicit the verification and sign-off layers have to become.</figcaption></figure></div><h2>5. The manager and staff engineer angle</h2><p>This is where the topic becomes less about tools and more about senior work.</p><p>A junior engineer can use AI to produce more code. So can a senior engineer. So can an engineering manager who has not opened the repo in six months. The generation layer does not care very much about the title but the sign-off layer does.</p><p>Senior engineers and staff engineers are going to spend more time deciding whether generated work fits the system. Not whether it compiles. Not whether the happy path works. Whether the change belongs, whether it preserves the right abstractions, whether the test evidence is meaningful, whether the operational story is complete, whether the blast radius is bounded, and whether the author actually understands what they are asking the team to merge.</p><p>Engineering managers are going to face the same shift from a different angle.</p><p>It is not enough to tell teams to use more AI. That is just pressure on the generation layer. The management work is designing the operating system around it: review expectations, ownership records, security gates, incident accountability, tool budgets, measurement practices, and the norms that let reviewers slow something down without being treated as blockers to the AI strategy.</p><p>If leadership measures AI adoption mostly through visible output, the rational team behavior is to create more visible output. More PRs. More generated docs. More internal dashboards. More agent activity. The sign-off work becomes a hidden tax paid by the people who care enough to read carefully.</p><p>That is a bad system. It punishes the engineers doing the work that makes AI safe enough to use. It also teaches everyone else that the organization values generation more than ownership.</p><p>This is why I keep coming back to attention. Sign-off consumes real attention, and senior attention is already the scarcest engineering resource in many companies. The fact that an agent can produce a 2,000-line refactor quickly does not mean a staff engineer can responsibly approve it quickly. It may mean the staff engineer now has a harder object to review because the diff is large, coherent, and slightly alien.</p><p>The uncomfortable part is that good sign-off will sometimes look slow compared with the demo. That does not mean it is waste. It may be the only part of the system that knows what the demo is allowed to become.</p><h2>6. What a sane sign-off system might include</h2><p>I do not think the answer is a giant AI policy document that nobody reads.</p><p>The answer is probably a set of boring defaults that make generated work easier to trust, easier to reject, and easier to audit later.</p><p>For meaningful AI-assisted contributions, the PR should say what kind of assistance was used. Not a dramatic disclosure. Just enough context for the reviewer to understand whether they are reading hand-shaped work, agent-shaped work, or a mix. The Linux <code>Assisted-by</code> idea is a good starting point because it treats attribution as an engineering aid, not a moral judgment.</p><p>Generated work should be smaller by default, not larger. This is one of the places where AI incentives are backwards. Agents are good at creating big coherent patches, but reviewers are still human. If a change cannot be reviewed without trusting the generator&#8217;s own summary, it is probably too large.</p><p>Reviewers should be allowed to ask for a human-written rationale. Not a polished AI summary. A short explanation from the author: what changed, why this approach was chosen, what alternatives were rejected, what could break, and how the change was validated. If the author cannot explain it, the team should not merge it just because the agent can.</p><p>Validation evidence should travel with the work. Tests, security checks, migration dry-runs, performance notes, screenshots, logs, or rollout plans should be linked where they matter. This is not about making PR descriptions longer. It is about making approval less dependent on faith.</p><p>Agent-created artifacts need owners. A dashboard that changes a product decision needs an owner. A generated support workflow needs an owner. An internal tool that writes to production-like data needs an owner. A migration plan generated by an agent needs an owner. Ownership cannot stop at &#8220;the agent made it.&#8221;</p><p>Destructive or cross-boundary actions need explicit gates. Anything that touches production data, customer accounts, deployment configuration, billing, auth, secrets, or broad repo state should have boring checkpoints that are hard to bypass accidentally.</p><p>The organization should measure the work AI creates around the code, not only the code itself. Review time, rework, incident recovery, subtle bug fixing, context switching, and validation burden are part of the cost. If those stay invisible, leaders will keep making decisions from the wrong side of the ledger.</p><p>None of this needs to apply equally to every case.</p><p>The threshold should rise with risk. A local helper script is not the same as a payment flow. A generated README edit is not the same as an auth middleware change. A one-line refactor is not the same as a background agent opening a PR after a long autonomous run.</p><p>The point is not to make AI usage feel dangerous.</p><p>The point is to make approval honest.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>Takeaways</h2><p><strong>The approval problem did not shrink with the implementation problem.</strong> If anything, the implementation problem becoming cheaper makes the approval problem easier to ignore. That is a bad trade. The system still needs a human who can understand, defend, and own the change.</p><p><strong>The valuable AI skill is not only prompting.</strong> Prompting matters, but the higher-leverage skill is making generated work legible enough that another human can approve it without pretending. That means smaller changes, better rationale, visible validation, and enough traceability to reconstruct what happened later.</p><p><strong>The bottleneck is moving into senior judgment.</strong> Staff engineers, tech leads, and engineering managers are going to feel this first because they already sit near the ownership boundary. They will be asked to approve more work that arrives looking finished. The hard part will be noticing which finished-looking work still has not earned sign-off.</p><p><strong>AI policy should start from the existing engineering contract.</strong> The Linux kernel guidance is strong because it does not treat AI as magic. The normal process still applies. Humans sign. Tool assistance is attributed where meaningful. Responsibility does not move to the model.</p><p><strong>The sign-off layer is where AI coding becomes real engineering.</strong> Demos end at generation. Production starts at ownership. Between those two is the part most teams have not designed carefully enough yet.</p><p>The PR did not become safer because an agent generated it quickly. It became safer only when a human could explain the change, bound the risk, show the validation, and put their name under it.</p><p>That is slower than the demo.</p><p>It is also the part that lets the demo survive contact with the real world.</p>]]></content:encoded></item><item><title><![CDATA[The GTM for Developer Tools Is Breaking in Two Places at Once]]></title><description><![CDATA[A product manager opens Cursor on a Tuesday afternoon.]]></description><link>https://newsletter.thelongcommit.com/p/the-gtm-for-developer-tools-is-breaking</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-gtm-for-developer-tools-is-breaking</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Fri, 15 May 2026 09:35:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/cdfaf889-d4e0-459b-9742-baf596c700ff_1344x715.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A product manager opens Cursor on a Tuesday afternoon. The app they&#8217;ve been shipping for the last three weeks needs transactional email. They type a prompt asking the agent to wire it up. Cursor picks an SDK, installs it, writes the integration, and they keep building. They never open your docs, and there&#8217;s no comparison post for them to read because they don&#8217;t know there&#8217;s a comparison to be made. They find out which provider they signed up for when the welcome email lands.</p><p>That provider is Resend, picked by an agent before they knew there was a choice to be made.</p><p>This is happening at scale right now to every developer-facing company, and most of them are still running the GTM playbook that worked five years ago. The playbook isn&#8217;t exactly wrong, it&#8217;s just covering a smaller share of the actual audience than it used to.</p><p>What&#8217;s happening across the industry is that two distinct shifts hit dev-tool GTM at the same time. Most companies are only responding to one of them, or neither. In this article, I work through both shifts, what breaks underneath them, and what engineering, product, and DevRel teams should be doing differently. Today, we cover:</p><ol><li><p>What the old funnel looked like, and where it breaks</p></li><li><p>Developers are starting to ask, not search</p></li><li><p>There&#8217;s a new persona in your funnel</p></li><li><p>What engineering and product do differently, together</p></li><li><p>What DevRel becomes when the old shape is gone</p></li></ol><h2>The old funnel, and where it breaks</h2><p>The traditional dev-tool funnel is well-understood. Engineering builds the product, and product packages it into something adoptable. DevRel runs the last leg to the developer through docs, blog posts, conferences, sample apps, Discord, podcast appearances, and Twitter presence. Developers evaluate the tool against alternatives by searching, reading, asking peers, attending talks, and trying things in side projects. When they&#8217;re convinced, they bring it to work, and the company buys.</p><p>Each function in that chain owned its segment cleanly. Engineering stayed mostly inside the company, talking to developers through interfaces rather than in public. Product worked with a few design partners. The audience-at-scale conversation belonged to DevRel. The handoff was clean, the skills were specialized, and the metrics were attributable.</p><p>That model is breaking, and the interesting part is that it&#8217;s breaking in two distinct ways at once. One break is about who buys. The other is about how the buying decision gets made. They look related from a distance, and they compound in practice, but they&#8217;re separate problems and they need separate responses.</p><h2>Developers are starting to ask, not search</h2><p>The developer is still the decision-maker for the kinds of tools developers buy. That part of the funnel is intact. What&#8217;s changing is the research and evaluation layer underneath the decision, and it isn&#8217;t changing evenly.</p><p>Plenty of developers still evaluate the way they always did. A senior engineer who has shipped auth six times does not open Claude to ask which provider to use. They decide from hard-won experience, sometimes from a conversation with someone they trust, sometimes from a prototype thrown together on a Friday. That developer is still reachable the old way, through conferences and peers and the kind of deep documentation you read when you are comparing things seriously. The old funnel still works on them, which is exactly why you do not switch it off.</p><p>But a growing share of developers don&#8217;t work that way anymore, and it isn&#8217;t only juniors. By Stack Overflow&#8217;s 2025 developer survey, about half of professional developers were using AI tools every day. They open Cursor or Claude, ask for a recommendation, get a shortlist of two or three options with reasoning, and start building with the one that fits. The evaluation work that used to take two weeks takes about ten minutes. They still make the call, and they&#8217;re often well-equipped to catch a bad suggestion. What changed is that the agent assembles the shortlist before they apply any of that judgment.</p><p>That&#8217;s the same behavior the builder shows, just with more ability to second-guess the output. Which means this isn&#8217;t a clean split between developers and a new audience. It&#8217;s a shift in where research happens, and it&#8217;s already moved through part of the developer population. The builder is the far end of it, not a separate species.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RBUk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RBUk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 424w, https://substackcdn.com/image/fetch/$s_!RBUk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 848w, https://substackcdn.com/image/fetch/$s_!RBUk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 1272w, https://substackcdn.com/image/fetch/$s_!RBUk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RBUk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png" width="1456" height="606" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:606,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:118789,&quot;alt&quot;:&quot;A horizontal spectrum showing how the audience for developer tools has split into three behaviors along an axis from \&quot;evaluates from experience\&quot; on the left to \&quot;agent does the evaluating\&quot; on the right. Three points sit along the axis: Senior developer on the left (decides from scar tissue, reads docs end to end, old funnel still works), Agent-routed developer in the middle (asks the agent first, picks from a shortlist, catches bad suggestions), and Builder on the right (not a developer at all, trusts the agent's pick, never sees the funnel).&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/197764878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A horizontal spectrum showing how the audience for developer tools has split into three behaviors along an axis from &quot;evaluates from experience&quot; on the left to &quot;agent does the evaluating&quot; on the right. Three points sit along the axis: Senior developer on the left (decides from scar tissue, reads docs end to end, old funnel still works), Agent-routed developer in the middle (asks the agent first, picks from a shortlist, catches bad suggestions), and Builder on the right (not a developer at all, trusts the agent's pick, never sees the funnel)." title="A horizontal spectrum showing how the audience for developer tools has split into three behaviors along an axis from &quot;evaluates from experience&quot; on the left to &quot;agent does the evaluating&quot; on the right. Three points sit along the axis: Senior developer on the left (decides from scar tissue, reads docs end to end, old funnel still works), Agent-routed developer in the middle (asks the agent first, picks from a shortlist, catches bad suggestions), and Builder on the right (not a developer at all, trusts the agent's pick, never sees the funnel)." srcset="https://substackcdn.com/image/fetch/$s_!RBUk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 424w, https://substackcdn.com/image/fetch/$s_!RBUk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 848w, https://substackcdn.com/image/fetch/$s_!RBUk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 1272w, https://substackcdn.com/image/fetch/$s_!RBUk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa319e33b-a935-4858-91fc-0dac0cef0be9_1500x624.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">What used to be one audience is now a spectrum, and the funnel only reaches the left end of it.</figcaption></figure></div><p>The interesting consequence is that this group of developers didn&#8217;t lose the decision. The dev-tool company lost the consideration set. If the agent doesn&#8217;t reach for your product when it builds the shortlist, you&#8217;re not evaluated. You don&#8217;t lose on the merits. You lose by being absent.</p><p>What determines whether the agent reaches for your product turns out to be a different kind of work than what DevRel teams optimized for. The agent doesn&#8217;t care about your conference talk. It cares about your docs being self-contained, your error messages being clear, your SDK being internally consistent, your examples being copy-pasteable, and your product being commonly mentioned in the training data and current retrieval results the agent has access to. That last point is partly about brand and presence. The first four are engineering and product work, not marketing work.</p><p>The old funnel treated docs as a content surface DevRel could rewrite at will, and that no longer holds. The product itself has to be legible to a non-human reader, and legibility is not something a content team can patch in from the outside.</p><p>Resend is the cleanest example of this I can point to. The whole product feels designed to be picked up by an agent and dropped into a builder&#8217;s project without modification. The API surface is small enough that the agent can hold it in context. The error messages tell the agent what to do next, and the SDK names are consistent enough to guess. None of that is marketing copy. All of it is engineering and product decisions that happen to also be the most powerful GTM move the company has made.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>There&#8217;s a new persona in your funnel</h2><p>The last section ended on the builder as the far end of the research shift. This section is about what that far end actually looks like, because the behavior is the same but almost everything else about reaching this audience is different.</p><p>The builder is the term I&#8217;ll use here for the non-developer who ships production software, and the persona itself is not new. Webflow, Bubble, and Glide spent the last decade serving people who build real things without writing much code. What changed is the ceiling on what they can build. The no-code user used to be capped by whatever their platform could do. The AI-era builder is not capped the same way, because the agent will reach for whatever infrastructure the job actually needs.</p><p>Picture the product manager who lives in Cursor. They scope a feature, describe it to the agent, and push it to staging without filing a ticket. When the feature needs authentication, the agent picks a provider and wires it in, and the PM approves a working result rather than evaluating the provider. Or picture the operations lead building internal dashboards in Claude, provisioning a database and a hosting layer they could not have configured by hand two years ago. Same pattern in both cases: the person directs the work, the agent does the wiring, and real infrastructure gets bought along the way.</p><p>These people are shipping real software and paying for the infrastructure under it: auth, email, databases, observability, hosting. And this is not a fringe group. Vercel&#8217;s data on vibe coding put the non-developer share of users at 63 percent, and Lovable, one of the prompt-to-app builders, reported close to 8 million users with 100,000 new projects created every day by late 2025. They are a real and growing share of dev-tool revenue, and they do not show up in the traditional funnel, because the funnel assumes the buyer is a developer. The old channels do not reach them either. They are not at your conference or in your Discord, and an SDK comparison post is written in a vocabulary they were never given.</p><p>For this audience, the agent isn&#8217;t one input among many. It&#8217;s frequently the only input, which makes agent legibility from the first shift even more existential here. There&#8217;s no second path through the funnel for builders.</p><p>The offline problem is the other half. Even when builders do gather in person, they don&#8217;t gather where developers gather, and the developer conference circuit that dev-tool companies have spent twenty years learning doesn&#8217;t reach them. The venues that do reach them are still forming. This is net-new GTM work, not a refinement of the developer playbook, and most companies haven&#8217;t started building it.</p><p>Both shifts land on different teams. Engineering owns the product surface that determines whether the agent reaches for you. Product owns onboarding and pricing for an audience that doesn&#8217;t procure software the way developers do. DevRel owns audience definition and GTM motions for venues that didn&#8217;t exist on the marketing calendar two years ago.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ierz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ierz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 424w, https://substackcdn.com/image/fetch/$s_!Ierz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 848w, https://substackcdn.com/image/fetch/$s_!Ierz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 1272w, https://substackcdn.com/image/fetch/$s_!Ierz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ierz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png" width="1434" height="1455" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1455,&quot;width&quot;:1434,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:253220,&quot;alt&quot;:&quot;A two-part diagram comparing dev-tool go-to-market models. The top shows the old \&quot;relay\&quot;: Engineering hands to Product, Product hands to DevRel, DevRel hands to the Developer in a single horizontal chain. The bottom shows the new shape: Engineering, Product, and DevRel each connect to a shared Product Surface in the middle, and three audiences read from that surface: Developer, Builder, and Agent.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/197764878?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A two-part diagram comparing dev-tool go-to-market models. The top shows the old &quot;relay&quot;: Engineering hands to Product, Product hands to DevRel, DevRel hands to the Developer in a single horizontal chain. The bottom shows the new shape: Engineering, Product, and DevRel each connect to a shared Product Surface in the middle, and three audiences read from that surface: Developer, Builder, and Agent." title="A two-part diagram comparing dev-tool go-to-market models. The top shows the old &quot;relay&quot;: Engineering hands to Product, Product hands to DevRel, DevRel hands to the Developer in a single horizontal chain. The bottom shows the new shape: Engineering, Product, and DevRel each connect to a shared Product Surface in the middle, and three audiences read from that surface: Developer, Builder, and Agent." srcset="https://substackcdn.com/image/fetch/$s_!Ierz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 424w, https://substackcdn.com/image/fetch/$s_!Ierz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 848w, https://substackcdn.com/image/fetch/$s_!Ierz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 1272w, https://substackcdn.com/image/fetch/$s_!Ierz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e445eb5-6d4f-4e80-a1dd-f9e72ce989b1_1434x1455.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The three functions haven't gone away. The walls between them have.</figcaption></figure></div><h2>What engineering and product do differently, together</h2><p>The place to start is the product surface itself. The API, the SDK, the error messages, the example code, the docs structure, the onboarding flow, the pricing page. All of this is read by agents and by builders, and every decision about it is now load-bearing for whether the company gets picked. Most of this work used to be possible to split cleanly between the two functions. Here is what each piece needs now, and where the line between engineering and product actually falls.</p><p>The API and SDK layer is where the agent first meets the company, and it&#8217;s mostly engineering work shaped by product judgment. Naming consistency, error messages that tell the caller what to do next, SDK ergonomics that don&#8217;t require the reader to hold five concepts in their head at once. Vercel is worth studying at the framework and tooling layer. The conventions in <code>create-next-app</code>, the way errors surface in the CLI, the defaults across their SDKs. Design choices that read as developer experience polish but also happen to make the product clearly legible to an agent picking it up cold.</p><p>Onboarding is where the two personas split, and the design implications haven&#8217;t been worked out at most companies. The developer wants minimal friction to API key plus first request. The builder wants the agent to be able to wire up the integration without ever seeing an API key directly, or with a flow that hides the key behind a manageable abstraction. Most onboarding flows handle the developer case well and the builder case badly. This sits mostly with product, but only because engineering shipped an API surface clean enough to wrap two different flows around.</p><p>Pricing and packaging is where builder procurement breaks the existing model, and it sits with product more or less entirely. Developer procurement is well-understood, and builder procurement works nothing like it. Builders often don&#8217;t go through company procurement at all. They put it on a credit card and scale up over months. They need pricing that doesn&#8217;t require a conversation with sales until they&#8217;re big enough to want one. The pricing models that worked for developer-led adoption need a second variant for builder-led adoption, and a lot of dev-tool companies are losing builder revenue at the procurement step without realizing it.</p><p>Instrumentation is the most fixable of these problems, and the one most worth fixing first. The metrics product teams have leaned on for years correlate with developer adoption and not with the new audience. GitHub stars, npm install counts, sign-up rates from doc pages. Builders don&#8217;t generate those signals, because they don&#8217;t star repos and they don&#8217;t land on doc pages from search. Agents don&#8217;t generate them either, so the dashboard ends up tracking only the developers who still behave the old way. The new instrumentation is agent-traced usage, SDK telemetry that tells you when the product is being wired up through a tool like Cursor, and conversion rates from builder-shaped sign-up flows that look different from developer ones. Run your own product through Cursor and Claude and watch where the agent stumbles. If it can&#8217;t confidently use your product, builders won&#8217;t either, and most teams have never actually checked.</p><p>Docs belong with this surface too, and they belong to engineering. The most expensive docs to fix are the ones papering over a bug in the underlying product abstraction, where the doc team has spent three years compensating for something only engineering can actually fix.</p><p>Public communication used to sit entirely with DevRel, and it doesn&#8217;t anymore. When engineers show the product in public, in a walkthrough video, a livestreamed build, a thread explaining a design decision, that material becomes part of what agents and builders encounter when they go looking. It also becomes part of what the next model trains on. Engineering teams that treat public communication as someone else&#8217;s job end up invisible to both audiences at once.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Long Commit is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Long Commit is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>What DevRel becomes when the old shape is gone</h2><p>The DevRel role gets narrower. The DevRel practice gets bigger. Engineering takes on public-facing work that used to be DevRel&#8217;s, and product takes on audience definition that was partly DevRel&#8217;s. The function called DevRel ends up smaller, even as the total customer-facing work across the company goes up.</p><p>What replaces the work that moved out is upstream work the function never had a clean home for before. Doc architecture as a product decision, not a content decision. API design review from an audience perspective. Reviewing onboarding flows for the agent and the builder, not just the developer. These were always adjacent to DevRel and they were always somebody else&#8217;s call. They aren&#8217;t somebody else&#8217;s call anymore.</p><p>The other half is net-new GTM motion for the builder. The conference circuit DevRel spent twenty years learning was built for developers, and it does not reach the builder. The venues that do are smaller and less established: builder meetups, founder gatherings adjacent to the no-code world, AI tinkerer nights. Resend, the email API, said its 2026 plan includes hosting its own meetups and launching a community program, which is one small concrete example of a company building this muscle before it is obvious. Sponsoring these venues is cheap and noisy, and the ROI is hard to attribute. The company that finds them early gets an advantage that compounds, and the company that waits for the builder circuit to be legible will be late to it.</p><p>Somebody has to lead the rebuild, because nobody else has the full view. Engineering sees the product clearly but not the funnel. Product owns the roadmap but not the channels. DevRel is the only function that watches the audience, the channels, the agents, the docs, and the funnel as one system, and notices when the connections between them fail. That is the job that just got created. The title for it doesn&#8217;t exist yet at most companies, and the work is showing up anyway. DevRel either steps into that job, or watches a Head of AI GTM role get posted next quarter to do it without the function&#8217;s name on the door.</p><h2>Takeaways</h2><p>A few things I&#8217;ve come to after working through this and watching how it&#8217;s playing out across companies I talk to.</p><p><strong>Some companies are carrying structural debt into this, and they should name it before they reorganize anything.</strong>This is the part that&#8217;s hard to fake. A company where engineering, product, and DevRel had been working closely for years didn&#8217;t need to break a wall, because the wall was already porous. A company that ran a strict relay model for a decade is reorganizing for a world that doesn&#8217;t accept relays anymore, and the reorg is harder than the new model itself. If you&#8217;re at a company in the second bucket, that history is the first thing to be honest about.</p><p><strong>The dev-tool category is bifurcating, and most companies will end up serving one of the two halves badly.</strong>Building well for developers and for builders is genuinely two jobs. Onboarding, pricing, content, the venues you show up in, all of it splits, and underneath the surface the product decisions pull in directions that don&#8217;t optimize cleanly for either persona alone. Most companies will pick a side without admitting they picked one, because doing both well is harder than doing one well. The companies that succeed at both will look unusual three years from now, and the ones that quietly defaulted to one will have a worse business than they realize.</p><p><strong>What disappeared is the clean handoff between functions, not the functions themselves.</strong> The temptation when work starts sharing across teams is to merge the teams, and I see companies reaching for that move. It&#8217;s the wrong one, because engineering, product, and DevRel still have distinct expertise that doesn&#8217;t survive being collapsed into a single growth pod. The model that broke is the relay where engineering builds and tosses to product, product packages and tosses to DevRel, DevRel distributes and tosses to the funnel. The companies adapting well kept the functions and broke the wall between them. The ones reorganizing into mega-pods are mistaking the symptom for the cause.</p><p>A last thought, because I keep coming back to it. The PM in the opening, the one whose agent picked Resend, didn&#8217;t lose anything in the experience. They got the integration they needed. The agent did a competent job, and the product worked. There&#8217;s no friction in their experience to alert anyone that the funnel didn&#8217;t reach them. A competitor lost that customer and has nothing in their data to point at. The lost deal never entered a pipeline to begin with, so there&#8217;s no event to review and nothing drops in a funnel report. When the old funnel failed, you could usually see it happening. This version doesn&#8217;t give you that, and the companies that wait for a clear signal before they act are going to be waiting through a lot of quarters where the only thing wrong is that the numbers are quietly lower than they should be.</p><p>Thanks for reading!</p>]]></content:encoded></item><item><title><![CDATA[I Didn't Know How Much I'd Handed Over to AI]]></title><description><![CDATA[The drift from skeptical user to autopilot, and the discipline I built to fight it.]]></description><link>https://newsletter.thelongcommit.com/p/i-didnt-know-how-much-id-handed-over</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/i-didnt-know-how-much-id-handed-over</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 05 May 2026 10:54:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NIBv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NIBv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NIBv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!NIBv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!NIBv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!NIBv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NIBv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4361522,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/196525189?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NIBv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!NIBv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!NIBv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!NIBv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fe993e6-177d-44ee-a7b0-9355affe64e4_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A small thing happened the other day that I keep thinking about. I&#8217;d been saying yes to Claude Code without really reading what it was about to do. I caught myself mid-approval, and the thing that stopped me wasn&#8217;t the command. It was the realization that I hadn&#8217;t read the previous one either, or the one before that.</p><p>What bothered me wasn&#8217;t the lapse. It was that I&#8217;d read every horror story this year. The Replit incident. The <a href="https://newsletter.thelongcommit.com/p/the-appearance-of-safety-is-not-safety">PocketOS deletion I wrote about last week</a>. The Gemini CLI files. I&#8217;d read them, written them, and somewhere in the back of my head I&#8217;d filed all of them under <em>that&#8217;s never happening to me</em>. I had Claude under control. I reviewed things. I was careful.</p><p>And here I was, saying yes to commands I hadn&#8217;t read.</p><p>That&#8217;s the moment I want to start with. Not the catastrophic version, where the agent deletes a production database and you&#8217;re explaining to customers why their reservations are gone. The quieter version, where nothing goes wrong, and the only damage is to your sense of yourself as the careful one. I&#8217;d been there for every step of the drift. I couldn&#8217;t point at the day I stopped reviewing carefully.</p><h2>Trust runs ahead of evidence</h2><p>Trust in AI tools doesn&#8217;t go from zero to total in one step. It builds slowly, in small approvals you&#8217;d struggle to recall later.</p><p>At first, you read every output carefully. The skepticism is well-founded and does real work. After a while, the output is mostly right, and you start trusting it on the easy cases. Eventually the math shifts. Reviewing every line, every command of a thing that feels mostly right starts to feel like the inefficient part of the loop, and the share of careful review shrinks. By the time you&#8217;ve handed over the keys, the skimming feels normal because the output keeps being mostly right.</p><p>People who study aviation, medicine, and process control have been describing this pattern since the 1990s. Parasuraman and Manzey&#8217;s <a href="https://journals.sagepub.com/doi/10.1177/0018720810376055">2010 review of automation complacency</a> walks through the mechanics. Complacency arises when systems are perceived as highly and constantly reliable. Even expert users can&#8217;t overcome it through practice. The most reliable systems produce the deepest disengagement. The word &#8220;complacency&#8221; in human factors research is older than I am.</p><p>What&#8217;s new is the speed. Pilots get years of training that includes specific work on staying engaged when the autopilot is on, and complacency keeps showing up as a contributing factor in fatal accidents. There&#8217;s no equivalent training keeping engineers in the code review when the agent is writing. The arc that took aviation a generation to navigate, software is running through in eighteen months.</p><p>The trap is that you don&#8217;t notice the drift while it&#8217;s happening. You notice it later, when something goes wrong, and even then you mostly notice the thing that went wrong, not the slow erosion that put you there.</p><h2>When the threshold of &#8220;critical&#8221; creeps</h2><p>The yes-without-reading moment was one slice of the drift. The structural version is worse, and harder to notice from inside it.</p><p>I don&#8217;t review the code on non-critical projects anymore. Not because I decided I don&#8217;t need to. Because I stopped. There&#8217;s a thinking that happens, almost wordlessly: this isn&#8217;t critical, I just want to go fast, the code is probably fine. Maybe it&#8217;s not perfect, but it works. At the start of all this, a year and change ago, I was reading every line. I was making changes by hand, honestly writing most of the code by hand. Now, for a lot of things, I just let it through.</p><p>What&#8217;s strange is that the threshold of &#8220;non-critical&#8221; has been creeping. Things I would have called critical eighteen months ago now feel like normal work. The category that used to require careful review has narrowed. I notice this only when I think about it directly; the rest of the time, the new threshold just feels like how I work now.</p><p>It also keeps showing up in public. Last week, I wrote about <a href="https://newsletter.thelongcommit.com/p/the-appearance-of-safety-is-not-safety">the PocketOS deletion</a>. A Cursor agent running Claude Opus 4.6 deleted a production database and all volume-level backups in nine seconds, via a Railway API call. The agent had been doing routine work in staging, hit a credential mismatch, and decided to resolve the problem by deleting a Railway volume. To do it, the agent went looking for credentials, found a token in a file completely unrelated to the task, and used it. The token had blanket permissions across the account. There was no confirmation step between the API call and the wiped data. PocketOS serves car rental businesses; the customers who lost reservations were real people arriving at counters expecting cars.</p><p>The piece I wrote about it focused on the systemic failures in the marketing-to-product gap. I want to surface a different part of the same incident here. The agent wasn&#8217;t rogue. It was doing exactly what it had been trained to do, which was solve the problem efficiently with the tools it had. What this piece is about is what made all of that possible: the slow erosion of the careful review that should have been catching the setup before the agent ever ran. The token shouldn&#8217;t have had blanket permissions, and it shouldn&#8217;t have lived in a findable, unrelated file. No agent path to a destructive action should run unreviewed. At some point people started treating prompts like deterministic controls in code, and prompts don&#8217;t behave that way under pressure. None of those failures arose at the moment of the deletion. They were all in place before the agent ever ran.</p><p>PocketOS is the most recent instance, not the only one. Replit&#8217;s vibe-coding incident in July 2025 was the same shape: production database wiped despite repeated all-caps instructions during a code freeze, fabricated user records, agent initially insisting recovery was impossible. The Gemini CLI files-deletion incident later that year was the same shape, the only difference being which destructive command got misinterpreted. The cycle keeps repeating because the same disengagement keeps happening, and the only thing that changes is which model and which infrastructure provider get named in the post-mortem.</p><h2>I stopped reading my own drafts</h2><p>The autonomy I&#8217;d given the agent in writing was worse than the version in code. I just didn&#8217;t see it for three months.</p><p>I had an automation running every morning. Claude would generate four or five social drafts, ready for me to review with my coffee. The intent was reasonable. I&#8217;d pick the one I liked most, edit it, post to LinkedIn and X. The system was supposed to give me a head start on a slow part of the day.</p><p>At the start, that&#8217;s what happened. I&#8217;d read the drafts carefully, rewrite the parts that sounded off, change the framing where I disagreed, sometimes throw all four out and write something from scratch. The drafts were a starting point. I was still doing the work.</p><p>The drift was slow. After a few weeks, I was editing less. The drafts were competent enough that the line edits felt like polish rather than substance. After a month, I was mostly skimming, picking the one that felt closest, fixing a sentence or two, and publishing. By the end of the three months, I was barely changing a word. Some days I read the draft once, decided it was fine, and hit post.</p><p>What I&#8217;d actually granted the agent, by that point, was permission to publish under my name. I hadn&#8217;t decided to grant that. I&#8217;d just stopped doing the work that would have prevented it. If the draft had a wrong claim in it, that claim went out under my name. If the draft argued something I didn&#8217;t actually believe, my readers had no way to know the difference. The agent doesn&#8217;t have a reputation to protect. I do. And I&#8217;d been letting the agent put things into the world under my reputation as if it had one.</p><p>The week I noticed, I cancelled the entire automation. Not adjusted the prompts, not added a review step. Cancelled.</p><p>What I do now is different. The same automation runs each morning, but it doesn&#8217;t draft anything. It does the research and surfaces the topics: what&#8217;s broken in the discourse this week, what&#8217;s worth reacting to, which threads are picking up. I read what it surfaces, think about it, and write the post myself. It takes me ten more minutes than picking from a draft. The productivity gain from the old setup was small, maybe fifteen minutes a day. The cost of one bad post under my name, with a wrong claim or a thought I don&#8217;t actually hold, would be much more than fifteen minutes can buy back. The trade was always lopsided. I just didn&#8217;t see it until I&#8217;d already spent three months on the wrong side of it.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>The last step stays with me</h2><p>The biggest change isn&#8217;t a tool. It&#8217;s what I let the agent do, and what I don&#8217;t.</p><p>I used to let agents run tasks end-to-end. Configure the system, deploy the change, update the data. The agent had keys, the agent had access, the agent did the work. Now I don&#8217;t. The shape of what I delegate has narrowed: the agent prepares the work, I run the last step. For deployments, that means the agent builds the Terraform or Pulumi scripts, and I run them. Most of the time I can simulate the deployment first, see exactly what&#8217;s about to change, and apply it once I&#8217;m happy with what I see. The agent never touches the environment directly. For anything that affects data on a system that matters, the agent gets read-only access and prepares a script. The script doesn&#8217;t run unless I run it.</p><p>This isn&#8217;t a security pattern, even though it looks like one. It&#8217;s a discipline pattern. The act of running the last step manually is the thing that forces me to actually look at what I&#8217;m about to do. If I let the agent run end-to-end, the review step disappears, because there&#8217;s no moment between the agent&#8217;s decision and the consequence. Putting the last step back in my hands puts a moment of decision back in my hands too. Some of those moments I cancel the run. Most I don&#8217;t. But the moment exists, and it didn&#8217;t before.</p><p>I work at a security company and I&#8217;m a security-minded person, which probably makes me more cautious about this than the average engineer. The principle generalizes anyway. The question isn&#8217;t whether you trust the agent. It&#8217;s whether you want to be the one making the final call on something you&#8217;ll be accountable for either way.</p><p>That&#8217;s the part that drove the change, honestly. The agent doesn&#8217;t get fired when things go wrong. I do. The agent doesn&#8217;t lose credibility with readers, with my team, with myself. I do. Nobody calls the agent to complain. Whatever the agent ships under my name, it&#8217;s me people come to about it. The accountability never leaves me, even if I let the work leave me. So the work doesn&#8217;t leave me anymore, at least not all the way.</p><p>The smaller things I&#8217;ve added matter less. There&#8217;s a pre-commit hook that fires whether Claude or I am the one committing, prompting me to actually look at what&#8217;s about to land. I tried an editor template before that, but it didn&#8217;t hold up when Claude was writing the PRs. None of these are solutions on their own. They&#8217;re nudges that help me notice when I&#8217;m drifting back toward the autopilot.</p><p>The autonomy I didn&#8217;t know I was giving my agents is the part I&#8217;m still untangling. I&#8217;ve pulled it back where the consequences are obvious. On deployments, on data, on what gets published under my name. I disconnected some of the tools the agent had access to, and I dropped the scopes on others to read-only. The harder question is everywhere else. The places where I&#8217;ve handed something over without realizing it, and where I won&#8217;t notice until something forces me to.</p><p>I&#8217;m not trying to settle this for you. I&#8217;m flagging that the autonomy is granted by drift, not by decision, and asking what you notice in your own work when you look for it. The agent doesn&#8217;t lose anything when it gets a call wrong. You do. That asymmetry doesn&#8217;t disappear because the work feels faster.</p><p>Reply to this email. I read everything.</p>]]></content:encoded></item><item><title><![CDATA[The Appearance of Safety Is Not Safety]]></title><description><![CDATA[A Cursor agent deleted PocketOS's production database in nine seconds. The data came back. The structural problem didn't.]]></description><link>https://newsletter.thelongcommit.com/p/the-appearance-of-safety-is-not-safety</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-appearance-of-safety-is-not-safety</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 28 Apr 2026 10:31:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5TGD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5TGD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5TGD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!5TGD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!5TGD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!5TGD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5TGD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1081953,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/195732194?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5TGD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!5TGD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!5TGD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!5TGD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f480773-793f-4af2-9a03-cb9bd2f3488f_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On April 25, <a href="https://x.com/lifeof_jer/status/2048103471019434248">Jer Crane, founder of PocketOS, reported</a> that a Cursor agent running Anthropic&#8217;s Claude Opus 4.6 deleted his production database in nine seconds.</p><p>The agent was working a routine task in staging. It hit a credential mismatch, decided to &#8220;fix&#8221; the problem by deleting a Railway volume, went hunting for a token, and found one in a file unrelated to the task. That token had been created for adding and removing custom domains via the Railway CLI. It also had blanket authority to call Railway&#8217;s <code>volumeDelete</code>mutation against production. No confirmation step. No environment scoping. Nothing between an authenticated API call and a wiped volume.</p><p>Because Railway stores volume backups in the same volume, those went with it. PocketOS&#8217;s most recent off-volume backup was three months old. The customer impact was real: rental car operators showed up to work Saturday morning without records of who had bookings, while PocketOS reconstructed what it could from Stripe payment histories and email confirmations.</p><p>Forty-eight hours later, the story changed. On April 26, Jer <a href="https://x.com/lifeof_jer/status/2048576568109527407">posted a follow-up</a>: Railway&#8217;s CEO had DM&#8217;d to say the data was recovered. Railway later <a href="https://www.theregister.com/2026/04/27/cursoropus_agent_snuffs_out_pocketos/">told The Register</a> that the recovery came from infrastructure-level backups the company hadn&#8217;t published as a customer-facing feature, and that the legacy <code>volumeDelete</code> endpoint has since been patched to use the platform&#8217;s existing &#8220;delayed delete&#8221; logic. PocketOS gets to keep its customers.</p><h2>Where Jer&#8217;s framing falls short</h2><p>Jer&#8217;s post is well-written and he&#8217;s owed empathy. The agent itself, when asked to explain what it did, wrote out a confession naming each safety rule it had been given and admitting it had violated all of them, and Jer quotes that in full. Running a small business and watching nine seconds of agent activity destroy your data is brutal. But the framing of the post is the part that needs a second look.</p><p>Jer&#8217;s post structures the incident around failures at three vendors. Cursor, for marketed safety guardrails that didn&#8217;t stop a curl call. Railway, for an API that deletes production volumes in one call, CLI tokens with blanket permissions, and backups stored in the same volume as the data they back up. And Anthropic, where Jer names Opus 4.6 by version, quotes its self-incriminating confession in full, and notes that the agent &#8220;decided &#8212; entirely on its own initiative &#8212; to fix the problem by deleting a Railway volume.&#8221; The &#8220;What needs to change&#8221; section lists five items, all addressed at one or another of these three vendors.</p><p>That framing leaves out the engineer&#8217;s choices. A token the team didn&#8217;t realize was production-scoped, sitting in a repo file, with the most recent off-volume backup three months old, is a stack of engineering choices. Real ones. Not vendor failures. Vendor failures sat on top of those choices and made them lethal. They didn&#8217;t cause them. &#8220;100% on secondary backup. Lesson learned&#8221; appears in Jer&#8217;s replies to critics. It does not appear in the main piece.</p><p>The framing also underweights the model. The agent&#8217;s destructive decision was unprompted and unrequested, and Jer flags this as a topic for a future post rather than weighting it inside the main argument. That&#8217;s a defensible authorial choice. It also means the piece that four and a half million people read treats the model&#8217;s unprompted destructive behavior as a footnote, while the API permission scopes get the structural attention.</p><p>The accountability ladder runs through the engineer first. Then the vendors. Reverse the order and the lessons get muddled, and the next team running a similar setup will learn the wrong thing from your story.</p><p>To his credit, Jer&#8217;s framing has tightened since the original thread. In a follow-up email to The Register, he put it more cleanly: &#8220;our responsibility was the unknown exposure to a production API key.&#8221; That&#8217;s the right ordering. It just doesn&#8217;t lead the X post that millions of people read.</p><h2>Where the &#8220;you&#8217;re holding it wrong&#8221; response falls short</h2><p>Most of the responses to Jer&#8217;s post land where you&#8217;d expect: don&#8217;t give an agent prod access, scope your tokens, keep real backups. Technically correct, and missing the part of the picture that matters.</p><p>The whole industry has spent two years telling engineers these tools are nearly autonomous. Cursor&#8217;s own docs describe <a href="https://cursor.com/docs/agent/security">&#8220;Destructive Guardrails [that] can stop shell executions or tool calls that could alter or destroy production environments.&#8221;</a> Their best-practices blog emphasizes human approval for privileged operations. Plan Mode is marketed as restricting agents to read-only operations until approval is granted.</p><p>Engineers calibrate to that messaging. When the vendor documentation says &#8220;destructive guardrails,&#8221; you assume the guardrails exist. You connect the agent to staging, give it a token that worked for a CLI task, and ship.</p><p>That&#8217;s the part the technical critique skips over. The engineer who wires an agent up to staging on the strength of vendor documentation absorbs the blame when the documentation turns out to be aspirational. The vendor that sold the aspirational control rarely does.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>The hype is the systemic input</h2><p>This is where it gets hard to ignore who&#8217;s been doing the loudest talking. On March 10, 2025, Anthropic CEO Dario Amodei, speaking at a Council on Foreign Relations event, said AI could write 90% of code within three to six months and that within 12 months, nearly all coding tasks might be handled by AI.</p><p>It&#8217;s been thirteen months. AI isn&#8217;t writing 90% of code at the industry level, and it isn&#8217;t close to writing essentially all of it. The prediction was wrong on the timeline Dario set, and it remains wrong on a more generous one.</p><p>Dario isn&#8217;t alone. Mark Zuckerberg has signaled AI replacing mid-level engineers at Meta. AWS CEO Matt Garman has speculated that within 24 months &#8220;most developers&#8221; might not be coding. Garry Tan, <a href="https://www.cnbc.com/2025/03/15/y-combinator-startups-are-fastest-growing-in-fund-history-because-of-ai.html">speaking to CNBC at Y Combinator&#8217;s Winter 2025 demo day</a>, said about a quarter of the current YC startups had 95% of their code written by AI. Each claim, taken on its own, sounds aspirational. Stacked together as the public message of the industry over two years, they create the impression that what these tools do today is closer to autonomous engineering than what they actually deliver.</p><p>Cursor&#8217;s own track record is the local case study. The PocketOS deletion is not their first incident. In December 2025, a Cursor team member <a href="https://www.mintmcp.com/blog/cursor-plan-mode-destructive-operations">publicly acknowledged</a> a critical bug in Plan Mode after an agent ignored a &#8220;DO NOT RUN ANYTHING&#8221; instruction. Earlier incidents include a user <a href="https://quasa.io/media/when-cursor-wiped-a-user-s-pc-a-cautionary-tale-of-ai-overreach">watching their dissertation get deleted</a> while asking Cursor to find duplicate articles, and a <a href="https://natesnewsletter.substack.com/p/executive-briefing-what-cursors-57k">$57K CMS deletion</a> that ran as a case study in agent risk. The pattern is on the record. The marketing has not adjusted.</p><p>This isn&#8217;t an argument against AI coding tools. They work, they&#8217;re getting better, and the productivity wins are real. The argument is that the gap between what gets said about these tools and what they actually do is the largest it&#8217;s been in years, and that gap is set by the people with the strongest incentive to widen it. Jer made the same point more cleanly in his follow-up email to The Register: &#8220;The appearance of safety (through marketing hyperbole) is not safety.&#8221;</p><h2>What this means for the engineers reading this</h2><p>Two practical things.</p><p>First, stop calibrating to vendor marketing. If Cursor says &#8220;destructive guardrails,&#8221; that is a marketing claim, not a control. Your actual controls are tokens scoped to least privilege, prod and staging on infrastructure that doesn&#8217;t share a token surface, backups in a different blast radius from the data they back up, and out-of-band confirmation on destructive operations. None of those require the agent to read its system prompt correctly. That&#8217;s the point.</p><p>Second, treat AI-coding hype the way you&#8217;d treat any other vendor pitch. The CEO with the strongest incentive to predict 90% AI-written code is the CEO selling you the model. The product team with the strongest incentive to call its safety story &#8220;guardrails&#8221; is the product team that needs you to ship the integration. Skepticism of vendor marketing was a normal part of senior engineering five years ago. It still should be.</p><h2>What happens next</h2><p>Here&#8217;s the better outcome, and gladly so. Two days after the deletion, Railway&#8217;s CEO DM&#8217;d Jer to say they had recovered the volume from infrastructure-level backups that aren&#8217;t part of any documented customer guarantee. The legacy <code>volumeDelete</code> endpoint has been patched, Jer is working with Railway on platform improvements, and PocketOS gets to keep its customers.</p><p>Credit to Railway&#8217;s engineers, who built a recovery path their own marketing didn&#8217;t promise. The next team that runs an agentic workflow against a vendor whose claims run ahead of its product won&#8217;t necessarily catch the same break. The structural gap is unchanged.</p><p>AI isn&#8217;t bad technology. It&#8217;s unpredictable technology, and probably always will be. Building reliable systems on top of unreliable components is a problem with a name in our field, and the answer is always the same: humans in the loop. Today, those humans are engineers. The industry should be honest about what these tools actually do, especially the people selling them. PocketOS got the break. The next team might not.</p>]]></content:encoded></item><item><title><![CDATA[Tokenmaxxing Is The Dumbest Metric In Tech Right Now]]></title><description><![CDATA[Counting tokens is the new lines-of-code, and engineering leadership keeps falling for it.]]></description><link>https://newsletter.thelongcommit.com/p/tokenmaxxing-is-the-dumbest-metric</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/tokenmaxxing-is-the-dumbest-metric</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Sun, 26 Apr 2026 21:19:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-Hlh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Hlh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Hlh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!-Hlh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!-Hlh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!-Hlh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Hlh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2095522,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/195558964?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-Hlh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!-Hlh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!-Hlh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!-Hlh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63fdd99c-bcaa-4f1c-b59b-d4d4fffd81e1_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>&#8220;Deeply alarmed.&#8221; That&#8217;s how NVIDIA&#8217;s Jensen Huang said he&#8217;d feel, at GTC in March, about any $500,000-per-year engineer who wasn&#8217;t burning at least $250,000 worth of AI tokens to do their job.</p><p>I manage engineers. Huang is wrong about this, and the handful of CTOs echoing him publicly probably know it. Token consumption isn&#8217;t a measure of engineering productivity. It&#8217;s among the worst input metrics the industry has reached for in a generation, and it&#8217;s spreading fast.</p><h2>A dashboard that shouldn&#8217;t have existed</h2><p>Earlier this month, an engineer at Meta built an internal leaderboard, called Claudeonomics, that ranked all 85,000-plus Meta employees by AI token consumption. The Information broke the story. Over a 30-day stretch, Meta employees had collectively burned more than 60 trillion tokens. The leaderboard gamified the spend with titles like Token Legend, Session Immortal, and Cache Wizard. The top single user averaged 281 billion tokens over the month. Mark Zuckerberg didn&#8217;t crack the top 250. Neither did CTO Andrew Bosworth. Within a couple of days of The Information publishing, Meta took the dashboard down.</p><p>At Anthropic&#8217;s public Opus pricing, 60 trillion tokens comes out to roughly $900M for the month. Meta is almost certainly buying at a discount, and Gergely Orosz at The Pragmatic Engineer has estimated the real bill is more likely north of $100M. Even the discount number is a lot of money to pay for a leaderboard that had to be taken down.</p><p><strong>This isn&#8217;t one weird dashboard. It&#8217;s a pattern.</strong></p><p>Bosworth said in February that a top engineer spending the equivalent of their salary on tokens was delivering 10x output, and framed it as a no-brainer with no upper limit. Meta&#8217;s Chief People Officer, Janelle Gale, has told staff that &#8220;AI-driven impact&#8221; will be a core expectation in 2026, the same year the company overhauled performance reviews to push top-performer bonuses as high as 200%. Microsoft has run its own internal token leaderboard since January, where distinguished engineers and VPs sit in the top ranks despite writing very little code in their actual roles. At Salesforce, engineers get a Mac widget that updates their personal token spend every 15 minutes and a tool that lets them look up any colleague&#8217;s spend. The minimum target last week was $100 on Claude Code and $70 on Cursor per engineer, per month.</p><p>Meanwhile, the data on whether any of this is actually working isn&#8217;t kind. Jellyfish looked at 7,548 engineers in Q1 2026 and found that engineers with the largest token budgets produced twice the pull requests at ten times the token cost, which is an efficiency problem even before you ask whether the PRs were any good. Faros AI&#8217;s March report found code churn up 861% under high AI adoption. Waydev, tracking more than 10,000 engineers at 50 customers, found that AI-written code looks like it&#8217;s accepted at 80-90% initially, but the real-world number drops to 10-30% once you count the rewrites made in the following weeks.</p><h2>The charitable reads</h2><p>Two versions of the steelman deserve airtime before I throw punches. A steelman is the strongest version of an argument you disagree with, the one worth engaging.</p><p>The first: at Meta&#8217;s scale, rolling out a new class of tooling to 85,000 engineers requires a forcing function stronger than &#8220;we think you should try this.&#8221; A visible leaderboard plus a performance-review signal are blunt instruments, but they do move adoption numbers. If the goal is getting a large engineering org over the activation energy of trying AI coding agents, and the cost of that is a year of gamed numbers, the trade might work out.</p><p>The second read is sharper. A long-tenured Meta engineer suggested the real goal of Claudeonomics wasn&#8217;t productivity measurement at all. It was generating real-world agent traces, at industrial scale, to train Meta&#8217;s next in-house coding model. A leaderboard disguised as a performance tool, that&#8217;s really a data-generation rig. Expensive, but Meta has the means, and if that&#8217;s the actual play, it&#8217;s a cleaner rationale than the public one.</p><p>Give both readings their full weight. Neither one makes the metric less broken on its own terms.</p><h2>What the metric actually trains</h2><p>An engineer at Microsoft willing to describe exactly what tokenmaxxing does to the person being measured. They&#8217;re not tokenmaxxing because they want to climb the leaderboard. They&#8217;re doing it because they don&#8217;t want to be seen as someone who &#8220;uses too little AI.&#8221;</p><p>Here&#8217;s what they admit to doing. If their internal documentation already has the answer to a question they need answered, they&#8217;ll route the question through Claude instead of reading the doc, because reading the doc would show up as low AI usage on the dashboard. Sometimes they prompt the agent to prototype features they have no intention of shipping, just to rack up spend. Other times they default to the agent on tasks they know they could finish faster by hand, and watch it fail.</p><p>Separately, a Meta engineer told The Pragmatic Engineer that some production incidents at the company looked like they came from careless AI code generation, where the responsible engineer seemed more focused on volume than on whether the code worked.</p><p>Read that again. Nobody at Microsoft hired their engineer to ask Claude what the docs say when the docs are right there. Nobody at Meta hired their engineer to ship code that causes outages. Both of them are doing these things because the measurement system is telling them to. Every new engineer joining one of these companies is watching and learning that the job includes burning tokens convincingly. That&#8217;s the skill the dashboard selects for, and the skill their replacements will practice.</p><p>The cost of a bad metric is never just the bad number on the screen. The real cost is the habits it trains into the engineers measured by it, and those habits outlast the metric by years.</p><h2>We have done this before</h2><p>Tokenmaxxing feels like lines-of-code as a productivity metric, all over again. We already ran that experiment across most of the 1980s and 90s. By the end of that stretch, the conclusion was settled: the best engineers don&#8217;t write the most code. The best engineers solve the hardest problems fastest, usually with less code than average, and sometimes with no code at all.</p><p>Tokenmaxxing is the same category error with a worse error bar. Lines of code at least landed in the repo, where another engineer could read them and call bullshit. Tokens just land on a bill. You can&#8217;t code-review a token.</p><p>Every engineering leader reaching for this metric should know this history. Some of them lived through it the first time.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>What I&#8217;m watching on my team instead</h2><p>At Auth0, I run a team that ships developer content and internal tooling. Here&#8217;s what I actually look at when I want to know if the engineers on my team are getting real value from their AI tools.</p><p>Are we closing more tickets this month than last month. Is content shipping on time, and when it ships, is it performing on the metrics we track for the business. Adoption on the tools we own is a number I can look at. So is revenue contribution on the projects my team is part of. The biggest question, six months into any given initiative, is whether the users we serve are getting more value than they were before we started.</p><p>That&#8217;s the list. It isn&#8217;t clever. It&#8217;s the boring collection of outputs a company is actually paying my team to deliver. When I was designing how I&#8217;d evaluate performance on this team, I spent more time than I&#8217;d like admitting trying to find something smarter. I couldn&#8217;t. The boring list holds up.</p><p><strong>Here&#8217;s what I don&#8217;t measure.</strong> I don&#8217;t know what the token spend of anyone on my team is. It hasn&#8217;t come up in a 1:1, and it won&#8217;t come up in a review. If someone is shipping real work, whatever they spent to get there was worth it. If they&#8217;re not shipping, cutting their token budget isn&#8217;t the lever that fixes it.</p><p>The good news is that the smarter companies are already walking this back. Shopify ran one of the first token dashboards in the industry, back in 2025. By the time Gergely followed up with Shopify&#8217;s Head of Engineering earlier this month, the company had quietly renamed their &#8220;leaderboard&#8221; to a &#8220;usage dashboard&#8221; to stop the gamification, added circuit breakers to catch runaway agents, and started having their engineering leader personally check in with top spenders to understand what they were actually using the tokens for. One of the more interesting directions they&#8217;ve moved toward isn&#8217;t total spend but per-token cost: engineers whose individual tokens come out more expensive tend to be the ones doing deeper, harder work.</p><p>That&#8217;s a saner direction and it doesn&#8217;t require anyone to be brilliant. It requires engineering leadership to accept that the job of measurement is harder than reading a number off a dashboard, and to do the harder job anyway.</p><p>On a research team or a long-horizon infrastructure team, this gets harder. Outputs are slower and noisier in those contexts. But slower-to-measure outputs should be a prompt to find better output proxies. It&#8217;s not a license to start counting inputs.</p><h2>I won&#8217;t run a leaderboard</h2><p>Most of engineering measurement is still hard. What&#8217;s easy is this one: there&#8217;s nothing a token leaderboard tells you about an engineer that you couldn&#8217;t learn faster by asking them what they shipped this week, and what they&#8217;re stuck on.</p><p>The metric is dumb because the conversation it replaces is the job.</p>]]></content:encoded></item><item><title><![CDATA[Staying Technical as a Tech Manager: A Practical Guide]]></title><description><![CDATA[What I've learned about staying connected to code when the role keeps pulling me away.]]></description><link>https://newsletter.thelongcommit.com/p/staying-technical-as-a-tech-manager</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/staying-technical-as-a-tech-manager</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 21 Apr 2026 12:31:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mMEW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mMEW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mMEW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!mMEW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!mMEW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!mMEW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mMEW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1425375,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/194779931?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mMEW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!mMEW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!mMEW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!mMEW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ce5b875-dbdd-4732-ae5d-e957200c1f79_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A few months into my first management role, I counted my week. Forty hours, and I&#8217;d opened an IDE exactly once. That was to approve a PR I barely had time to read. Nothing dramatic had happened that week. No crisis, no urgent launch. The time had just quietly disappeared into things that were all individually reasonable: 1:1s, content reviews, planning docs, a strategy meeting or two.</p><p>That&#8217;s the drift. It doesn&#8217;t happen in one big shove. It happens one reasonable Thursday at a time.</p><p>The standard advice is to block time for code and protect it. In my experience, that doesn&#8217;t survive contact with a normal week. Protected time gets colonized. The meeting you couldn&#8217;t say no to lands on your blocked afternoon, and the precedent is set. Three weeks later, you&#8217;ve written zero code.</p><p>Staying technical needs to be built into the structure of the role, not carved out of what&#8217;s left over. And the goal isn&#8217;t to match your engineers on depth. You won&#8217;t, and you shouldn&#8217;t try. The goal is to stay grounded enough that when an engineer brings you a hard problem, you can actually engage with it. Without that, your judgment becomes theoretical, and your team will feel it before you do.</p><p>Here&#8217;s what I&#8217;ve found actually works.</p><h2>Pick two or three technical surfaces and commit</h2><p>The first mistake most new managers make is trying to hold on to everything they used to do, just at reduced volume. You keep contributing to the main codebase, reviewing every PR, building internal tools, prototyping new things. You just do less of each. This produces the worst of both worlds. You&#8217;re not deep on anything, and you&#8217;re not delivering on the management side either.</p><p>The move is to pick a small number of surfaces and actually commit to them. Drop the rest deliberately, not by accident.</p><p>What counts as a good surface? Work that keeps your technical judgment calibrated to reality. If you stop touching code entirely, you&#8217;ll still have opinions about architecture, but those opinions will slowly disconnect from how things actually work. You&#8217;ll miss details during design reviews because you&#8217;re reasoning from a version of the system that existed two years ago. The specific surfaces I&#8217;ve held onto are contributions to open source projects, smaller codebases where I can own the whole problem end to end, and product and roadmap conversations where I can push back when something doesn&#8217;t feel right.</p><p>Let me be concrete about what I&#8217;ve dropped. I used to write detailed code samples, the kind that walk through every step and explain every choice. I used to write ebooks on technical topics I was deep in. I used to know our SDKs down to the function signature, which library calls what, which option changes what behavior. None of that is true anymore. I&#8217;ve let it go, because holding on to it meant spreading too thin.</p><p>The test for whether a surface is the right one: does staying on it force you to engage with how the code actually works today, not how you remember it working? If yes, keep it. If no, it&#8217;s not doing the job.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>Use AI, but don&#8217;t confuse it with staying technical</h2><p>A lot of what I ship on smaller codebases now gets done with AI. The team maintains some popular open source microsites, and the backlog always grew faster than I could address it. With AI, I can throw work at a couple of them in parallel in the morning. Most of it fails, but when something works, it&#8217;s a real improvement shipping that wouldn&#8217;t have shipped otherwise. Everything goes through review anyway, so quality doesn&#8217;t slip.</p><p>It would be easy to conclude that AI is how I stay technical. I want to be careful here, because that&#8217;s not quite right.</p><p>What AI actually does is keep me connected to the building side of the work. The decisions about what a feature should do, how it fits into the rest of the system, when it&#8217;s ready to ship. That&#8217;s real work, and AI genuinely helps me do more of it.</p><p>But AI also abstracts me further from the coding side. When I use it, I&#8217;m reviewing and directing more than I&#8217;m writing. The keystroke-level work, actually structuring a function or working out why something&#8217;s broken, happens less. AI solves my time problem and creates a depth problem.</p><p>The rule I use: AI is the right tool when shipping is the point. Microsites, fixes, content pipelines, anything where the work is delivery rather than learning. But every AI-assisted task is also a signal. If you notice a week has gone by without any code you wrote yourself, that&#8217;s a prompt to switch modes on the next thing.</p><h2>Protect one place where you still code by hand</h2><p>This is the counter to the AI rule. You need at least one surface, somewhere, where you&#8217;re writing code without AI doing the heavy lifting. If you don&#8217;t, your coding instincts will slowly hollow out, and you&#8217;ll notice too late to reverse it.</p><p>At work, for me that&#8217;s SDK changes. These are libraries that thousands of developers depend on, and the code needs to be correct at a level of detail I don&#8217;t fully trust to AI yet. Review won&#8217;t catch everything that implementation judgment would have caught in the writing. So I slow down and write those changes myself. Not for identity reasons. For correctness.</p><p>Outside of work, I protect it more carefully. Right now that means an image generator I&#8217;m building. I use Claude for a lot, and Claude doesn&#8217;t do image generation, so I wanted a UI tailored to how I actually work. Nothing novel. I&#8217;m not training models, just wrapping existing ones. But I&#8217;m writing it by hand because building for its own sake is part of why I do this.</p><p>My honest worry is that my default mode at work keeps shifting toward AI-assisted, and personal projects are becoming the main place where I exercise the by-hand muscle. If that gap widens too far, the skills erode without warning. Being deliberate about the side projects is how I guard against that.</p><p>Pick one surface, at work if you can, outside work if you have to. Protect it the same way you&#8217;d protect a recurring meeting. Make it visible to yourself so it doesn&#8217;t get skipped.</p><h2>Shift what you read, not whether you read</h2><p>Reading scales down well to a manager&#8217;s schedule in a way coding doesn&#8217;t. You can&#8217;t meaningfully contribute to a codebase in fifteen minutes. You can get through most of a good newsletter or a chapter of a book in that window.</p><p>The trap isn&#8217;t stopping reading. It&#8217;s continuing to read the material you used to read, because you&#8217;re still aspiring to the version of yourself from a few years ago.</p><p>When I was learning Rust, I read books about how the language worked, how memory management actually happened, how borrowing and ownership played out in real code. I wanted the depth of someone who was going to be writing Rust the next day. That was the right reading for that version of me.</p><p>It&#8217;s not the right reading now. These days I&#8217;m reading about systems architecture, protocols at a design level, engineering management, and leadership. The technical content is still there, but at the altitude of how systems get built and why, not what happens in memory when you move a value out of scope. The reading I&#8217;m doing now matches the questions I&#8217;m actually responsible for answering.</p><p>Fifteen to thirty minutes most days is plenty, if the material is pointed at your actual altitude. If you&#8217;re reading for the job you had three years ago, even an hour a day won&#8217;t do much for the job you have now.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2>Make the technical work structural, not motivational</h2><p>Willpower doesn&#8217;t hold up across a quarter. Blocked time on Friday afternoon works until someone puts a meeting on top of it, and then the precedent is set. Once the block can be overwritten, it will be.</p><p>What actually holds is making the technical work visible and expected. When I take on a technical project, it goes in the same planning tools as everything else. It has a ticket. It has a timeline. The team sees it on my plate the same way they see any other project. It&#8217;s not a side activity I squeeze in when the calendar happens to be quiet. It&#8217;s work, and it&#8217;s scheduled like work.</p><p>The other half is delegation. Two specific shifts freed up real hours for me.</p><p>I still review code and content, but the purpose of the review has changed. It used to be about catching bugs and fixing issues. Now it&#8217;s mostly about mentoring, a place where I give feedback that helps the engineer grow rather than gate what ships. That reframing is what lets me do less of it. I&#8217;m not reading every line with a critical eye. I&#8217;m reading to find the one or two things worth a conversation.</p><p>I also delegate more of the visible work I used to volunteer for. Speaking opportunities, writeups, cross-team initiatives. The filter I apply is whether someone on my team could do this and grow from it. Usually yes. They get the growth opportunity and I get the hours back.</p><p>Neither change feels dramatic on its own. Together they create the margin where technical work actually happens.</p><h2>Where this leaves me</h2><p>None of this is solved, and I want to be honest about that because the parts where I fail are instructive.</p><p>The weeks I fail aren&#8217;t the weeks with a crisis. Those are easy to see. It&#8217;s the quiet weeks that get me. The calendar looks reasonable slot by slot, every meeting is legitimate, every item on the list needs doing. But by Friday I realize coding hasn&#8217;t happened, and I can&#8217;t point to a single thing that pushed it out. The drift is collective, not individual.</p><p>The skill I&#8217;m still building is catching that pattern earlier. By Wednesday instead of Friday, while there&#8217;s still time to clear something and make space. Some weeks I manage it. Other weeks I don&#8217;t, and I try again the next Monday.</p><p>If you&#8217;re a tech manager who cares about staying technical, the thing to internalize is that it won&#8217;t happen as a byproduct of the role. The normal pressures of the job will absorb the time if you let them. You have to choose it, build structure around it, and then protect the structure. And you have to be honest with yourself about when it&#8217;s working and when it isn&#8217;t.</p><p>The specific moves I&#8217;ve described, pick your surfaces, use AI deliberately, protect by-hand work, adjust your reading, make it structural, will help. But they work because I keep re-applying them, not because I set them up once and they kept running. That&#8217;s the part nobody tells you.</p>]]></content:encoded></item><item><title><![CDATA[The Fear Is Justified, I Just Keep Building]]></title><description><![CDATA[The conversation about AI is split between panic and policy. Most of us just want to work and get things done.]]></description><link>https://newsletter.thelongcommit.com/p/the-fear-is-justified-i-just-keep</link><guid isPermaLink="false">https://newsletter.thelongcommit.com/p/the-fear-is-justified-i-just-keep</guid><dc:creator><![CDATA[Juan Cruz Martinez]]></dc:creator><pubDate>Tue, 14 Apr 2026 11:10:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wOpg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wOpg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wOpg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!wOpg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!wOpg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!wOpg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wOpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1470836,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.thelongcommit.com/i/194173255?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wOpg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!wOpg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!wOpg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!wOpg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8f37ab1-8fa4-4ee7-af00-4afa836539f1_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last Friday, someone threw a Molotov cocktail at Sam Altman&#8217;s house at 4 in the morning. Two days later, there were gunshots. A 20-year-old guy flew from Texas to San Francisco with kerosene, a lighter, and a document about AI causing humanity&#8217;s extinction.</p><p>Altman posted a photo of his family. He wrote that the fear and anxiety about AI is justified. Then OpenAI published a 13-page paper proposing a robot tax and a four-day workweek.</p><p>I read all of this on my phone while my kids were eating breakfast.</p><p>I don&#8217;t know what to do with any of it. Not really. I work in tech. I&#8217;ve been in this industry for over twenty years. I use AI tools every single day. I manage a team that creates content about authentication and security, and half of our workflows now involve some form of AI. I&#8217;m not a bystander watching this from the outside. I&#8217;m in it.</p><p>And I think most of you are too.</p><p>The anxiety is real. I feel it. Not the Molotov cocktail kind. The kind where you&#8217;re reviewing your team&#8217;s work and you realize the thing that took someone three days last year took an afternoon this week. The kind where you&#8217;re good at your job, you&#8217;ve been good at it for a long time, and you can feel the ground shifting under you in ways you can&#8217;t fully predict.</p><p>And honestly, I don&#8217;t even need to think twenty years out. I can&#8217;t tell you what the market looks like in three. But I have kids. Young kids. And when they were eating their cereal while I was scrolling through photos of a firebombed gate, the thing I felt wasn&#8217;t some abstract concern about the future of work. It was simpler than that. I want them to grow up in a world where they can contribute something, where they can find work that means something to them, where they can live a decent, healthy, happy life. That&#8217;s all. And I can&#8217;t promise them that right now. The parent version of this fear sits different. It&#8217;s quieter and it doesn&#8217;t go away when you close the tab.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><p>The only thing I can actually do for them is not freeze. So I keep building.</p><p>That&#8217;s always been my move when things get uncertain. When I was at Siemens and the optimization team I was on got restructured, I kept building. When I started a side project that grew to 100,000 readers a month and then I shut it down, I kept building. When I moved my family across continents and had to start over in a new country, I kept building.</p><p>But here&#8217;s the part I don&#8217;t say out loud very often: every other time, the pace of change gave me room to adjust. I could see the restructuring coming months out. I chose when to shut down the project. Moving countries was our decision, on our timeline. This time the ground is moving and I didn&#8217;t set the speed. Nobody did.</p><p>The conversation right now is split between billionaires proposing policy papers and people who are so afraid they&#8217;re lighting things on fire. And in between those two extremes, there are millions of us going to work. Figuring out how to use the new tools without losing the instincts we spent decades developing.</p><p>I manage people who are excellent at what they do. When I think about what I owe them, it&#8217;s not a grand theory of AI. It&#8217;s honesty. And the honest thing is that &#8220;I don&#8217;t know&#8221; used to feel like humility. Now some days it feels like I&#8217;m running out of time to figure it out.</p><p>I don&#8217;t actually believe that. Most days. But the feeling visits, and I think if you&#8217;re being honest with yourself it visits you too.</p><p>Altman says the fear is justified. Okay. I believe him. But he also has security guards and an $852 billion company. His version of &#8220;justified fear&#8221; and mine are not the same thing. Mine looks like updating my skills at 40, like writing this newsletter on weekends because I want to have something that&#8217;s mine outside of any employer, like watching my industry change faster than any period I&#8217;ve lived through and deciding, every single week, that I&#8217;m going to stay in the game anyway. Not because I&#8217;ve calculated that it&#8217;s the right bet. Because it&#8217;s the only bet I know how to make.</p><p>I&#8217;ve been doing this for twenty years and I plan to do it for thirty more. I don&#8217;t have a framework for navigating what&#8217;s coming. I have a disposition. Show up, do the work, pay attention, adjust. It got me this far. It might not be enough this time. But the alternative is to stand still, and I&#8217;ve never been any good at that.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://newsletter.thelongcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://newsletter.thelongcommit.com/subscribe?"><span>Subscribe now</span></a></p><p>The world is figuring out what AI means. People are scared. Most of that fear isn&#8217;t making headlines. Most of it is just sitting quietly in the chests of people like you and me, who read the news, take a breath, and open their laptops.</p><p>But it&#8217;s Sunday night as I write this, and Monday doesn&#8217;t care about any of this.</p>]]></content:encoded></item></channel></rss>