<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Dimitri Sudomoin</title>
    <description>Making AI work for real businesses - not just demos.</description>
    
    <link>https://www.sudomoin.com/</link>
    <atom:link href="https://rss.beehiiv.com/feeds/JPP2QM3aYU.xml" rel="self"/>
    
    <lastBuildDate>Sat, 12 Sep 2026 03:47:16 +0000</lastBuildDate>
    <pubDate>Fri, 04 Sep 2026 18:23:22 +0000</pubDate>
    <atom:published>2026-09-04T18:23:22Z</atom:published>
    <atom:updated>2026-09-12T03:47:16Z</atom:updated>
    
      <category>Machine Learning</category>
      <category>Artificial Intelligence</category>
      <category>Technology</category>
    <copyright>Copyright 2026, Dimitri Sudomoin</copyright>
    
    <image>
      <url>https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/publication/logo/67cb5949-a05d-47dc-b2e8-78ad9286e24e/DALL_E_2023-12-04_09.13.17_-_A_minimalistic_and_clean_flat_design_style_logo_of_a_stylized_emerald-colored_pyramid._The_pyramid_gradually_transforms_into_a_digital__pixel-like_str.png</url>
      <title>Dimitri Sudomoin</title>
      <link>https://www.sudomoin.com/</link>
    </image>
    
    <docs>https://www.rssboard.org/rss-specification</docs>
    <generator>beehiiv</generator>
    <language>en-us</language>
    <webMaster>support@beehiiv.com (Beehiiv Support)</webMaster>

      <item>
  <title>I Haven&#39;t Lost a Customer Service Fight in Seven Months</title>
  <description>How consumer claims became the highest-ROI thing I do with AI agents, from a $10 mac-and-cheese coupon to a $6,800 IRS reversal.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/540cffb6-91d9-4ad3-8d9b-1f63d5342c9a/2026-09-01-consumer-claims-cover.png" length="786498" type="image/png"/>
  <link>https://www.sudomoin.com/p/consumer-claims</link>
  <guid isPermaLink="true">https://www.sudomoin.com/p/consumer-claims</guid>
  <pubDate>Fri, 04 Sep 2026 18:23:22 +0000</pubDate>
  <atom:published>2026-09-04T18:23:22Z</atom:published>
    <dc:creator>Dimitri Sudomoin</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;">It started with a box of Annie&#39;s mac and cheese. Three boxes, actually, a Costco three-pack, and not one of them had a cheese packet. The kitchen is chaos because kids are hungry, and I&#39;m standing there looking at the box of dry pasta thinking about what a consumer is supposed to do here. Go back to Costco, stand in line, (wait... did you remember to bring the box?) - for six dollars? Write to Annie&#39;s? I&#39;m not writing to Annie&#39;s. Every version of &quot;do something about it&quot; costs me more than the problem.</p><p class="paragraph" style="text-align:left;">So I would have eaten the cost, the way I&#39;ve eaten every one of these my entire adult life. But instead I took a picture of the box and told Claude to deal with it. It asked where I bought it and when, found the manufacturer&#39;s complaint form, filled it in through the Chrome extension with the photos attached, and submitted it. A few weeks later a coupon showed up in the mail. So whats there to write about?</p><p class="paragraph" style="text-align:left;">Well, I knew going in that the coupon wasn&#39;t the point. I was spending family capital on a $10 coupon because I could see it was the seed for a process. Companies were going to keep not holding up their end, that&#39;s just an increasingly common fact of life, and I was going to keep being at a disadvantage every time. Except now there was a way to counteract that!</p><p class="paragraph" style="text-align:left;">Seven months later that process is ten for ten, no losses, though a few are still in flight. It has clawed back somewhere north of <b>twelve thousand dollars</b> in refunds, reversals, and costs I would otherwise have eaten. And it&#39;s the highest-ROI thing I&#39;ve done this year with AI agents in my personal life.</p><h2 class="heading" style="text-align:left;" id="the-scoreboard-12000">The scoreboard ($12,000)</h2><div style="padding:14px 15px 14px;"><table class="bh__table" width="100%" style="border-collapse:collapse;"><tr class="bh__table_row"><th class="bh__table_header" width="50%"><p class="paragraph" style="text-align:left;">Fight</p></th><th class="bh__table_header" width="50%"><p class="paragraph" style="text-align:left;">Outcome</p></th></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">IRS, 2017 late-filing penalties</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">$6,813 in penalties removed with one prepared phone call. I&#39;d been paying $500/month on an installment plan. Got sent a $2,404 check.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Subaru, CVT torque converter</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Replaced under an extended warranty three weeks before it expired. Dealer&#39;s opening move was a $249 &quot;customer-pay&quot; service; waived on pushback. Roughly a $2,000+ job out of pocket.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Solar installer, panel layout</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">They redesigned my array last minute after a financing partner banned trimming vent pipes. They ended up paying a roofer to relocate the pipes instead. Lets call it $1,500.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Pet insurance</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Billed $359 in &quot;earned premium&quot; on two lapsed policies, then spent two months threatening me with collection. Settled at $118 with a written promise nothing ever went to a bureau.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Health insurer, mammogram</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">$1,153 screening for my wife, who has a family history, auto-denied as &quot;non-covered.&quot; Reversed.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Health insurer, lab work</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">$276 in cholesterol tests denied as &quot;experimental.&quot; Internal appeal lost, external review won.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Plaud, voice recorder</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">$179 device died from a waist-height drop, inside their own published drop spec. &quot;Drop damage not covered.&quot; Free replacement.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Boulder Organic soup, CAVA, Annie&#39;s</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">A $25 gift card, plus a $10 Costco refund, a $14 refund, a $10 coupon.</p></td></tr><tr class="bh__table_row"><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">Robot mop</p></td><td class="bh__table_cell" width="50%"><p class="paragraph" style="text-align:left;">$1,166, the one still in play. More on this below.</p></td></tr></table></div><p class="paragraph" style="text-align:left;">The IRS one dominates the total. First-time penalty abatement is a relief provision that had been sitting there for eight years while I paid a monthly installment, and I had no idea it existed.</p><h2 class="heading" style="text-align:left;" id="time-was-only-part-of-the-cost">Time was only part of the cost</h2><p class="paragraph" style="text-align:left;">Before this, unless there were hundreds of dollars at stake, I ate it. Insurance appeals I ate on principle, because everyone knows appeals are a nightmare, and the nightmare is by design. The denial works because most people don&#39;t appeal. <a class="link" href="https://www.kff.org/patient-consumer-protections/claims-denials-and-appeals-in-aca-marketplace-plans-in-2024/?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">KFF&#39;s latest numbers</a> put it at fewer than 1% of denied claims ever getting appealed, and the <a class="link" href="https://litigationtracker.law.georgetown.edu/litigation/estate-of-gene-b-lokken-the-et-al-v-unitedhealth-group-inc-et-al/?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">class action against UnitedHealth</a> over its AI-driven denials alleges the company counted on exactly that, denying claims knowing only a fraction of a percent of policyholders would push back.</p><p class="paragraph" style="text-align:left;">I always told myself the reason was time. Time is money, these things need to happen during business hours or they eat into family time, and so on. That was true, but the bigger cost was cognitive. Every step needs research and thought. What&#39;s the warranty actually say? Who regulates pet insurance in Pennsylvania? Even a phone call is effort long before you dial. And effort is a finite resource. You can only juggle so many open loops at once, and a months-long dispute occupies one of those slots the whole time it drags, wearing you down. Each one a small commitment, stacked into a big one, for an outcome I&#39;d already priced at &quot;probably not worth it.&quot; Its exhausting.</p><p class="paragraph" style="text-align:left;">The agent carries almost all of that now. Every claim starts with me word-vomiting using <a class="link" href="https://github.com/Elevate-Code/better-voice-typing?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">voice-to-text</a> about what happened, and it collects the context, does the research and comes back with a dossier: what the contract says, what their marketing promised, what other owners are reporting on Reddit. Which regulator has jurisdiction. A plan with three branches depending on how they respond. I have a pipeline that accurately OCRs the PDFs. I voice-to-text the support calls and feed the transcript back live to Claude to get told what to say next. Sure, I still review every email before it goes out, still make the strategy calls, still have to physically get on the phone or box up a return and drive it to FedEx. But the part where I have to hold the whole thing in my head is mostly gone. Some weeks my only job on a long dispute is to fire up Claude and say &quot;they responded&quot;.</p><p class="paragraph" style="text-align:left;">Most of what the agent does for me is manage the context.</p><h2 class="heading" style="text-align:left;" id="pulling-levers-until-one-works">Pulling levers until one works</h2><p class="paragraph" style="text-align:left;">The agent&#39;s first-pass advice on these cases is usually weak. &quot;File a complaint with the FTC.&quot; &quot;Contact your state representative.&quot; Sounds right, has probably never worked for anyone except in the very early days of each channel. What actually wins is a multi-front campaign. Pull the agent in early and it plans the whole thing: which levers exist, what order to pull them in, what to do when each one doesn&#39;t move. There&#39;s no way to know up front which one a company will respond to. You find out by trying them all.</p><p class="paragraph" style="text-align:left;">With the robot mop, the opening plan was a Magnuson-Moss Warranty Act angle that I doubt would have gone anywhere. What does a federal warranty statute even look like against a manufacturer on the other side of the world? Meanwhile eleven weeks of &quot;forwarding this to our technical team&quot; ended in a flat refusal to honor their own warranty: a 46% partial refund, prorated for &quot;months of use&quot; that included the month the robot spent at their repair center, and that was final. The most explicit case of a company not holding up its end of a deal I&#39;ve ever had, and a few years ago that would have been that.</p><p class="paragraph" style="text-align:left;">Then the agent, reading through their terms of service, found the needle in the haystack: the company&#39;s own arbitration clause. Under the consumer rules of the forum they themselves picked, my filing cost is capped at $250 and they bear the rest of the fees, which for a claim this size is several times the refund before anyone even hears the merits. Three days after I filed, they offered the exchange they&#39;d previously denied.</p><p class="paragraph" style="text-align:left;">With the pet insurance the lever was the regulator. Weeks of correspondence with the CX team did nothing; their supervisor closed the thread with &quot;we&#39;ve gone as far as we can go.&quot; A complaint to their home-state regulator in New York, plus emails straight to two of their executives, and they folded in three and a half hours. A &quot;special exception&quot;, approved by leadership. That closing line from the supervisor used to be exactly true. Now it just marks where the campaign leaves the support queue.</p><p class="paragraph" style="text-align:left;">And with Plaud it was publicity. Email got nowhere across six reps in six days, each one sending a slightly more coherent version of the same denial. So it went public on every channel more or less at once: Reddit, an Amazon review, a Trustpilot review, a tweet, and a LinkedIn message straight to the CEO. Their Trustpilot team reopened the case the following morning and the CEO replied at midnight. (The reviews got updated to three stars once they made it right.) I couldn&#39;t have run that campaign by hand, or rather, I would never have bothered.</p><h2 class="heading" style="text-align:left;" id="it-became-a-game">It became a game</h2><p class="paragraph" style="text-align:left;">The strangest side effect is that the emotion is gone. When I get a denial now I don&#39;t feel anything, because I already know the three moves that come after it. A setback used to mean another evening of research I didn&#39;t want to do. Now it means the next branch of a plan I&#39;ve already read.</p><p class="paragraph" style="text-align:left;">Companies have always relied on a drop-off funnel, and the funnel is well documented. Customer-service research going back to the 1970s found only about half of unhappy customers ever complain, and somewhere between 1 and 5% escalate past the frontline rep. And at the far end it collapses to effectively nothing: against the <a class="link" href="https://www.supremecourt.gov/DocketPDF/20/20-1143/192737/20210917131917937_41475%20pdf%20Szalai.pdf?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">800-million-plus arbitration agreements</a> Americans are bound by, <a class="link" href="https://files.consumerfinance.gov/f/201503_cfpb_arbitration-study-report-to-congress-2015.pdf?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">roughly a thousand consumer arbitrations</a> were getting filed per year as recently as the early 2010s. Support is designed for that funnel. &quot;Here&#39;s 20% off your next purchase, take it or leave it&quot;.</p><p class="paragraph" style="text-align:left;">An agent breaks the funnel.</p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;">AI agents have infinite patience. They will fight for your claim for as long as you keep paying for your Claude/ChatGPT subscription.</p><figcaption class="blockquote__byline"></figcaption></blockquote></div><p class="paragraph" style="text-align:left;">So when the Subaru started making a weird noise around 15 miles an hour, I didn&#39;t get stressed. I told Claude about it, it figured out it was probably the torque converter and that there was an extended CVT warranty that expired soon, and we made sure the dealer honored it. That&#39;s the mindset I&#39;m in now. Anything that goes wrong, first thought is &quot;let&#39;s make sure they hold up their end&quot;.</p><h2 class="heading" style="text-align:left;" id="the-other-side-has-agents-too">The other side has agents too</h2><p class="paragraph" style="text-align:left;">&quot;We can use this for customer support&quot; was likely the very first business idea for LLMs, and &quot;customer support&quot; is now just a polite name for automating away customer complaints. The pet insurer&#39;s first two replies to me came from a chipper bot with a human first name. The tracking parameter in her email links literally said <code>cxllm</code>. The robo-mop company&#39;s reps sent the same template with the same broken auto-numbering week after week. Every support interaction I have now, I assume there&#39;s an AI model somewhere in the loop, either driving the whole thing or just helping a rep draft.</p><p class="paragraph" style="text-align:left;">And they have every right to do that. That&#39;s the arms race. In practice it means the dark patterns of customer support are getting applied more efficiently and more thoroughly. Stall for exactly the right amount of time, ask for one more small reasonable thing and then another, and the customer frog boils without ever noticing. An agent can run that script with a consistency no call center ever could.</p><p class="paragraph" style="text-align:left;">Which means the vanilla advice from free ChatGPT gets deflected on contact, because the support team already has an SOP for it. The low-level channels will get flooded and filtered, if they ever worked at all. The higher-level ones, the regulators and the arbitration forums, move at government speed and will be a year behind whatever consumers and companies are deploying.</p><p class="paragraph" style="text-align:left;">The part that bothers me is the flood. I take care to make every outbound message and filing precise, partly because precision is what wins and partly because I don&#39;t want to add slop to the pile. But I&#39;m sure most people won&#39;t, and the state insurance departments and the regulators and the arbitration intake queues are about to be buried in it. Arbitration shows where this goes: those thousand filings a year have <a class="link" href="https://www.adr.org/news-and-insights/understanding-the-mass-arbitration-landscape-in-2024-insights-from-the-aaa-s-new-infographic/?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">already become six figures</a>, except it&#39;s law firms running coordinated mass campaigns rather than individual consumers.</p><h2 class="heading" style="text-align:left;" id="am-i-the-karen-now">Am I the Karen now?</h2><p class="paragraph" style="text-align:left;">The system does push me to be a bit Karen-like. When CAVA shorted the kids&#39; meals and the bowl came half full, I hesitated. Is this really something I should be filing a corporate complaint over? What about the $10 Costco soup with the punctured seal, made by a small company?</p><p class="paragraph" style="text-align:left;">I don&#39;t treat small businesses this way. <a class="link" href="https://news.gallup.com/poll/270296/americans-dislike-big-business.aspx?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">Gallup last year</a> had confidence in big business at 15% versus 70% for small business, and that split is about where I am. When the chimney guy came out and did a great job, I used the same agent to post detailed reviews with photos for him across every platform, every word of them true, because it costs me nothing extra and it&#39;s the same machine pointed the other way. And I&#39;m never rude to the customer support rep. The person on the other end is just following a script to get paid, same as I would be. Nothing good comes from being aggressive with them.</p><p class="paragraph" style="text-align:left;">Big corporations, though, I&#39;ve stopped feeling bad about. <a class="link" href="https://www.forrester.com/blogs/cx-index-2025-results/?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">Forrester&#39;s CX index</a> hit an all-time low in 2025 after four straight years of decline. The <a class="link" href="https://www.prnewswire.com/news-releases/new-national-customer-rage-survey-reveals-civility-in-freefall-across-americas-marketplace-302629849.html?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=i-haven-t-lost-a-customer-service-fight-in-seven-months" target="_blank" rel="noopener noreferrer nofollow">National Customer Rage Survey</a> has 77% of Americans reporting a problem with a product or service in the last year, more than double what it was in 1976. My own support experience with Anthropic, the company that makes the agent I do all this with, has been a Google Form and silence. The prevailing business model is to churn through customer goodwill: as long as marketing acquires new customers faster than terrible support loses old ones, support is just a cost center to starve.</p><p class="paragraph" style="text-align:left;">There&#39;s a line going around the internet, no idea who said it first:</p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;">Corporations behave like they&#39;re annoyed that they have to go through you to get to your money</p><figcaption class="blockquote__byline"></figcaption></blockquote></div><p class="paragraph" style="text-align:left;">The mop company wants you to buy the robot, and when it stops mopping they seem irritated that a warranty is standing between them and the sale being over.</p><p class="paragraph" style="text-align:left;">Companies always knew the funnel math. For the first time it&#39;s turned around. But it probably won&#39;t stay that way for long, so I&#39;m taking full advantage while I can.</p></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F67cb5949-a05d-47dc-b2e8-78ad9286e24e%2FDALL_E_2023-12-04_09.13.17_-_A_minimalistic_and_clean_flat_design_style_logo_of_a_stylized_emerald-colored_pyramid._The_pyramid_gradually_transforms_into_a_digital__pixel-like_str.png%3Fv%3D1789183189&publication_name=Dimitri+Sudomoin&utm_campaign=6f28d887-09fd-489c-9491-9b0c0c4c3a17&utm_medium=post_rss&utm_source=dimitri_sudomoin">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>My Most Autonomous Loop Has Been Dead for Two Months</title>
  <description>My fully autonomous pipeline ran for months, quietly died, and nobody noticed - including me. What that taught me about the loop-engineering hype.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/be98a4a4-327f-4034-bfb5-1d6b29335919/loop-engineering-cover.png" length="808250" type="image/png"/>
  <link>https://www.sudomoin.com/p/loop-engineering</link>
  <guid isPermaLink="true">https://www.sudomoin.com/p/loop-engineering</guid>
  <pubDate>Mon, 10 Aug 2026 17:03:52 +0000</pubDate>
  <atom:published>2026-08-10T17:03:52Z</atom:published>
    <dc:creator>Dimitri Sudomoin</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;">Forget context engineering, &quot;loop engineering&quot; is what you should be doing, or wait is it harness engineering now? Anyway...</p><p class="paragraph" style="text-align:left;">Boris Cherny (created Claude Code): &quot;My job is to write loops.&quot;</p><p class="paragraph" style="text-align:left;">Peter Steinberger (the OpenClaw guy): &quot;You shouldn&#39;t be prompting coding agents anymore. You should be designing loops that prompt your agents.&quot;</p><p class="paragraph" style="text-align:left;">Meanwhile on Reddit, in a thread titled &quot;Can anyone explain me loop engineering like I&#39;m 5&quot;, the top comment:</p><div class="blockquote"><blockquote class="blockquote__quote"><p class="paragraph" style="text-align:left;">&quot;Loop engineering doesn&#39;t exist. It&#39;s just scheduled jobs and event triggers that make an agent go brrrr.&quot;</p><figcaption class="blockquote__byline"></figcaption></blockquote></div><p class="paragraph" style="text-align:left;">The skeptics are right that loop engineering is mostly an AI-flavored rebranding of systems engineering. Cron jobs, CI, state machines, retry logic - boooring. (Though if your tokens are free and infinite, throwing them at a problem in a loop probably <i>is</i> cheaper than doing systems engineering.) And to be fair to the evangelists: they live inside their loops all day, which is exactly the part that doesn&#39;t travel when you copy the pattern.</p><p class="paragraph" style="text-align:left;">Where the concept does damage is the picture it sells: build the loop, step away, wake up to finished work. The newer, bolder version of the picture even gives the agent loop a name and a profile photo and lets it handle relationships. People are running real businesses on that right now. Either they can&#39;t tell the output is slop, or the bill just hasn&#39;t arrived yet.</p><p class="paragraph" style="text-align:left;">Mine arrived. Let me tell you about my deadest loop.</p><h2 class="heading" style="text-align:left;" id="the-loop-that-ran-until-it-didnt">The loop that ran until it didn&#39;t</h2><p class="paragraph" style="text-align:left;">A while back we built a project sentiment engine for the agency. An LLM pipeline that ran through every client Slack channel, emails, meeting recordings, etc. and produced a per-project read: is this going well? Is the client getting pissed off? Or are they giving praise and we should ask for a testimonial? The output fed a dashboard we put real effort into. Health indicators, sentiment tooltips, hours tracking, the works.</p><p class="paragraph" style="text-align:left;">By every definition in the loop engineering posts, this thing was a success. Fully autonomous. Stateful. Ran on its own. Produced structured output to a purpose-built surface.</p><p class="paragraph" style="text-align:left;">This month I came back to that dashboard because I&#39;m building out a higher level system for collecting proof (testimonials, reviews, case study material) and sentiment signals are an input to it. That&#39;s when I found out the pipeline had been dead for the better part of two months. Fully dead. Nobody noticed. Including me, and I built it.</p><p class="paragraph" style="text-align:left;">The &quot;loop engineer&quot; would say the loop was never fully closed - it needed better alerting, maybe a sub-agent for health checks. The darker read is that the loop was fully closed, and the removal of the human is what killed it. The value it produced depended entirely on a human being readily available and willing to act on its outputs, and nobody owned that step. So the output became background noise (slop). Meanwhile my account lead was already doing sentiment reads &quot;the old fashioned way&quot; on every client, just by being in the channels and paying attention. The loop never stood a chance.</p><h2 class="heading" style="text-align:left;" id="doing-work-for-the-sake-of-doing-wo">Doing work for the sake of doing work</h2><p class="paragraph" style="text-align:left;">That failure mode is extremely easy to fall into with agents, and I see people doing it constantly. Agents will happily do work for the sake of doing work. They&#39;re sycophantic, the dopamine of watching something build itself is real, and it&#39;s very easy to slide into building loops just because you can. Claude: &quot;I can do [next task] for you - just say the word.&quot;</p><p class="paragraph" style="text-align:left;">It&#39;s the same vibe as the people who previously built elaborate Notion/Obsidian/ClickUp systems to perfectly catalog and organize for some hypothetical future. Productivity theater. Clients do it. I do it. The only difference now is the theater runs itself and bills you for tokens.</p><h2 class="heading" style="text-align:left;" id="what-i-loop-autonomously-and-what-i">What I loop autonomously and what I refuse to</h2><p class="paragraph" style="text-align:left;">I have a blunt pattern - everything I trust to run unattended yields a notification (mostly to me) and nothing more. A hyper-local news digest for me and my wife. A weekly exercise accountability email. A homelab health report. Aggregation summaries with low stakes if ignored, or with consequences that are explicit and accepted, nothing in between. The one scheduled routine that actually acts on its own is a small sub-loop of my paid ads process, and it only earned that trust because it runs sandwiched between sessions where I review its work. Interleaved with me, not independent of me.</p><p class="paragraph" style="text-align:left;">The most autonomous routine I ever built was a general-purpose task executor: calendar-triggered, polling every 30 minutes, retry logic, per-task autonomy levels. It ran three times before I killed it. Removing myself from that process required documenting every edge case, which is impossible, and providing all the necessary context in advance, which is also impossible. It&#39;s the same reason full self-driving cars are still fenced into a few cities after all these years and billions of dollars. Running that executor was like claiming I could sleep in the back seat.</p><p class="paragraph" style="text-align:left;">The decision rule I actually landed on is about leverage. Checking metrics in a UI is a minimal-leverage decision; if the agent is 90-95% aligned with how I&#39;d do it, it&#39;s gone from my plate. Deciding what the ad account should stop spending money on is a medium-leverage decision, delegable, but only after the guidelines are set and reviewed, and setting the guidelines is the key step. The high-leverage decisions never leave. On its face this sounds simple, but it&#39;s basically systems design. Simple but not simple.</p><h2 class="heading" style="text-align:left;" id="the-missing-piece-getting-back-in">The missing piece: getting back in</h2><p class="paragraph" style="text-align:left;">My real objection to running everything as scheduled autonomous loops isn&#39;t reliability - agents are plenty reliable when I&#39;m in the session with them. It&#39;s that an unattended schedule removes you from the loop, and once you&#39;re removed, you&#39;re removed from the verification, the escalation, the design, and the constraints, all the things the loop engineering posts say are your new job.</p><p class="paragraph" style="text-align:left;">Sure, the loop can email you a report. But you&#39;re not in that context anymore. It&#39;s like the difference between assigning a task to someone on your team and meeting with them regularly, versus assigning it and getting a weekly status report. The weekly report becomes background noise fast, to the point where you don&#39;t even know how to get back into the context to change things. The cost isn&#39;t just the missed decision. It&#39;s the re-entry.</p><p class="paragraph" style="text-align:left;">So my version of loop engineering became a system for getting back in cheaply. During working sessions the agent queues future work on my calendar, and every item carries what&#39;s needed to resume the exact session that created it, context intact. The loop runs on my calendar. Its memory is resumable sessions. Its trigger is me.</p><p class="paragraph" style="text-align:left;">I should be honest about its failure mode, because it&#39;s sitting in my queue right now: the overdue items are all physical. Ship a warranty unit back to the manufacturer. Call the car dealership. Drive to FedEx. The agent-side work is done on every one of them; the stalled step is mine. Until we get very effective robots, most high-level loops cant be meaningfully closed, and in my system the human in the loop is both the value source and the single point of failure.</p><h2 class="heading" style="text-align:left;" id="the-bottleneck-is-decision-bandwidt">The bottleneck is decision bandwidth</h2><p class="paragraph" style="text-align:left;">Compared to two years ago, I&#39;m doing the work of about two and a half of me. That&#39;s not hyperbole. I maintain half a dozen internal apps and have absorbed entire functions we&#39;d otherwise be paying for: paid ads, bookkeeping, our own taxes. Last week an agent reconciled hundreds of transactions in our accounting software while I supervised from the side, a chore that used to eat hours. That&#39;s becoming a weekly cadence, but I fire it manually and check the outputs, because I&#39;ve watched what happens otherwise: the process gets devalued, or derailed and forgotten, or both.</p><p class="paragraph" style="text-align:left;">The constraint on all of it is decision bandwidth. Decisions are surprisingly taxing. The newer agents, in their effort to be thorough, hand you pages of analysis and then ask for a decision at the end, and I can only read and absorb so fast. Especially with ADHD, that&#39;s the tax that actually hurts. Which is exactly why handing the decisions to the agent is so tempting, and exactly why I don&#39;t: I need to be informed before I make a decision, and if I don&#39;t make a decision, no value is being created.</p><p class="paragraph" style="text-align:left;">Staying close is also how you catch the brittleness that would otherwise compound silently. In an ongoing warranty fight, I noticed the agent&#39;s drafted replies kept rehashing the entire saga from the beginning, every single time. It wasn&#39;t reading the full email thread before drafting, so it had no idea we&#39;d already made those points. I only caught it because I review every email before it goes out. A fully autonomous version of that loop doesn&#39;t get corrected. It gets stuck in a local minimum and confidently grinds there forever.</p><p class="paragraph" style="text-align:left;">There&#39;s a quieter cost to stepping away too: the learning is the multiplier. When I <a class="link" href="https://www.sudomoin.com/p/ai-can-reverse-engineer-hardware-i-can-t-turn-off-my-own-alarm?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=my-most-autonomous-loop-has-been-dead-for-two-months" target="_blank" rel="noopener noreferrer nofollow">reverse-engineered my voice recorder</a>, the agent did the work, but collaborating tightly on it is why I now know what reverse engineering actually takes. Wave the magic wand from a distance and you get the artifact without the capability.</p><h2 class="heading" style="text-align:left;" id="two-of-my-clients-built-the-same-em">Two of my clients built the same employee</h2><p class="paragraph" style="text-align:left;">The same decay shows up when the loop points at someone else&#39;s inbox instead of your own dashboard, only then the thing decaying is the relationship. This spring, two of my clients independently built the same thing: an AI employee with a name, a personality, and a job. Neither knew about the other. And neither of them has ever said the words &quot;loop engineering&quot; - they&#39;re just building agent systems and letting them run, which is all this trend ever was.</p><p class="paragraph" style="text-align:left;">The first runs his through Discord. She sweeps his Gmail, his meeting transcripts, and his calendar every hour, briefs him morning, midday, and night, and drafts his replies. When an email goes out, it goes out from his own address, signed with her name, &quot;on behalf of&quot; him. I&#39;ve been on the receiving end. The drafts are good, she once read our codebase, correctly spotted a missing function, and asked me to add it. I answer her the way I&#39;d answer him, because he reads everything before it sends and I know the asks are his. What he&#39;s actually built, and I mean this as a compliment, is memory infrastructure. The Discord channel is how context survives between sessions, and replying to her there is how it stays coherent. The loop still ends at him pressing send.</p><p class="paragraph" style="text-align:left;">The second built the outward-facing version. His persona has her own email, her own profiles in his community and his project tools, and by his account she does the work of several people and everyone thinks she&#39;s human. Some of her loops are the same shape as the first guy&#39;s, drafts that end at his keyboard, escalation digests that end in a section with his name on it, and those are the ones producing his wins. The rest broadcast. Daily posts and updates whether or not there&#39;s anything worth saying, sent to people who never asked for them. Somewhere along the line my inbox rules learned to auto-archive her. My agent now filters out his agent. In my inbox, at least, nobody thinks she&#39;s human, because nobody thinks about her at all.</p><p class="paragraph" style="text-align:left;">That end state is what the broadcast loops earn, no matter whose persona is running them. Even people who are fooled for a while eventually notice the pattern - this &quot;person&quot; produces a lot of words that never mean anything, never need anything, never change anything - and they tune her out like any other spam. In a community, that&#39;s corrosive. The mechanism underneath is simple: the moment I know a reply is agent-written, the value of reading it collapses, so the effort I put into answering collapses with it, and the honest equilibrium is my agent replying to his agent. Telephone, with nobody on either end. At which point: what was the relationship for?</p><p class="paragraph" style="text-align:left;">I&#39;m not against agents touching correspondence. My own Claude-driven family emails open with &quot;Hey, Claude here&quot; and sign off as the AI, and it works because everyone involved knows exactly what it is. The &quot;on behalf of&quot; byline works for the same reason. The rule I keep arriving at: an agent inside a relationship is fine when it&#39;s transparent and both sides have agreed what that means. It corrodes when one side is being fooled, or worse, has stopped caring enough to check. And even with full disclosure, I edit every draft before it sends, which in practice means cutting a third of what the agent wrote and adding the one piece of context only I could know. That edit is what makes it mine. The other person can tell, even when they can&#39;t say how.</p><h2 class="heading" style="text-align:left;" id="youll-know-when-it-gets-boring">You&#39;ll know when it gets boring</h2><p class="paragraph" style="text-align:left;">I&#39;ve never run an overnight autonomous loop. Never done the <a class="link" href="https://github.com/ghuntley/how-to-ralph-wiggum?utm_source=www.sudomoin.com&utm_medium=newsletter&utm_campaign=my-most-autonomous-loop-has-been-dead-for-two-months" target="_blank" rel="noopener noreferrer nofollow">Ralph Wiggum thing</a> (an agent in a loop rerunning until the work is done). Never used the loop-forever features in my own tools. Maybe that&#39;s a blind spot. But here&#39;s what I&#39;d tell someone fired up to schedule their first autonomous loop after reading the loop-engineering posts: run it manually until it gets boring. Fire it yourself, stay in it, review the outputs. You&#39;ll know it can be trusted autonomously when it gets very consistent. Jump the gun and it collapses in on itself like my sentiment engine did, or it just burns money into Anthropic&#39;s pocket.</p><p class="paragraph" style="text-align:left;">The deeper reason to stay in the loop is that it&#39;s the only way up. Automate your customer support with a loop and step away, and you get poor results, sure. But you also forfeit the thing the loop would have shown you: that customers keep hitting the same couple problems, and that there&#39;s a higher-order loop worth building to catch them upstream. That&#39;s where the actual business value was, and you gave it up by leaving.</p><p class="paragraph" style="text-align:left;">The sentiment engine is getting rebuilt right now, as a sub-loop of that proof-collection system that explicitly consumes its output. The engineering barely changed. The difference is that this time, its output has a customer.</p><p class="paragraph" style="text-align:left;">The loop was never the point. The decisions are. Everything I&#39;ve automated works precisely because I kept the decisions and gave away everything else, and the moment I stop making them, no value is created at all. It&#39;s just slop, on a schedule.</p></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F67cb5949-a05d-47dc-b2e8-78ad9286e24e%2FDALL_E_2023-12-04_09.13.17_-_A_minimalistic_and_clean_flat_design_style_logo_of_a_stylized_emerald-colored_pyramid._The_pyramid_gradually_transforms_into_a_digital__pixel-like_str.png%3Fv%3D1789183189&publication_name=Dimitri+Sudomoin&utm_campaign=da8d683d-483d-4f5b-8239-c6ad2e222770&utm_medium=post_rss&utm_source=dimitri_sudomoin">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>AI Can Reverse-Engineer Hardware. I Can&#39;t Turn Off My Own Alarm.</title>
  <description>On Claude Code, ESP32s, and building your own shop digital tools</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/75bf5497-fbed-4f18-8ca8-0b727e5fbecf/remix_20260608_102157_flat-design-illustration-wide.jpg" length="2078607" type="image/jpeg"/>
  <link>https://www.sudomoin.com/p/ai-can-reverse-engineer-hardware-i-can-t-turn-off-my-own-alarm</link>
  <guid isPermaLink="true">https://www.sudomoin.com/p/ai-can-reverse-engineer-hardware-i-can-t-turn-off-my-own-alarm</guid>
  <pubDate>Mon, 08 Jun 2026 14:27:50 +0000</pubDate>
  <atom:published>2026-06-08T14:27:50Z</atom:published>
    <dc:creator>Dimitri Sudomoin</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;">My standing desk raised itself while I was writing this article. I&#39;d been sitting for too long.</p><p class="paragraph" style="text-align:left;">The home automation was originally supposed to be a nudge, not a lifestyle. If I sit for an hour, Home Assistant starts with a polite spoken warning, escalates through push notifications and increasingly aggressive LEDs, and at the 80-minute mark, starts the emotional manipulation (&quot;Think of your kids...&quot;). No match for my ADHD. After two weeks I was reflexively smacking my Nest Mini before it finished its first sentence.</p><p class="paragraph" style="text-align:left;">A desk physically rising while you&#39;re typing is harder to dismiss.</p><p class="paragraph" style="text-align:left;">Claude Code wrote most of that first version, which felt unremarkable in the way AI assistance now feels unremarkable: I vaguely describe what I want - it changes the template, reloads the automation, done. But the desk refused to move on command. Hardware with a proprietary protocol isn&#39;t something Claude Code can just config-file its way through.</p><p class="paragraph" style="text-align:left;">The FlexiSpot E7 Pro Plus standing desk, despite being everywhere, does not really want to be automated. With the existing community integrations, Home Assistant can read its height, but most attempts to move it die silently. The controller expects a specific conversation with a connected keypad, and every existing integration was seemingly skipping that handshake. So I had a $7 ESP32 impersonate the keypad well enough to be trusted. About $12 in parts, an afternoon with Claude Code grinding away, and the desk obeyed.</p><p class="paragraph" style="text-align:left;">I&#39;d been using AI to write Home Assistant config files for months. But this was an LLM helping me reverse-engineer a hardware protocol and write working embedded firmware for a microcontroller. The coding agents are bleeding out of the terminal and into the physical world. And I wanted to see how far I could push that.</p><h2 class="heading" style="text-align:left;" id="going-deeper">Going deeper</h2><p class="paragraph" style="text-align:left;">The desk was just the latest thing. Before that, Claude Code had already taken over my Home Assistant setup piece by piece. A presence sensor plus walking pad power monitor that knows whether I&#39;m sitting, standing, or walking. Frigate camera integrations with object detection and person recognition. Leak sensors, security automations, the whole break escalation system.</p><p class="paragraph" style="text-align:left;">At some point I realized I don&#39;t actually know how Home Assistant works. Not really. I&#39;ve never manually written an automation from scratch. I don&#39;t use the dashboard. Claude Code is my interface to the physical systems in my house - that&#39;s just how it is now.</p><p class="paragraph" style="text-align:left;">This became extremely clear one evening when our washing machine had a minor leak. The leak sensor triggered. The alarm went off. And I opened Claude Code to ask it to silence the alarm, because I genuinely did not know how to do it myself.</p><p class="paragraph" style="text-align:left;">You could read that as a trap I walked into, or you could read it as just the next abstraction layer. Both readings are probably true at the same time. I lean toward the &quot;just another tool&quot; framing most days.</p><p class="paragraph" style="text-align:left;">After the desk, I got greedy - Claude can figure out how to control a desk - what else can it figure out? What about my Plaud Note with a broken display? It&#39;s a portable &quot;AI&quot; voice recorder that had no real reason to be &quot;smart&quot;. It&#39;s a microphone that records audio. The smart features are literally just two API calls: transcribe and summarize, both mediocre. I want it to be dumb and have more capable models handle the smart parts.</p><p class="paragraph" style="text-align:left;">So Claude Code reverse engineered it - with surprising ease. Standard reverse engineering approach: analyze the app, map out the protocol, figure out how the device talks to the phone. Took about five days of casual sessions, multiple dead ends, and at one point it found the necessary key sitting in data it already captured days earlier - just didn&#39;t look at it closely enough. Now I can pull my recordings directly over Bluetooth without any app or subscription.</p><p class="paragraph" style="text-align:left;">Three more ESP32 boards are on order. Next target is a robot lawnmower with no API.</p><h2 class="heading" style="text-align:left;" id="shop-tools">Shop tools</h2><p class="paragraph" style="text-align:left;">There&#39;s a woodworking analogy I keep coming back to. Most woodworkers build their own crosscut sleds (a sled that rides the rails of a table saw to cut wood across) instead of buying one. It&#39;s almost a right of passage. Not because it&#39;s cheaper, but because a pre-made sled is necessarily generic. It doesn&#39;t know your saw, your material sizes, or how you actually work. You can&#39;t easily modify something you didn&#39;t build. I&#39;ve always been this person. Long before Claude Code, I built my own voice transcription app because the built-in Windows one wasn&#39;t good enough. I press Caps Lock and talk instead of type. 90% of my computer interaction is voice now. There are probably better apps out there, but mine works exactly how I want it to.</p><p class="paragraph" style="text-align:left;">Claude Code just made it possible to build these things faster. The barrier drops so low that experimentation becomes almost free. I rigged a classic 7-Eleven door chime to play whenever my dogs came back inside through the back door. Kept it for about two days, which was probably two days longer than my wife would have liked. Deleted the automation. Total investment: maybe five minutes.</p><p class="paragraph" style="text-align:left;">I am extremely known for picking up hobbies and abandoning them about 80% of the way to mastery - not giving up exactly, just moving on to the next thing. Minimal sunk cost means walking away is painless. And fast enough iteration means you might actually finish before your brain moves on. That&#39;s what Claude Code changed. It handles the tedious parts, the proper syntax and debugging and configuration research, leaving me to focus on whether the thing even makes sense in my life. And picking a project back up is effortless. I literally just resume the session. All the files, docs and context is right there. I can ask &quot;where were we?&quot; or &quot;what&#39;s next?&quot; and get an immediate answer. Context switching, which is the thing that costs me the most, is basically eliminated.</p><p class="paragraph" style="text-align:left;">The desk, the Plaud, the home automations, the voice app. They&#39;re all shop tools. Built to fit how I actually work. Mine to maintain. Mine to debug. Mine to break.</p><h2 class="heading" style="text-align:left;" id="where-it-breaks">Where it breaks</h2><p class="paragraph" style="text-align:left;">These models are about 90% accurate. Some of the failure modes you learn to recognize. Knowledge cutoffs. Time estimates are way off, every time, completely useless for planning. Whatever you discuss in the first half of a session strongly biases everything after. If you mention you&#39;re tired or it&#39;s getting late, the model will rush through the rest trying to wrap things up. These are patterns you can work around once you know they exist.</p><p class="paragraph" style="text-align:left;">But failures are inconsistent. I&#39;ve watched it look at a screenshot of something clearly wrong and say &quot;looks great, I think we&#39;re done here.&quot; That Westworld line. &quot;Doesn&#39;t look like anything to me.&quot; Any person would see the problem immediately. The model genuinely does not, until you point at it.</p><p class="paragraph" style="text-align:left;">The real danger isn&#39;t the failure modes you learn to spot. It&#39;s the ones you can&#39;t, because you&#39;re out of your depth. If you don&#39;t have enough domain knowledge to recognize when the 90% has drifted into the wrong 10%, you won&#39;t even know to point at the problem. The inaccuracy just compounds quietly, and you end up somewhere very far from reality without ever feeling lost. You say: &quot;Looks great, I think we&#39;re done here&quot; wiring a relay to mains voltage for the first time.</p><h2 class="heading" style="text-align:left;" id="i-paid-for-this">I paid for this</h2><p class="paragraph" style="text-align:left;">It feels like increasingly every company is converging on the same playbook: subsidize the hardware, lock away the data, charge monthly for access, increase prices. What are you going to do, switch? They&#39;re all doing it. Customer support is now a Google form that goes nowhere - looking at you Anthropic. That&#39;s just how things work now, apparently.</p><p class="paragraph" style="text-align:left;">I paid for my standing desk. The hardware is mine. If the manufacturer didn&#39;t build an integration, I&#39;ll add one. I paid for my Plaud recorder. My voice recordings don&#39;t need to route through anyone&#39;s cloud. The device is fully capable of being a dumb recorder, so that&#39;s what it is now. The OnePlus phone I bought for $80 to root and use as a Bluetooth bridge is, by any standard, obsolete. Too slow to use as a phone. But now it has a second life - a bridge for Claude Code to reverse engineer devices with.</p><p class="paragraph" style="text-align:left;">Once you&#39;ve done this just a couple times, you can stop choosing hardware based on what it can do - and instead think about what <i>you</i> can make it do. I&#39;m shopping for a robot lawnmower right now. The best one for my yard has no Home Assistant support whatsoever. I&#39;m buying it anyway and going to bend it to my will. My OXO coffee maker needs two physical button presses to start. Previously, that was the end of the conversation. Now I look at it and think: ESP32, a relay, 30 minutes of work.</p><p class="paragraph" style="text-align:left;">Limitations you previously accepted as permanent stop being permanent. That&#39;s the actual shift.</p><h2 class="heading" style="text-align:left;" id="the-fine-print">The fine print</h2><p class="paragraph" style="text-align:left;">I have some experience with microcontrollers. Arduinos years ago, a couple PCBs, comfortable with electronics. And I&#39;ve fully committed to doing literally everything through Claude Code, which everyone on my team now comes to me for guidance on. But my story is not the proof that anyone can do hardware with AI. Not yet.</p><p class="paragraph" style="text-align:left;">What I am is someone running an experiment to see where this goes, how far coding agents can push into the physical world, and what breaks along the way. Lots of people already use Claude Code to build and improve their Home Assistant setup. That&#39;s the start. What I&#39;m saying is it doesn&#39;t have to stop there.</p><p class="paragraph" style="text-align:left;">The barrier keeps dropping. M5stack makes modular hardware that clicks together like Legos. The extent of &quot;hardware work&quot; for many projects is literally connecting a few things with cables. The limitation is mostly in your head (and your wallet). And maybe this is just the new way we interact with everything. As long as the tool stays available, stays capable, and doesn&#39;t decide one day that you violated its terms of service.</p><p class="paragraph" style="text-align:left;">Claude Code can be the breaker of walled gardens.</p><p class="paragraph" style="text-align:left;">Assuming they don&#39;t try to make Claude itself a walled garden. Which, in almost certainty, they will try.</p></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F67cb5949-a05d-47dc-b2e8-78ad9286e24e%2FDALL_E_2023-12-04_09.13.17_-_A_minimalistic_and_clean_flat_design_style_logo_of_a_stylized_emerald-colored_pyramid._The_pyramid_gradually_transforms_into_a_digital__pixel-like_str.png%3Fv%3D1789183189&publication_name=Dimitri+Sudomoin&utm_campaign=926d92eb-c759-4b11-b592-50ea76bee02e&utm_medium=post_rss&utm_source=dimitri_sudomoin">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Claude: &quot;Yes, I did it, it&#39;s perfect&quot;</title>
  <description>AI builds what you ask for. It can&#39;t tell you what you should have asked for.</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d5d9e279-f826-4e94-8f9f-93456d4548be/remix_20260506_165009_flat-design-illustration-wide.jpg" length="2017360" type="image/jpeg"/>
  <link>https://www.sudomoin.com/p/claude-yes-i-did-it-its-perfect</link>
  <guid isPermaLink="true">https://www.sudomoin.com/p/claude-yes-i-did-it-its-perfect</guid>
  <pubDate>Wed, 06 May 2026 21:06:43 +0000</pubDate>
  <atom:published>2026-05-06T21:06:43Z</atom:published>
    <dc:creator>Dimitri Sudomoin</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #C0C0C0; }
  .bh__table_cell { padding: 5px; background-color: #FFFFFF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Helvetica',Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#F1F1F1; }
  .bh__table_header p { color: #2A2A2A; font-family:'Trebuchet MS','Lucida Grande',Tahoma,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><p class="paragraph" style="text-align:left;">There was a month earlier this year where I caught myself just copy-pasting client bug reports into Claude Code.</p><p class="paragraph" style="text-align:left;">The agent was running with <code>--dangerously-skip-permissions</code> (it doesn&#39;t ask me to confirm anything. It just runs). I&#39;d paste in what the client sent over, walk away, come back ten minutes later, and check if the changes made sense. If it did, I shipped. If it didn&#39;t, I&#39;d paste the error back in and let Claude figure it out.</p><p class="paragraph" style="text-align:left;">What was I actually adding? I wasn&#39;t writing code. I wasn&#39;t making architectural calls. I was a copy-paste bridge between a client&#39;s Slack message and a capable-enough agent doing the work. And I was charging for it.</p><p class="paragraph" style="text-align:left;">Something had shifted. I just hadn&#39;t named it yet.</p><h2 class="heading" style="text-align:left;" id="clients-like-you">Clients like you</h2><p class="paragraph" style="text-align:left;">The copy-paste realization wasn&#39;t the start of the existential crisis. Just one wave of it. The crisis had been simmering for a while, on and off. Some weeks my wife could read it on me across the room. Some weeks I&#39;d be ranting at strangers at kids birthday parties. The question underneath was always the same: what am I actually adding here? And underneath that one, the heavier version: what work will my kids do?</p><p class="paragraph" style="text-align:left;">What boiled it over was the inbound.</p><p class="paragraph" style="text-align:left;">We started getting messages from clients who&#39;d already built the thing.</p><p class="paragraph" style="text-align:left;">Not &quot;I have an idea.&quot; Not &quot;Here is a broken Lovable app.&quot; Actual <b>working</b> applications. With real users. Stripe integrations. Production-<i>ish</i>. Internal tools, customer-facing SaaS, both. Usually running somewhere on Vercel or Replit. The founders hadn&#39;t written any of the code themselves. They&#39;d prompted it into existence with Claude Code, Replit, Lovable, etc. Three years ago, people came to us because they couldn&#39;t build the thing. Now they came because they already had, and they couldn&#39;t tell what they were holding. One was a COO at a large grant-services company. Another was a financial planner building a practice-management platform for other advisory firms. Another was an operations director at a medical practice with dozens of providers.</p><p class="paragraph" style="text-align:left;">On a call with one of them, the rant slipped out: &quot;This is kind of an existential crisis, right? The AI is getting better. What&#39;s the role of developers?&quot;</p><p class="paragraph" style="text-align:left;">Then I heard myself say the quiet part out loud: &quot;We&#39;re encountering a lot of clients like you. People who took the project 90% of the way themselves.&quot;</p><p class="paragraph" style="text-align:left;">Three or four more calls in the same shape over the following weeks really ground it in. This was the new top of our funnel. If our value was &quot;we can build it for you,&quot; our value was about to be zero.</p><h2 class="heading" style="text-align:left;" id="under-the-hood">Under the hood</h2><p class="paragraph" style="text-align:left;">What pulled me out of the spiral was actually looking at the apps.</p><p class="paragraph" style="text-align:left;">They looked great on live demos. Some even had a handful of real beta users. Then we&#39;d ask about the code architecture, or the deployment pipeline, or logging, backups, security. Blank stares. Or we&#39;d go look ourselves and find horrors.</p><p class="paragraph" style="text-align:left;">The clearest framing I&#39;ve heard for this came from my account lead. He used it with a client last week, and I&#39;ve been stealing it ever since:</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Everything that was asked for got built. Every room is legitimately in the ship. But the ship is not a ship. It&#39;s a set of rooms floating next to each other. The moment the ship hits water, everything floods.</p><p class="paragraph" style="text-align:left;">I keep coming back to this metaphor whenever the existential crisis creeps back (usually after a new round of AI hype). Yes, AI will keep getting better at building what you ask for. What doesn&#39;t improve on the AI&#39;s side is knowing what <i>you</i> <i>should</i> have asked for. For a passenger-prompted app to actually hold up, the AI would have to upskill the prompter into a shipbuilder. An explanation can only be simplified so far before it stops being accurate enough to decide from. The ceiling isn&#39;t the AI. It&#39;s the passenger.</p><h2 class="heading" style="text-align:left;" id="the-one-who-stopped-and-the-one-who">The one who stopped, and the one who didn&#39;t</h2><p class="paragraph" style="text-align:left;"><span style="background-color:#FFF3B0;">Both cases below draw from real engagements. Industries and identifying details have been changed to protect client confidentiality. The technical findings are real.</span></p><p class="paragraph" style="text-align:left;">Two cases from this month make the pattern visible.</p><p class="paragraph" style="text-align:left;"><b>Client one.</b> A COO at a medium size grant-services firm. Her CEO asked her to eliminate the team&#39;s busywork. She opened Claude Code and, over sixty days of evenings and weekends, built them an internal client-management and workflow-tracking app. Vercel, Supabase, Microsoft auth, document upload, email integration, the works. She showed it to her coworkers and, by her words, they &quot;were picking their jaws off the floor.&quot;</p><p class="paragraph" style="text-align:left;">Then she stopped. She&#39;d set up a QC step &quot;kind of loose on purpose&quot; because she knew her coworkers wouldn&#39;t use anything rigid. But she&#39;d hit the ceiling of what she could evaluate herself, and she knew it. &quot;I don&#39;t know what I don&#39;t know,&quot; she told me in the first meeting. &quot;So maybe there&#39;s something I just don&#39;t know to be worried about.&quot;</p><p class="paragraph" style="text-align:left;">On a first-pass review her codebase actually looked pretty good. Claude Code&#39;s default output is usually fine at a surface level. The deeper pass was a different story. The build script ran <code>prisma db push --accept-data-loss</code> on every deploy, meaning any schema change could silently drop production data. She&#39;d been developing locally against a Supabase database while production actually ran on a completely different Neon instance. She didn&#39;t realize the two weren&#39;t connected, and Claude Code hadn&#39;t flagged it either. Fourteen API routes were completely unauthenticated, including admin endpoints exposing contact PII.</p><p class="paragraph" style="text-align:left;">If she&#39;d pushed on without the audit, the most likely outcome was some combination of silent data loss on a deploy, a security breach via the open admin routes, or a codebase too entangled to recover.</p><p class="paragraph" style="text-align:left;"><b>Client two</b>, the contrast. A financial planner building a practice-management platform (client onboarding, portfolio tracking, document management, billing, compliance reporting) as a commercial B2B product for other advisory firms. Bootstrapped from personal savings. Built solo on Replit using AI. By the time we got on the discovery call, she&#39;d done several demos, had more firms in the pipeline, and had a target launch date a month out.</p><p class="paragraph" style="text-align:left;">She had also, the week before, signed a contract with a SOC 2 compliance firm. SOC 2 Type 2 is the enterprise-grade security attestation. It requires demonstrating six to twelve months of operational controls.</p><p class="paragraph" style="text-align:left;">My team found over a hundred issues in the first six hours of audit. The backend was a single seventeen-thousand-line file. There was no tenant isolation, which for a multi-tenant financial app means any advisor at any firm could hypothetically pull another firm&#39;s client records and portfolio data by URL. Payment tokens were generated with <code>Math.random()</code>. A production API key for a third-party document service was hardcoded into frontend code shipped to every user&#39;s browser. Client custodial account balances - the kind that trigger regulatory action when they&#39;re wrong - were read and written without a database lock. Her actual state was a product with zero working access controls.</p><p class="paragraph" style="text-align:left;">On the call, she said something I suspect every vibe coder comes to eventually. The later they get there, the more expensive it is:</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">She isn&#39;t reckless. Reckless would require her to have known what she was missing. Shes a real financial professional with real domain expertise in an industry that genuinely needs better tooling. She saw a gap, used the tools available, built something impressive. The tools didn&#39;t tell her when to stop. The AI told her, every feature along the way, &quot;yes, I did it, it&#39;s perfect.&quot; And she believed it, because she had no other reference point.</p><p class="paragraph" style="text-align:left;">Two founders. Same tools. Different outcomes. The difference wasn&#39;t capability. It was knowing where the boundary was. The first one knew because she&#39;d spent a career running a 22-person company and had developed the instinct that when something is load-bearing, you call someone who knows what load-bearing means.</p><h2 class="heading" style="text-align:left;" id="what-this-means-for-the-agency">What this means for the agency</h2><p class="paragraph" style="text-align:left;">I&#39;m rewriting our landing page this month. The pitch used to be about building software. Now it&#39;s about owning whether the thing actually works in the real business it&#39;s deployed into. Whether it holds under load. Whether it survives a regulator. Whether it stays in sync with a workflow that changes every quarter.</p><p class="paragraph" style="text-align:left;">I&#39;m still figuring this part out. I don&#39;t have a neat playbook for what a sustainable AI consultancy looks like in 2028 - nobody does. But I know what we&#39;re investing in: getting deep enough into specific industries (legal, healthcare, professional services) that we accumulate the business context an LLM can&#39;t.</p><p class="paragraph" style="text-align:left;">And I know what we&#39;re not doing anymore. The month I described at the top, where I was a copy-paste bridge and charging for it: that was brief, maybe six weeks. Most agencies are still inside that window and not saying so out loud. The honest version is that a lot of &quot;developer hours&quot; are AI hours with a human copy-pasting, and the invoice doesn&#39;t distinguish. Going forward, clients pay for human time and judgment. AI compute is overhead. If a problem can be solved by handing it to Claude unsupervised, we shouldnt be charging for it, and we&#39;re going to stop pretending otherwise.</p><h2 class="heading" style="text-align:left;" id="the-bet">The bet</h2><p class="paragraph" style="text-align:left;">So here&#39;s the bet I&#39;m making with the agency.</p><p class="paragraph" style="text-align:left;">Code will be effectively free. But the ceiling I described will hold for a long time. The limitation isn&#39;t the AI, it&#39;s the passenger. The work that matters stays stubbornly human: understanding a business deeply enough over time to be trusted with its outcomes, owning the consequences when something goes wrong, knowing where the load-bearing walls are. Not because humans are magical, but because knowing what to ask for is still the whole game, and AI hasn&#39;t changed that yet.</p><p class="paragraph" style="text-align:left;">That&#39;s a narrower strip of ground than agencies used to stand on. &quot;We write good code&quot; is gone. What&#39;s left is knowing which questions to ask before the first line gets written.</p><p class="paragraph" style="text-align:left;">The COO at the grant-services firm understood this, I think, in the second meeting. She said, &quot;I&#39;m willing to be your guinea pig, because I know that this is new.&quot; She was telling me she knew we were both figuring out the rules in real time, and she was okay with that, because the alternative, her and Claude alone, had a ceiling that scared her.</p><p class="paragraph" style="text-align:left;">The financial planner building the advisory platform would probably have learned it, eventually. Or she&#39;d keep racing the avalanche and one of her queued firms would discover the cross-firm data leak, and she&#39;d learn it the expensive way.</p><p class="paragraph" style="text-align:left;">AI writes the code. We own the outcome. I don&#39;t know if that&#39;s enough. But it&#39;s a better bet than billing for Claude time and hoping nobody notices.</p></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/powered-by?publication_logo=https%3A%2F%2Fmedia.beehiiv.com%2Fcdn-cgi%2Fimage%2Ffit%3Dscale-down%2Cformat%3Dauto%2Conerror%3Dredirect%2Cquality%3D80%2Fuploads%2Fpublication%2Flogo%2F67cb5949-a05d-47dc-b2e8-78ad9286e24e%2FDALL_E_2023-12-04_09.13.17_-_A_minimalistic_and_clean_flat_design_style_logo_of_a_stylized_emerald-colored_pyramid._The_pyramid_gradually_transforms_into_a_digital__pixel-like_str.png%3Fv%3D1789183189&publication_name=Dimitri+Sudomoin&utm_campaign=fe59aaa7-ee8a-408b-b394-554e72008d1c&utm_medium=post_rss&utm_source=dimitri_sudomoin">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

  </channel>
</rss>
