<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>TheWhiteBox by Nacho de Gregorio</title>
    <description>The newsletter to stay ahead of the curve in AI</description>
    
    <link>https://thewhitebox.beehiiv.com/</link>
    <atom:link href="https://rss.beehiiv.com/feeds/lBAH7wWFox.xml" rel="self"/>
    
    <lastBuildDate>Mon, 20 Jul 2026 03:38:15 +0000</lastBuildDate>
    <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
    <atom:published>2026-07-19T00:00:00Z</atom:published>
    <atom:updated>2026-07-20T03:38:15Z</atom:updated>
    
      <category>Machine Learning</category>
      <category>Artificial Intelligence</category>
      <category>Technology</category>
    <copyright>Copyright 2026, TheWhiteBox by Nacho de Gregorio</copyright>
    
    <image>
      <url>https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/publication/logo/9b355c4a-4dbc-4f53-b676-a166fee6812c/Isotipo_negro.png</url>
      <title>TheWhiteBox by Nacho de Gregorio</title>
      <link>https://thewhitebox.beehiiv.com/</link>
    </image>
    
    <docs>https://www.rssboard.org/rss-specification</docs>
    <generator>beehiiv</generator>
    <language>en-us</language>
    <webMaster>support@beehiiv.com (Beehiiv Support)</webMaster>

      <item>
  <title>Kimi K3 🚀, President Xi, a mosquito-killing drone walk into a bar</title>
  <description>.
</description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6ca0993f-fba5-4523-890c-9b99cf2b3f4f/ChatGPT_Image_Jul_18__2026__05_12_51_PM.png" length="1944491" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar</guid>
  <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
  <atom:published>2026-07-19T00:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #FBFBFB; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;">Welcome back! This week, we talk about <b>Kimi K3</b>, the first Chinese model that might have closed the gap with the US, great models running on smartphones, market data, <b>OpenAI’s first physical device</b>, a <b>mosquito-killing AI drone</b>, and many more.</p><p class="paragraph" style="text-align:left;"><b>Enjoy!</b></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Kimi K3 is Here. And… wow</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/37ae066c-cc55-4433-8127-c39879b563e9/image.png?t=1784278214"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.kimi.com/blog/kimi-k3?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">Kimi K3 was released yesterday</a>, and some believe <b>it has finally closed the gap with the US frontier</b>. While GLM-5.2 narrowed the gap with Opus 4.8 and GPT-5.5, this model goes head-to-head with Fable and GPT-5.6 Sol.</p><p class="paragraph" style="text-align:left;">It<span style="background-color:transparent;"> has 2.8 trillion parameters (50 billion activated, an extreme 1.7% sparsity, making it fast), establishing it as the largest Chinese model released to date by a wide margin.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">It is larger than Opus 4.8 or Grok 4.5 </span>(both at 1.5 trillion), <b>and yet the economics are quite embarrassing for OpenAI/Anthropic</b>.</p><p class="paragraph" style="text-align:left;">Kimi K3 is priced at $3 per million input tokens and $15 per million output tokens, <b>roughly Sonnet/GPT Terra-level pricing</b>, while claiming performance comparable to the big ones, Fable and Sol, which are priced at $10/$50 and $10/$45, respectively.</p><p class="paragraph" style="text-align:left;">In the BrowseComp benchmark, which measures how good agents are at finding hard-to-find information, not only is the model best-in-class, it’s outrageously cheaper, at least relative to Fable.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I expect a lot of industry talk over the coming days because, at least on a per-token basis, <b>OpenAI and Anthropic now look horribly mispriced.</b></p><p class="paragraph" style="text-align:left;">However, we must also acknowledge the other variable determining cost, token count, which shows OpenAI can be very competitive in terms of pricing. Nonetheless, as shown by Artificial Analysis, Kimi K3 is only slightly cheaper than GPT-5.6 Sol.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7b6704e4-6646-4d10-bfa7-877d59ba6099/image.png?t=1784278848"/></div><p class="paragraph" style="text-align:left;">But there’s no way I can save Anthropic from the burn; <b>they are really completely out of band relative to the rest</b>. TogetherAI, a US inference company, ran several tests <a class="link" href="https://x.com/ZainHasan6/status/2078289171807150312?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">showing that Kimi K3 offers the same performance at 65% lower overall cost than Fable</a>.</p><p class="paragraph" style="text-align:left;">Seeing this, it’s not surprising that Anthropic is being forced to increase rate limits and usage of Fable despite previous claims, <b>which tells you all you need to know about whether China is doing all of us a favor or not by open-sourcing its models</b>.</p><p class="paragraph" style="text-align:left;">In their defense, the reason Anthropic’s prices are so high isn&#39;t just margin-hoarding; it’s that Kimi and OpenAI’s models are much less “dense” mixture-of-experts (MoE) models.</p><p class="paragraph" style="text-align:left;">MoEs (basically all models today) only activate a portion of the model for any given prediction. This lets you combine the advantages of having a larger model (they are smarter) with the latency of a smaller one.</p><p class="paragraph" style="text-align:left;">However, you do pay a price. The fact that Kimi K3 is so incredibly sparse (only 1.7% of the model activates for any single prediction) is necessary to make models run faster, as Chinese chips are worse,<b> but the trade-off is that you lose per-token compute</b>.</p><p class="paragraph" style="text-align:left;"><b>Anthropic models are known to be much denser</b>, which may explain why some people still consider Fable the best model in the world in terms of raw performance (at the expense of cost).</p><p class="paragraph" style="text-align:left;">Which is to say, <b>Anthropic’s Fable might still be, overall speaking, the best model in the world</b>, but with a nominal superiority that in no way justifies the price gap. It’s the best model, but one that no longer makes sense to choose except for intelligence-maximizing use cases, which, as I insist time and time again, are a tiny percentage of global use.</p><p class="paragraph" style="text-align:left;">Needless to say, <b>one has to wonder how long Anthropic and OpenAI will sustain those prices</b>, and I&#39;m less confident they&#39;ll IPO this year because we know for a fact they did not expect “Fable-level” Chinese AIs until the end of the year.</p><p class="paragraph" style="text-align:left;">Hence, Anthropic prices have to drop; I just don’t see how they don’t unless regulation eliminates their competition.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">I will say the picture is blurrier because, at an overall cost level, they seem to be well-priced simply because OpenAI models need way fewer tokens per response than Kimi K3; having better intelligence-per-token is just as effective in reducing overall spending as cutting per-token prices.</p><p class="paragraph" style="text-align:left;">However, <b>perception matters just as much as reality</b>, and commanding higher per-token prices gives customers the (somewhat wrong) impression that your model is more expensive than it really is, which might force you to drop prices either way.</p><p class="paragraph" style="text-align:left;">Another important aspect of all this I’m hearing a lot of nonsense about is hardware. <b><i>Can China serve this model?</i></b></p><p class="paragraph" style="text-align:left;">It&#39;s a good question many are asking, <b>but one that is being asked in the wrong way</b>.</p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">While you can cast reasonable doubt on their capacity (whether they have enough servers to scale), they definitely have the server size to deploy this model. Chinese servers are huge, </span><b>with enormous scale-up domains and therefore perfectly capable, in terms of memory and memory bandwidth, of serving such a model</b><span style="background-color:transparent;">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">In fact, just a few hours ago, Huawei presented its corridor-scale pod with 1 ExaFLOP of FP8 compute and 2 EFLOPs of FP4 compute. In layman’s terms, this gigantic server could theoretically output 2 quintillion operations per second.</span></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/caa4d83c-6b28-4509-b0db-6300acc8e904/image.png?t=1784279171"/></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Such a cluster, with 256 Terabytes of memory, is more than enough to serve a model as large as Kimi K3 with considerably large batches.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>SMALL MODELS</b></span><br>Great Models on Smartphone Hardware?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4848b06c-08d9-425e-b515-73b0204bc81a/image.png?t=1784290052"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://prismml.com/news/bonsai-27b?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://prismml.com/news/bonsai-27b?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">reported by PrismML</a>, the company has released <b>Bonsai 27B</b>, a compressed multimodal model based on Qwen3.6 27B, <b>which</b><b> it says is the first model in its capability class to run on a phone</b>.</p><p class="paragraph" style="text-align:left;">The secret is that, despite having a number of parameters similar to those of other models that don’t run on smartphones, the model has been compressed to ternary and 1-bit precision, <b>dramatically reducing its size to the point that it can be served on an iPhone.</b></p><p class="paragraph" style="text-align:left;">In layman’s terms, while most models have FP8/FP4 precision, meaning each parameter can weigh either 4 bits or 8 bits (one byte), allowing for larger granularity, these models can take values of either 1, -1, or 0 for the ternary weights, (1.58 bits per weight) or 1, -1 for 1-bit (in reality, it’s either the plus sign or the negative sign, meaning each weight either switches the sign of the compuation or it doesn’t).</p><p class="paragraph" style="text-align:left;"><i>And what’s the impact? </i></p><p class="paragraph" style="text-align:left;">For a 10-billion-parameter model, an FP8 model means each parameter weight is 1 byte, so the total size is 10 gigabytes. But for the same model in 1-bit form, the size is 8 times smaller, at 1.25 gigabytes.</p><p class="paragraph" style="text-align:left;">The implication is that, unlike the former, the latter can be run on smartphones, which are usually much more memory constrained (around 6-16 gigabytes of RAM), preventing them from accessing model sizes larger than 4 gigabytes or less, and only for the most ‘beefy’ smartphones with up to 16 GB.</p><p class="paragraph" style="text-align:left;">This could very well be revolutionary.</p><p class="paragraph" style="text-align:left;">Bonsai 27B comes in two versions: a <b>5.9 GB ternary model</b> designed for laptops and a <b>3.9 GB 1-bit model</b> that fits within the available memory of an iPhone 17 Pro. Both support reasoning, vision, structured tool calls, agentic workflows, and a 262,000-token context window.</p><p class="paragraph" style="text-align:left;">Despite the aggressive compression, according to PrismML’s 15-benchmark evaluation, <b>the ternary version retained about 95% of the full-precision model’s overall performance</b>, while the 1-bit version retained about <b>90%</b>.</p><p class="paragraph" style="text-align:left;">The models run on Apple devices through MLX and on NVIDIA GPUs through CUDA. PrismML has released the weights under the <b>Apache 2.0 license</b> and is also offering a limited developer-preview API.</p><p class="paragraph" style="text-align:left;">You can try the models for free <a class="link" href="https://huggingface.co/spaces/webml-community/bonsai-webgpu-kernels?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">using this link</a> (the AI runs in your browser, fully locally). In my personal computer, the model runs at 60 tokens/second, way above the threshold of utility (feels very fast).</p><p class="paragraph" style="text-align:left;">At this rate, I predict we could have Fable-level AIs on smartphones by the end of next year. Or earlier.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Small models are an extremely undervalued category in AI, despite, ironically, <b>my belief that they will represent the vast majority of token generation in the future</b>, leaving data center workloads for the more challenging work.</p><p class="paragraph" style="text-align:left;">Every day that passes, <b>Apple’s decision not to commit to AI in the same way the others did turns closer and closer to being a good choice for them</b>.</p><p class="paragraph" style="text-align:left;">With the ‘AI God’ idea out of vogue, Apple can be comfortable with simply being a distribution play; not creating its own models from scratch (they continuously rely on partners like Google in the US or, just announced, <a class="link" href="https://techcrunch.com/2026/07/16/apple-intelligence-approved-for-launch-in-china-with-alibabas-qwen-ai/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">Alibaba in China</a>), while having direct distribution access to the one billion richest people on Earth.</p><p class="paragraph" style="text-align:left;">The more the industry evolves, the more likely it seems that value once again accrues to the platform, not the models. To the infrastructure and hardware companies, <b>not the increasingly commoditized AI Labs</b>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>[schema] Scores 99% on ARC-AGI 3</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/72df0322-0759-4888-97e0-462d9c82d370/image.png?t=1784383515"/></div><p class="paragraph" style="text-align:left;">ARC-AGI is perhaps the hardest AI benchmark in the world. For reference, <b>the highest frontier score is GPT-5.6 Sol with a 7% score and a benchmark cost of $21k</b>. Not only are models really expensive to run on these tasks, but the scores are terrible.</p><p class="paragraph" style="text-align:left;">And now, <a class="link" href="https://schema-harness.github.io/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">a company claims to have pushed GPT-5.6 and Fable to 95% and 99% scores</a>, respectively, with 0 fine-tuning, just running them on an improved harness.</p><p class="paragraph" style="text-align:left;"><i>But how is that possible?</i></p><p class="paragraph" style="text-align:left;">Recall that the AI products you use these days are not just an AI; they are systems with AI models at the core but a wide range of additional components around them, components that, <a class="link" href="https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-part-ii?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">as we saw a few days ago</a>, not only enhance an AI’s performance (mostly by improving the context they are given) but also allow them to execute actions.</p><p class="paragraph" style="text-align:left;">One of the most powerful components is the ability to write and run code, which enables the AI to execute actions, perform math operations deterministically, and more.</p><p class="paragraph" style="text-align:left;">But what’s new about this solution called [schema] is that <b>it uses code to maintain state</b>. <i>But what does that mean?</i></p><p class="paragraph" style="text-align:left;">All three ARC-AGI benchmarks share one design principle:<b> they are designed to test models in situations they couldn’t have memorized beforehand</b>. This is done to clearly distinguish responses that have been simply memorized from those that have been reasoned.</p><p class="paragraph" style="text-align:left;">Think of this as taking a kid’s cheat sheet from them before they take a maths exam. If they can simply copy the answers from the cheat sheet, we can’t tell whether the student really understands the problem; we aren’t testing reasoning anymore; we are testing copying capabilities.</p><p class="paragraph" style="text-align:left;">ARC-AGI problems are therefore designed to be presented as unique to the AIs. This has not prevented AIs from beating ARC-AGI benchmarks one and two, but the third adds an extra layer of complexity:<b> it lacks clear instructions or objectives.</b></p><p class="paragraph" style="text-align:left;">As explained on the website:</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Tests look like the one below, <a class="link" href="https://arcprize.org/tasks/ls20?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">which you can play too, by the way</a>, like a game, <b>but without clear instructions or goals</b>. Thus, the AI has to test the environment, familiarize itself with it on the fly, and solve an unstated problem by figuring out first what the problem and the constraints are.</p><p class="paragraph" style="text-align:left;">For models trained on clear instructions and objectives, this is most likely very similar to hell, and unsurprisingly, it’s hell for them.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b1376618-bbc5-4222-ab21-92355ccc3f23/image.png?t=1784382690"/></div><p class="paragraph" style="text-align:left;">Therefore, <i>how does [schema] take a model struggling with these problems and turn it into one that can solve these games?</i></p><p class="paragraph" style="text-align:left;">Ironically, using the same pattern we’ve been using for agents for years now, ReAct, but with a twist.</p><p class="paragraph" style="text-align:left;">The ReAct pattern is one where AIs solve problems by reasoning, acting, observing, and repeating. The AI reasons what it thinks it should do, acts based on that, observes the result, and repeats, using the feedback to improve its reasoning-action chain.</p><p class="paragraph" style="text-align:left;">It draws parallels to Bayesian inference, widely believed to be the learning engine our brains use, in which the actor holds certain beliefs, acts based on those beliefs, receives feedback, and updates its priors for the next run. You jump from up high, you hurt yourself, and you learn not to jump from high places anymore.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">However, this is not new, <i>so what changed?</i></p><p class="paragraph" style="text-align:left;">The key difference is between remembering everything that happened and maintaining a clear model of what is happening.</p><p class="paragraph" style="text-align:left;">A standard ReAct agent (how standard models are run on these benchmarks) stores its understanding across the conversation history. Its beliefs are spread through many observations, actions, and pieces of reasoning. Each time it needs to act, it must effectively reread that history and reconstruct what it currently believes.</p><p class="paragraph" style="text-align:left;">A [schema] agent instead turns that scattered understanding into a single, explicit object: <b>a piece of code describing how it thinks the environment works</b>. When new evidence arrives, it updates the code. It can then run that code to predict what will happen next.</p><p class="paragraph" style="text-align:left;">The advantage is similar to the difference between keeping every receipt from a business and maintaining up-to-date accounts. The receipts contain all the information, <b>but the accounts turn that information into a usable state for the business. </b></p><p class="paragraph" style="text-align:left;">The key is that forcing the model to explain ‘what is going on’ and ‘what I’m seeing’ into something tangible, in itself, improves performance.</p><p class="paragraph" style="text-align:left;">It’s like thinking you understood something but realizing you quite didn’t once you put it into words; <b>the act of forcing yourself to describe something, in itself, improves understanding</b>.</p><p class="paragraph" style="text-align:left;">But it’s nonetheless surprising to see that this simple idea translates into models that range from single-digit performance to nearly 100%.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is the latest proof that harnesses have become increasingly important and are a necessary variable for improving performance.</p><p class="paragraph" style="text-align:left;">In a way, it seems we have built something (AIs) that we don’t yet know how to squeeze performance from. This is interesting and also very dangerous for AI Labs, <b>which could soon find themselves building wrenches and startups upstream of them, building the actual plumber that gets sold to customers</b>.</p><p class="paragraph" style="text-align:left;">Once again, we confirm one of the biggest questions in AI: <b>we only know for sure that we don’t know where the most value will accrue</b>.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MEMORY</b></span><br>What is Going on, Google?</h2><p class="paragraph" style="text-align:left;">Alphabet shares fell 4.44% on July 16 <a class="link" href="https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">after reports that Google had delayed Gemini 3.5 Pro</a>, its next flagship AI model.</p><p class="paragraph" style="text-align:left;">Google had previously said the model would arrive in June, <b>but it reportedly failed to meet internal performance targets</b>, particularly for coding tasks.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Google has a problem. And in the most Google way possible, <b>it’s a them problem</b>. They are their own enemy.</p><p class="paragraph" style="text-align:left;">They have it all, and I mean all. More data than anyone. Most compute than anyone (and it’s not particularly close). Verticalized stack, as they own both the hardware and the software, which means they can co-design their AIs to get the most out of their hardware and vice versa. They have more cash than any other model developer.</p><p class="paragraph" style="text-align:left;">And they have a top 3 frontier Lab. Well, allegedly.</p><p class="paragraph" style="text-align:left;">It’s crazy to say this, but Google has a DeepMind problem. Nobody does pretraining better than they do (the phase where AIs capture knowledge), which requires careful engineering, scale across several data centers… and engineering marvel.</p><p class="paragraph" style="text-align:left;"><b>But they are falling </b><b>far short in the area that matters most today: Reinforcement Learning</b>, the training phase that determines who’s in the lead and by how much.</p><p class="paragraph" style="text-align:left;">The best way I can think of to describe what this means is the following:</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Failing where it matters the most has put them at an <b>uncomfortable sixth or even seventh place in the race</b>, after Anthropic, OpenAI, Moonshot, Zhipu, SpaceXAI, and Meta (and you could even perhaps squeeze Minimax there too).</p><p class="paragraph" style="text-align:left;"><b>This is unforgivable, knowing they have more data, compute, and cash (the key variable trifecta) than literally all of them</b>; I would argue they probably have more compute and data than all these companies combined. Unforgivable.</p><p class="paragraph" style="text-align:left;">I’m a Google shareholder, and I will remain one, but God was I right about Google two years ago when I wrote that Google was its own biggest enemy.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>GEOPOLITICS</b></span><br>President Xi’s AI Speech</h2><p class="paragraph" style="text-align:left;">President Xi Jinping <a class="link" href="https://www.reuters.com/world/asia-pacific/chinas-xi-promotes-chinas-commitment-ai-access-speech-shanghai-conference-2026-07-17/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">addressed a crowd in Shanghai </a>regarding China’s AI goals. Key takeaways were:</p><ul><li><p class="paragraph" style="text-align:left;">Started the speech by referring to his signature maxim, <i>&quot;great changes unseen in a century are unfolding across the world&quot;</i></p></li><li><p class="paragraph" style="text-align:left;">Said that the world has &quot;entered an unprecedented period of active innovation on AI technology&quot;, which means &quot;great opportunities as well as challenges for governance”</p></li><li><p class="paragraph" style="text-align:left;"><b>Reaffirmed commitment to open source to promote AI &quot;openness and win-win&quot;</b></p></li><li><p class="paragraph" style="text-align:left;">Warned against &quot;over stretching&quot; the concept of national security as applied to AI, where one country&#39;s national security is prioritized over others</p></li><li><p class="paragraph" style="text-align:left;">Mentioned China opposes the emergence of “new historical injustices”  in AI (one of the most strongly worded parts of the speech)</p></li><li><p class="paragraph" style="text-align:left;">In the next 5 years, China will provide 5000 opportunities to developing countries in &quot;AI training and seminar programs&quot; and &quot;cooperation centers&quot;. He named ASEAN, League of Arab States, African Union, CELAC, SCO, and BRICS</p></li></ul><p class="paragraph" style="text-align:left;"><i>But what should we take away?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">There were fears that China was considering a shift from its open-source stance</a>. But for now, <b>it seems they remain committed to open models</b>, and Kimi K3’s open-weight release confirms that.</p><p class="paragraph" style="text-align:left;">However, nothing prevents them from changing their opinions tomorrow, so who knows.</p><p class="paragraph" style="text-align:left;">To me, their open-source stance has little to do with “win-wins” and “openness,” as he claimed, <b>but rather is a geopolitical weapon; nobody hurts Anthropic and OpenAI’s IPO prospects more than Chinese Labs and their open models.</b></p><p class="paragraph" style="text-align:left;"><b>Hurting Anthropic and OpenAI hurts the US as a whole</b>, given that its largest companies have significant exposure to these Labs. China is not dumb and knows this, and they are actively pursuing it.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>COMMODITIES TRADING</b></span><br>AI Compute, the next commodity?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.bloomberg.com/news/articles/2026-07-14/kalshi-ramps-up-effort-to-build-markets-for-ai-computing-power?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc4NDA4ODMyOCwiZXhwIjoxNzg0NjkzMTI4LCJhcnRpY2xlSWQiOiJUSFo5WjFUOTZPU0cwMCIsImJjb25uZWN0SWQiOiIwOThFNzNDQTE5QTA0RDkxODEyQzQ4MjcwRDZERTI0QiJ9.PAzaf2TkKPqXXceW122TAQk0I0l-S6RQJaDtgkyQurg&utm_source=tldrai&leadSource=article-gifting" target="_blank" rel="noopener noreferrer nofollow">As published by Bloomberg</a>, prediction-market operator <b>Kalshi</b> has launched a forward curve tracking the expected future cost of renting AI computing power. The tool combines weekly and monthly event contracts to estimate GPU rental prices for periods extending up to one year.</p><p class="paragraph" style="text-align:left;">Kalshi says the curve could provide a pricing reference for future derivatives, including futures, options and swaps, allowing AI companies and computing providers to hedge changes in infrastructure costs. Chief Risk Officer <b>Udesh Jha</b> said the company is using prediction markets to show expected prices across different GPU grades and time periods.</p><p class="paragraph" style="text-align:left;">The initiative comes as <b>computing capacity increasingly resembles a tradable commodity</b>, although GPU products differ by model, location, and performance, and become outdated quickly.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">In the same way we trade oil these days, AI compute will become an essential commodity whose price will impact the global economy and directly impact inflation.</p><p class="paragraph" style="text-align:left;"><b>I do envision a future where AI is present in every digital product in some way</b>, and these products will account for a large share of the economy. This means that each single one of these products’ prices will be largely determined by the price of AI compute.</p><p class="paragraph" style="text-align:left;">This is why I always insist that AI compute will be an essential component for the US to maintain the dollar as the world’s reserve currency and, in consequence, <b>I expect the US to follow a similar strategy they followed with oil</b>, forcing Gulf countries to denominate oil in dollars, thereby securing the global importance of the dollar (giving way to the petrodollar). If the world needed oil, and oil was denominated in dollars, the world needed dollars.</p><p class="paragraph" style="text-align:left;">Now, I expect AI compute to follow the exact same path, and initiatives like<a class="link" href="https://www.state.gov/pax-silica?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow"> the Pax Silica</a>, which the EU has already signed in paper, <b>aim to guarantee US global compute dominance.</b></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Thinking Machines’ First Great Models</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/63ef988d-fbdf-41ec-9de0-6f02e9c527a1/image.png?t=1784282625"/></div><p class="paragraph" style="text-align:left;">Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, <a class="link" href="https://thinkingmachines.ai/news/introducing-inkling/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">has introduced Inkling</a>, a new open-weight model that has rapidly emerged as the leading US-based entrant in the sector.</p><p class="paragraph" style="text-align:left;">Independent trackers, including Artificial Analysis, <b>currently rank Inkling ahead of other major US open-weight models such as Nvidia’s Nemotron and OpenAI’s gpt-oss</b>.</p><p class="paragraph" style="text-align:left;">Inkling distinguishes itself through several high-performance features designed for efficiency and versatility:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Architecture:</b> It employs a mixture-of-experts (MoE)—how surprising—design with 975B total parameters, utilizing only ~41B active parameters per query for cost-effective performance.</p></li><li><p class="paragraph" style="text-align:left;"><b>Multimodality:</b> Unlike many open peers, Inkling is natively multimodal, capable of reasoning across text, image, and audio inputs.</p></li><li><p class="paragraph" style="text-align:left;"><b>Context and Control:</b> The model supports a 1M-token context window and features a &quot;controllable thinking effort&quot; dial, allowing users to balance reasoning depth against speed and cost.</p></li><li><p class="paragraph" style="text-align:left;"><b>Performance:</b> Reports suggest the model achieves results comparable to rivals while using roughly one-third to one-half as many tokens, with improved calibration and lower hallucination rates.</p></li><li><p class="paragraph" style="text-align:left;"><b>Decision Support:</b> Notably, the model demonstrates unusual strength in forecasting and calibration (meaning its confidence levels more accurately reflect the probability of its answers being correct), making it a strong candidate for decision-support applications.</p></li></ul><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">It is a great model and an incredible contribution to national security because the US needs strong, open models. There are other important insights to acknowledge:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Thinking Machines also offers the Tinker API</b>, the fine-tuning platform that Bridgewater Associates used to train its “news-filtering” model, which reportedly outperformed Opus 4.8 and GPT-5.5. Until now, Tinker supported third-party open-weight models. Inkling will now become available for fine-tuning as well.</p></li><li><p class="paragraph" style="text-align:left;">The Lab’s entire strategy can be summarized as a single bet: <b>enterprises will increasingly train and customize open models rather than rely exclusively on frontier model</b><b> providers</b>. Thinking Machines therefore wants to provide both the models and the infrastructure required to train them. This lab, filled to the brim with prominent former OpenAI and Anthropic researchers, <b>is making a direct bet against those companies’ current business models.</b></p></li><li><p class="paragraph" style="text-align:left;"><b>Inkling may be one of the few strong open models that was not heavily distilled from existing frontier systems</b>. The company claims that it used no frontier-model distillation. If true, the model may behave noticeably differently from its peers, which would be valuable given the degree of behavioral convergence now visible across leading models.</p></li><li><p class="paragraph" style="text-align:left;"><b>It appears particularly strong at forecasting</b>, potentially frontier-leading according to the reported benchmark results below. That may indicate unusually good calibration, meaning its expressed confidence is more closely aligned with the real probability that its answer is correct (e.g., if the model tells you &quot;I&#39;m 70% confident the answer is x&quot;, the real probability is much more likely to actually be 70%). That could make it a more trustworthy model for decision-support applications.</p></li><li><p class="paragraph" style="text-align:left;">The smaller version, Inkling Small, looks terrifically cost-efficient.</p></li></ol><p class="paragraph" style="text-align:left;">We’ll have to see how things evolve, but for now, this is arguably my favorite AI Lab in the US, and it’s not particularly close.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>HARDWARE</b></span><br>First Agent Device?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8a840ae0-81dc-4eab-8ee5-675750ed1c40/image.png?t=1784283495"/></div><p class="paragraph" style="text-align:left;">OpenAI and hardware company <b>Work Louder</b> have launched the <a class="link" href="https://openai.com/supply/co-lab/work-louder/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">Codex Micro</a>, a compact programmable control pad designed for working with Codex coding agents.</p><p class="paragraph" style="text-align:left;"><span style="background-color:rgba(0, 0, 0, 0);">The device includes </span><b>13 mechanical keys, a rotary dial, a touch sensor, and a joystick</b><span style="background-color:rgba(0, 0, 0, 0);">.</span> Its keys can display live RGB status signals showing whether individual agents are thinking, running, waiting, or finished.</p><p class="paragraph" style="text-align:left;">Users can also assign shortcuts for actions such as accepting or rejecting changes, starting a chat, using push-to-talk, reviewing pull requests, debugging errors, and refactoring code.</p><p class="paragraph" style="text-align:left;">The rotary dial adjusts Codex’s reasoning level, while the joystick can trigger commonly used workflows. The device supports <b>Bluetooth and USB-C</b>, works with Mac and Windows, and includes 32 custom Codex keycaps.</p><p class="paragraph" style="text-align:left;"><b>OpenAI lists the Codex Micro at $230</b>, with clicky or silent switches, although it’s currently out of stock.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Incredibly nerdy device that, in reality, makes sense. People who interact with Codex as much as I do can really benefit from having a small set of buttons to instantly switch models, give instructions verbally, commit PRs easily, and so on.</p><p class="paragraph" style="text-align:left;">It does look like a delightful experience. <i>But will I buy it?</i> Hell no. Not for $200 plus almost $100 in shipping costs; <b>I’ve never ever seen such absurd shipping costs for any product</b>. Ever.</p><p class="paragraph" style="text-align:left;">It also feels somewhat suboptimal; I want to make a clearly suboptimized agentic experience feel less painful. I get it,<b> but I do hope we find better ways to interact with AI in the future.</b></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>STARTUPS</b></span><br>Mosquito-Killing AIs</h2><p class="paragraph" style="text-align:left;">In one of the craziest ideas I’ve ever seen, a YCombinator startup is trying to build tiny, 40-gram AI drones that kill mosquitoes. </p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/alextoussss/status/2077086243632873540?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar" target="_blank" rel="noopener noreferrer nofollow">As seen in this video</a>, they have recorded the first-ever execution of a mosquito using a drone, and their goal is to scale this so that we can, as they put it, <b><i>“eradicate mosquitoes.”</i></b></p><p class="paragraph" style="text-align:left;">The drones patrol your home incessantly, identifying targets and going for the kill. I swear I feel I’m taking the piss saying all this, <b>but they are literally building this and now have proof of execution to show for it.</b></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I don’t know what to think. I’m no biology expert, but as much as I hate mosquitoes,<i> I assume they must have some sort of role to play in nature?</i></p><p class="paragraph" style="text-align:left;">That said,<b> eradicating mosquitoes is not something new; </b>this is just the latest “ChatGPT, but to eradicate mosquitoes” AI play.</p><p class="paragraph" style="text-align:left;">Bill Gates’ famous mosquito programs that aim to eradicate malaria, dengue, and other diseases <b>by genetically modifying female mosquitoes to make them sterile</b>.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">It seems like the industry is turning against OpenAI and Anthropic, at least in terms of sentiment about their futures. I<b> always insisted that open-source, if not regulated away, would prevail</b>, for the very simple reason that AIs are trained on data, and companies would eventually want to train models with their data without having to gift it to the Labs.</p><p class="paragraph" style="text-align:left;">It was a long time coming,<b> but that time is now</b>.</p><p class="paragraph" style="text-align:left;">I’ve now grown much more concerned about the stability of the entire trade. I feel like this industry really needs Anthropic and OpenAI to go public and raise liquidity from retail investors, <b>but I’m not sure the appetite for those stocks is growing. If anything, it might be falling.</b></p><p class="paragraph" style="text-align:left;">And talking about stocks having a bad time on the markets lately, <b>we have the memory companies, arguably the most important stocks on the planet right now</b>, paying the price for taking all the blame (and the fall) amid growing fears around this industry collapsing any day now.</p><p class="paragraph" style="text-align:left;">But let me explain to you below why this is just wrong.</p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="background-color:#FF5632;" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=kimi-k3-president-xi-a-mosquito-killing-drone-walk-into-a-bar"><span class="button__text" style=""> Upgrade to Full Premium to continue reading </span></a></div></div><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/53f8bf15-00e7-4c53-851b-306c6da49f9f/image.png?t=1759853205"/></div></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=e73a85e6-2f0e-4137-96d5-34981103e48f&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The Ultimate Guide for AI, Part II</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9870eb38-7e70-4fe1-beae-4eac9ad050fc/ChatGPT_Image_Jul_15__2026__04_58_53_PM.png" length="2284990" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-part-ii</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-part-ii</guid>
  <pubDate>Thu, 16 Jul 2026 00:00:00 +0000</pubDate>
  <atom:published>2026-07-16T00:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>The Ultimate Guide for AI, Part II</h2><p class="paragraph" style="text-align:left;">Last week <a class="link" href="https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-in-2026?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii" target="_blank" rel="noopener noreferrer nofollow">I wrote my longest newsletter ever</a>, 8,500 words long, which covered the essentials of AI software and hardware, from the very basic key intuitions about AI models (what they are, how they learn), to the <b>key intuitions in AI hardware</b> that helped readers understand why GPUs and other accelerators are used, <b>why memory is so important</b>, and other <b>key ideas the industry holds dear</b>.</p><p class="paragraph" style="text-align:left;">And finally, we discussed the <b>“interesting” world of AI finance</b>, particularly some numbers to understand how this entire industry aims to make money.</p><p class="paragraph" style="text-align:left;">Today, I bring you the second part of the guide. In this one, I’ve focused more on the product, detailed AI inference maths, and future trends.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">What is an agent really?</h2><p class="paragraph" style="text-align:left;">Definitions of an agent are like opinions; everyone has one. Interestingly, despite everyone talking about it, most people don’t even understand what an agent is, when, in fact, it’s quite simple.</p><h3 class="heading" style="text-align:left;">Loops and tools</h3><p class="paragraph" style="text-align:left;"><a class="link" href="https://thewhitebox.beehiiv.com/p/the-agent-bible-first-act-context-engineering?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii" target="_blank" rel="noopener noreferrer nofollow">I’ve talked about agents in detail in the past</a>, so I won’t dwell too much on definitions, only the essentials. An agent has three components:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>The AI model</b>. Almost always a Large Language Model (LLM). Receives tokens, outputs tokens. Send new tokens; return new output tokens. It’s a continuous back-and-forth. Sometimes, the AI returns a tool call, a request to use a certain tool.</p></li><li><p class="paragraph" style="text-align:left;"><b>Memory.</b> Also known as the context system, it’s responsible for handling the input tokens the AI receives. <b>It determines what the AI receives</b> and, based on the AI’s outputs, <b>what must be remembered in future interactions</b>.</p></li><li><p class="paragraph" style="text-align:left;"><b>Tools</b>. These enable the AI to execute tools or actions.</p></li></ol><p class="paragraph" style="text-align:left;">At this point, it’s mandatory that we ask: <i>why do agents need external memory in the first place?</i></p><p class="paragraph" style="text-align:left;">We can understand the need for external tools (AIs can’t use tools directly; they are simply next-word predictions), but they do have internal memory and knowledge (known as parametric memory), <i>so why the need for external systems providing extra context?</i></p><p class="paragraph" style="text-align:left;"><i>Doesn’t the AI know it all?</i></p><h3 class="heading" style="text-align:left;">Statelessness and the continual learning problem</h3><p class="paragraph" style="text-align:left;">When you think about it, it’s pretty ironic that AIs, which have “seen it all,” need support from an external memory component; much like a human has Google search a smartphone unlock away, <b>AIs need external memory too</b>.</p><p class="paragraph" style="text-align:left;"><i>But why?</i> And the answer is twofold: AIs are <b>stateless</b> and have a <b>continual learning problem</b>.</p><p class="paragraph" style="text-align:left;">The first one is crucial to understanding the dynamics of what is going on under the hood when you interact with Gemini, Claude, or any LLM for that matter.</p><p class="paragraph" style="text-align:left;">LLMs don’t have state. Every single interaction is like a new blank page in a book. Unless they’ve been trained on you, if the ChatGPT app doesn’t provide some context about you, when you click ‘New chat’ and you send something to the model,<b> it’s like the first interaction it’s ever had with you</b>. And if you click ‘New chat’ again, the model loses context from the previous conversation and starts anew with you.</p><p class="paragraph" style="text-align:left;">What this means is that it <b>doesn’t have an ‘internal state’ that carries from one conversation to the next</b>. It knows everything, but can’t remember what happened 5 seconds ago. It’s as if an encyclopedia and Dory from <i>Finding Nemo</i>, who forgets everything every three seconds, had a baby.</p><p class="paragraph" style="text-align:left;">That lack of state forces AI Labs to add it externally. When the previous conversation ends, unbeknownst to you, a background process may run, identify key details you shared in the conversation, <b>and add them to all future prompts</b>.</p><p class="paragraph" style="text-align:left;">If you mentioned you liked Tom Brady a lot, the harness will catch this, add a <i>“the user loves Tom Brady”</i> snippet to future prompts, and if you ever ask it <i>“Who’s my favorite player?”</i> <b>the LLM will “magically” know despite not really knowing it</b>; if the harness decides to eliminate that memory and no longer add it to the prompt, suddenly, the model doesn’t know any more who’s your favorite sports star.</p><p class="paragraph" style="text-align:left;">Of course, that also means that the memory system must work well. That process (which, by the way, <a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/anthropics-models-can-now-dream-but-how-fd13858a2a0d?sk=68613dd2b21f6c36b6c0bfd99b008ead&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii" target="_blank" rel="noopener noreferrer nofollow">is called ‘Dreaming’,</a> which I wrote about here in case you want to understand it further) isn’t perfect and may not pick the right memories.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But the takeaway I want you to, well, take away,<b> is that this is very much an external process</b>, and no matter how <i>much </i>better an AI becomes at “knowing you”, <b>it doesn’t really know you at all unless it’s trained on your data</b>, which is never the case as things stand today.</p><p class="paragraph" style="text-align:left;">Therefore, to this day, LLMs do not have state, and the quality of their knowledge about you depends entirely on the harness. But I know what you’re thinking: <i>can’t we just train AIs on the user’s outputs so that they become knowledgeable about us by default?</i></p><p class="paragraph" style="text-align:left;">And this, my dear reader, <b>is the continual learning problem.</b></p><p class="paragraph" style="text-align:left;">While we have cracked the code for training excellent models, we have yet to crack the code for retraining them at scale.</p><p class="paragraph" style="text-align:left;">In other words, <b>we don’t know yet how to retrain general models without breaking them</b>. Careful, I’m not saying we don’t know how to fine-tune them. We do, of course, but we’re doing so knowing that we’re breaking them in one way or another due to what the industry calls <b>“catastrophic forgetting”</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But companies like OpenAI or Anthropic do care about breaking the model (i.e., training on something while making it forget other stuff), because consumers can one day ask about cake recipes and the next about quantum physics.</p><p class="paragraph" style="text-align:left;">Worse, it’s not predictable;<b> the only thing we can predict is that it will happen, but not where</b>. As shown below by research from Thinking Machines, training a model on internal docs improved the performance for that use case, while also catastrophically degrading the model’s ability to follow instructions, essentially killing the model.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8d31b51e-630b-4c55-8ba9-a89b6ca28181/image.png?t=1784121856"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://thinkingmachines.ai/blog/on-policy-distillation/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;"><i>But why does this happen?</i></p><p class="paragraph" style="text-align:left;">Because you’re modifying the model and neurons are polysemantic, a certain neuron may be in charge of storing information about maths principles and also about turtles, so modifying that neuron could make it ‘drop’ the maths principles to absorb more ‘turtle data’ even if you didn’t intend that outcome.</p><p class="paragraph" style="text-align:left;">In the meantime, <b>our best bet is to simply have a good harness on top of models</b>, providing that additional context in the prompt.</p><p class="paragraph" style="text-align:left;">And what about the third component of the modern AI harness, <i>tools? </i>It’s important that you understand this often-neglected part of the AI story.</p><h3 class="heading" style="text-align:left;">Tools and the CPU story</h3><p class="paragraph" style="text-align:left;">As the name suggests, <b>the ‘tools’ component of the LLM harness “offers” tools to the AI to choose and leverage towards the assigned task</b>. As we have discussed, AIs can’t use tools; they don’t have either a physical or a digital body (the harness is the body, in a way); they just predict tokens.</p><p class="paragraph" style="text-align:left;">Therefore, when it wants to use a tool, it doesn’t respond with words; <b>it responds with a ‘tool call,’</b> a declaration of intent for the harness to run that tool with those specifications and return the result to the AI.</p><p class="paragraph" style="text-align:left;">Tools have numerous implications.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>They allow AIs to execute actions on external systems</b>. Tools are needed for your AI to register leads for you or send emails.</p></li><li><p class="paragraph" style="text-align:left;"><b>They severely increase sequence length</b>. Not only does the AI need to have tool definitions in the prompt to know what tools it can use, but also which ones to use, but the outputs of those tools have to be fed back to the AI too.</p></li><li><p class="paragraph" style="text-align:left;">They need to be lightweight, which, by the way, <b>is a death sentence for most SaaS companies today</b>, as I believe it is for most non-AI software, which will end up being agentic tools, which will be required to be cheap, compressing margins unless you lay off 80% of your workforce. In other words, <b>I believe most SaaS companies will be massive relative to how small they’ll need to be to compete in the future</b>.</p></li><li><p class="paragraph" style="text-align:left;">They move the focus (and latency) from GPUs to CPUs.</p></li></ol><p class="paragraph" style="text-align:left;">Regarding the fourth, <a class="link" href="https://thewhitebox.beehiiv.com/p/the-time-of-the-cpu-has-come?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii" target="_blank" rel="noopener noreferrer nofollow">I discussed it at length in this piece</a> (I also compare CPU vendors), but the short explanation, as mentioned above, <b>is that tools shift latency from the GPU to the CPU.</b></p><p class="paragraph" style="text-align:left;">As shown by Intel below, for workloads that include tool calling, most of the model’s latency is attributable to the CPU running tools, not to the AIs generating tokens on the GPUs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/853a501e-9303-4a6d-8f69-4e42499d920f/image.png?t=1784106402"/><div class="image__source"><span class="image__source_text"><p>Source: Intel</p></span></div></div><p class="paragraph" style="text-align:left;">Today, with much heavier tool-heavy workloads, the picture is way worse. This is one of the primary reasons NVIDIA introduced an NVLink-type interconnect between the GPU and the CPU for Blackwell onward (which are traditionally connected via a PCI bus), <b>effectively increasing the communication bandwidth between the host and the device by seven times </b>(CPU and GPU).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/91f3571f-1d93-4722-bc6f-de63c291e33b/image.png?t=1784106591"/><div class="image__source"><span class="image__source_text"><p>Source: <a class="link" href="https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii" target="_blank" rel="noopener noreferrer nofollow">NVIDIA</a></p></span></div></div><p class="paragraph" style="text-align:left;">And talking about tools lets us perfectly bridge the conversation with the effect all of this is having on AI as a business.</p><p class="paragraph" style="text-align:left;">And it’s a lot.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-part-ii">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=f0871533-89db-46a2-8329-92988d1e7d1a&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Price Wars are Nigh, SensorFM, Harness Matter, &amp; More</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/74e2d37e-a662-4237-a2f0-b727005476a6/image.png" length="293602" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/price-wars-are-nigh-sensorfm-harness-matter-more</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/price-wars-are-nigh-sensorfm-harness-matter-more</guid>
  <pubDate>Sun, 12 Jul 2026 13:00:00 +0000</pubDate>
  <atom:published>2026-07-12T13:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/><div class="image__source"><span class="image__source_text"><p> </p></span></div></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back! </b>This week, we have a lot to discuss. The last few days were probably the most active for top model releases from the US in perhaps a year, <b>with two Labs that had fallen from grace somehow making a comeback</b>.</p><p class="paragraph" style="text-align:left;">We also discuss really cool tech from Google in <b>SensorFM</b>, a robotics hand that is the stuff of nightmares, a <b>Databricks</b> research showing that AI is much more than the model itself, and more.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>A Universal AI Sensor Model</h2><p class="paragraph" style="text-align:left;">In one of the most heated weeks so far, with multiple incredible model releases, <a class="link" href="https://research.google/blog/sensorfm-towards-a-general-intelligence-and-interface-for-wearable-health-data/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">my favorite news has been a new paper from Google and its model SensorFM</a>.</p><p class="paragraph" style="text-align:left;">This model has been trained on data from 5 million people and has learned a general-purpose representation of human physiology that transfers across<span style="background-color:rgb(255, 255, 255);"> 35 health prediction tasks.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Ok,</span><span style="background-color:rgb(255, 255, 255);"><i> but what does that even mean?</i></span><span style="background-color:rgb(255, 255, 255);"> Well, it means I’m excited now </span><span style="background-color:rgb(255, 255, 255);"><b>because it opens the path to personalized health</b></span><span style="background-color:rgb(255, 255, 255);">. Let me explain why.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Currently, billions of wearables track sensory signals from the wrists of millions of people. However, </span><span style="background-color:rgb(255, 255, 255);"><b>these devices are extremely noisy, and translating that sensor data into meaningful health insights is debatable at best</b></span><span style="background-color:rgb(255, 255, 255);">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Each company measures differently, each device registers the data differently, and, importantly, each human is different. This means that, in reality, </span><span style="background-color:rgb(255, 255, 255);"><b>the signal-to-noise ratio is very low</b></span><span style="background-color:rgb(255, 255, 255);">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);"><i>But what if we trained an AI to process the data across millions of humans?</i></span><span style="background-color:rgb(255, 255, 255);"> That is what Google did, training the AI on one trillion minutes of data, </span><span style="background-color:rgb(255, 255, 255);"><b>or roughly 2 million years</b></span><span style="background-color:rgb(255, 255, 255);">, and the results are pretty incredible.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">But first: </span><span style="background-color:rgb(255, 255, 255);"><i>how do we train a model with that data?</i></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">If you read </span><span style="background-color:rgb(255, 255, 255);"><a class="link" href="https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-in-2026?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">my last newsletter</a></span><span style="background-color:rgb(255, 255, 255);">, which I recommend you do, </span><span style="background-color:rgb(255, 255, 255);"><b>I explained that an AI can only learn what can be measured</b></span><span style="background-color:rgb(255, 255, 255);">. And by &quot;measure,&quot; I mean models learn by making predictions and comparing them to a ground truth (what they should actually have predicted).</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">For these one trillion minutes, we don’t have a ground truth (officially known as ‘label’); we don’t know what each of these one trillion minutes is showing. </span><span style="background-color:rgb(255, 255, 255);"><i>Depression? Anxiety?</i></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Ideally, you would want a dataset that describes the following: if sensor data looks like this, the human has heart problems. If the data looks like this, the human is fine. Eventually, </span><span style="background-color:rgb(255, 255, 255);"><b>the AI makes the connection and learns to associate certain data patterns with certain outcomes</b></span><span style="background-color:rgb(255, 255, 255);">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">But if there’s no way to measure the outcome, </span><span style="background-color:rgb(255, 255, 255);"><i>how are they supposed to learn?</i></span><span style="background-color:rgb(255, 255, 255);"> And the answer is unsupervised learning, </span><span style="background-color:rgb(255, 255, 255);"><a class="link" href="https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-in-2026?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">which I did not include in an already 8,000-long piece</a></span><span style="background-color:rgb(255, 255, 255);">, </span><span style="background-color:rgb(255, 255, 255);"><b>but it’s essentially a rare type of training where the data itself is the signal</b></span><span style="background-color:rgb(255, 255, 255);">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Essentially, we’re telling the AI: </span><span style="background-color:rgb(255, 255, 255);"><i>“We can’t really associate the patterns you’re going to find with particular outcomes. However, we still want you to find them.”</i></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);"><i>But why would we want this?</i></span><span style="background-color:rgb(255, 255, 255);"> For things like clustering. Although the AI doesn’t fully understand the implications of each pattern, it can still classify them.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">If we give an AI a huge unlabeled dataset of flowers, the model might know what each flower actually is—it might not even know what a flower is—but it learns to classify them nonetheless; </span><span style="background-color:rgb(255, 255, 255);"><b>it will still learn to separate roses from carnations even if it doesn’t know what those are</b></span><span style="background-color:rgb(255, 255, 255);">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Here, they’ve done the same with wearable sensor data.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">Okay, got it. At this point, we have a great-but-useless model. We have a global representation of sensor data, a model that receives sensorial input and captures key patterns. However, </span><span style="background-color:rgb(255, 255, 255);"><b>we have no way to decode those patterns, to associate them with actual outcomes.</b></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);">For that, they add extra training phases, but ones focused on selected, well-labeled data from real humans with real metabolic, mental, and sleep-related signals, basically saying: “this is the sensorial data for this human, and here’s how they actually feel”. </span></p><p class="paragraph" style="text-align:left;">With that, <i>we can now associate the pattern with the outcome!</i></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/51b34ce5-c8c9-49a9-b689-a7013097a0b4/image.png?t=1783680516"/></div><p class="paragraph" style="text-align:left;">Interestingly, for these new training phases, <b>they had AIs design them, using the idea of an ‘LLM classroom’</b>. They would give the AI the training constraints, and the groups of LLMs would design the experiments, leading to training regimes that beat human-designed experiments by a long shot.</p><p class="paragraph" style="text-align:left;">And the results are pretty good, I have to say.</p><p class="paragraph" style="text-align:left;">For starters, this universal representation, this one-size-fits-all model, <b>can capture dependencies across 35 health domains</b> and shows great prediction accuracy potential across many of them:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/68ce3e75-71cc-441d-843b-cec06b117d60/image.png?t=1783681070"/></div><p class="paragraph" style="text-align:left;">Importantly, the AI&#39;s potential struggles in some areas are not the model’s fault per se; they may simply reflect what many already believe to be true: <b>wearable devices produce very noisy data.</b></p><p class="paragraph" style="text-align:left;">But perhaps a more fascinating outcome was testing whether SensorFM could serve as a good context engineer for an LLM acting as a health agent.</p><p class="paragraph" style="text-align:left;">In other words, the Gemini agent would receive predictions from the SensorFM about the user in particular based on sensor data; things like age, predicted BMI, or anxiety scores (see below for an example), and the Gemini agent not only provided much more meaningful recommendations,<b> but using the actual ground truth data from the user (actual age, BMI, or insulitn resistance) did not improve recommendation quality relative to using SensorFM</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/27f1bde1-ba98-4a40-8653-916f8dbfdccb/image.png?t=1783681191"/></div><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I’m a staunch believer that healthcare is one of the key domains where AI will change our world the most.</p><p class="paragraph" style="text-align:left;">The option of offering personalized health recommendations to every human at scale, something our current system can’t for the life of it offer (healthcare systems around the world are completely broken), <b>is something society can’t possibly allow not to happen.</b></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ROBOTICS</b></span><br>1X’s New Robot Hand is the Stuff of Nightmares</h2><p class="paragraph" style="text-align:left;">1X, a US robotics company, has announced its new robotic hand. And let me tell you, <a class="link" href="https://x.com/1x_tech/status/2075252899442204952?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">it’s genuinely incredible</a>. This new hand has <b>25 degrees of freedom</b> (meaning it can independently control 25 different joint movements), making it incredibly dexterous and gentle.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.youtube.com/watch?v=QRyXV3csReA&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">As shown in this video</a>, the robot can pick up glass and grapes without breaking them, play video games, and more.</p><p class="paragraph" style="text-align:left;">They now claim Neo (the robot) can perform any task with its hands that a human can, and that the robot is also strong (enough to pick up weights).</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Moreover, the hand’s skin serves as the sensor channel, allowing the robot to measure force and even detect when an object is slipping. Truly alien technology.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I’m happy to see Western robotics labs putting focus on the hardware, <b>not just on the “brains.”</b> This is the most common approach in China, with examples like Unitree, because many people believe (and so do I) <b>that hardware is way harder than software in robotics.</b></p><p class="paragraph" style="text-align:left;">Of course, the big question is how long it will take to transition from marketing art to actual usable products. <b>My opinion is that we’re yet to hit robotics ‘ChatGPT moment’ </b>and that we must remain optimistic but realistic about timelines.</p><p class="paragraph" style="text-align:left;">I know we were recently talking about a robot coming this fall, but I won’t believe it until I see it.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ENGINEERING</b></span><br>How Much Does the Harness Matter?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/74e2d37e-a662-4237-a2f0-b727005476a6/image.png?t=1783841410"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">Databricks has published some results</a> on agents that may surprise some people.</p><p class="paragraph" style="text-align:left;">As shown in the graph above, the choice of harness (e.g., Claude Code from Anthropic, a third-party option called Pi, and Codex, OpenAI’s harness) not only significantly impacts performance on the same underlying models <b>but can also considerably elevate the performance of inferior models.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">The best example is GLM-5.2 using the Pi harness, which beats Opus 4.8 using Anthropic&#39;s own harness, Claude Code. This result might be surprising, but the interesting insight for me is that i<b>t&#39;s perfectly reasonable that third-party harnesses beat the harnesses of the model Labs because the latter group has misaligned incentives</b>; they want to build a good harness for you but, at the same time, they want their harnesses to push models into consuming as many tokens as possible, so it&#39;s reasonable to assume that <b>if you&#39;re running your AIs on harnesses built by the same Lab that created the model, you&#39;re going to pay more</b>.</p><p class="paragraph" style="text-align:left;">This is palpable in the image below, where the third-party harness running OpenAI and Anthropic models required between 2 and 4 times fewer tokens than those same models running in the &quot;official&quot; harnesses.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/707581e0-816d-42d9-a310-e4ae2e1b89bb/image.png?t=1783841573"/></div><p class="paragraph" style="text-align:left;">In short,<b> you have to be really careful about what system you choose</b>, not just the underlying models, as it can really impact your wallet.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">An interesting point is that the Databricks team proves something most people struggle to understand: <b>pricier models aren&#39;t necessarily more expensive overall</b>, because larger models (especially the new ones like Fable or GPT-5.6) require far fewer tokens.</p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">For example, they show that Sonnet 5, despite being 1.7x cheaper </span>per token than Opus 4.8, <b>was more expensive overall </b>because it required 1.9x as many<span style="background-color:transparent;"> tokens for the task. </span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">This should remind companies to decide which model to use based not just on token price,</span><span style="background-color:transparent;"><b> but also to test models for overall expenditure</b></span><span style="background-color:transparent;"> (token price x number of tokens used) before deciding.</span></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And to be clear, the results from running the Pi harness are obtained using the exact same underlying model as Codex/Claude Code, <b>so this is a pure harness-cost difference</b>.</p><p class="paragraph" style="text-align:left;">To me, the immediate question an investor must ask is: <i>if harnesses play such a vital role in both performance and cost, and third parties can beat Labs at their own harness game, </i><b><i>where&#39;s the moat?</i></b></p><p class="paragraph" style="text-align:left;">I have my thoughts on what I believe is the real moat in AI. To me, <b>the answer is something like what Anthropic is trying to do with </b><b><a class="link" href="https://www.anthropic.com/news/introducing-claude-tag?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow" style="color: rgb(17, 85, 204)">Claude Tag</a></b><b>,</b> a product that lets users in your company interact with Claude via Slack. People are taking it as just another AI product these days, but that&#39;s, in fact, <b>incredibly wrong</b>.</p><p class="paragraph" style="text-align:left;"><b>This product is a Trojan Horse that will &quot;marry&quot; you to Anthropic forever</b>. The reason is that Claude Tag is not just “another harness”. It’s a system comprising the harness and a model trained on Slack&#39;s synthetic environments.</p><p class="paragraph" style="text-align:left;">The point here is that the underlying model is purpose-built for Slack, so it offers performance that is basically unmatched. Furthermore, it is built to become entrenched in your organization, parsing, processing, and learning from your data.</p><p class="paragraph" style="text-align:left;">As I said once, <b>there’s a trade-off between privacy and personalization</b>. If you want a truly personalized AI, you need it to see your secrets. <b>That is why it would be hilariously naive to build your personalized AI agents on third-party software</b>.</p><p class="paragraph" style="text-align:left;">For these companies, it’s a combination of personalized AI training in your environments, a strong harness, and, especially, personalization by knowing everything about you, that creates the moat.</p><p class="paragraph" style="text-align:left;">It’s not the AI per se; <b>the real moat is the entire product.</b></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>SEMICONDUCTORS</b></span><br>Seeing the bottlenecks clearly</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e09b6f00-4bdc-4386-95e4-bb8975738c1d/image.png?t=1783842062"/></div><p class="paragraph" style="text-align:left;">This interesting exhibit from Goldman Sachs shows the price changes across the semiconductor chain, which can explain vendor pricing power.</p><p class="paragraph" style="text-align:left;">As you can see, the price hikes in memory and fast storage (DRAM and NAND, respectively) are something to behold, <b>with GS seeing considerable supply tightness at least throughout H1 2027</b>, although the companies in those sectors, mostly the Big 3 (SK Hynix, Samsung, and Micron), <b>believe the supply tightness could remain very strong through 2028</b>.</p><p class="paragraph" style="text-align:left;">Beyond memory, you could argue almost everything is tight.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">A note of caution: this can be easily misinterpreted as a “bottleneck ranking” to suggest which companies have the highest pricing power.</p><p class="paragraph" style="text-align:left;">And although supply tightness and pricing power are indeed deeply correlated, some companies may simply decide not to raise prices even if they could.</p><p class="paragraph" style="text-align:left;">For example, <b>companies like TSMC and ASML are known to be particularly unwilling to raise prices</b>, even though they could easily do so (they are basically monopolies in their respective markets).</p><p class="paragraph" style="text-align:left;">In fact, I maintain that the biggest bottleneck in semis is not memory, <b>but advanced packaging</b>, the part of the process that packs the compute and memory chips into a single package, giving you the actual accelerator (e.g., GPUs, TPUs, etc.), which is largely dominated by TSMC.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>OpenAI Releases GPT-5.6 Luna, Terra, and Sol</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c2beaf2e-3279-4901-ac60-9502399286bc/image.png?t=1783760520"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://openai.com/index/gpt-5-6/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">OpenAI has finally released its newest models</a>, Sol, the strongest, Terra, and Luna, the smaller one.</p><p class="paragraph" style="text-align:left;"><b>The most remarkable thing seems to be its much-improved cost-per-intelligence</b>, being considerably cheaper than previous OpenAI models and especially compared to Anthropic’s models (take the performance scores with a pinch of salt). And as we’ll see below in the Grok 4.5 article, OpenAI’s models are now all three on the “Pareto frontier,” <b>meaning they offer the best performance per cost amongst all models, at their respective sizes.</b></p><p class="paragraph" style="text-align:left;">But without a doubt, the most impressive part of the release was that, according to OpenAI, Luna was trained solely by Sol. In layman’s terms, <b>the smaller model was built by the larger model</b>, with zero human participation and in a zero-shot fashion.</p><p class="paragraph" style="text-align:left;">In fact, they shared a snippet of the prompt they gave the larger model, trying to convey the idea <b>that RSI</b> (Recursive Self-Improvement) or AIs helping or autonomously creating better AIs, <b>is starting to happen</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b31a50f7-11a1-44a5-8802-0937cc178da9/image.png?t=1783759692"/></div><p class="paragraph" style="text-align:left;">The real breakthrough will happen once AIs are capable of training better AIs, not smaller versions of themselves, because in reality, although still impressive, <b>this is just an AI model running the training run of another model; the tests, experiments, and all</b>.</p><p class="paragraph" style="text-align:left;">I insist it’s really cool, but in reality, it’s just an AI running a bunch of scripts.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Considering Fable’s guardrails and all, <b>these are quite possibly the best models you can use today.</b> Personally, after several days of using them, I would argue that GPT-5.6 Terra in high mode is the best bang for your buck.</p><p class="paragraph" style="text-align:left;">However, overall, I haven’t felt much of a change. <b>This could be the first indication that I might be reaching my personal ceiling of capabilities</b>, meaning, yes, for some stuff like coding, these models feel genuinely better, but for many tasks I use AIs for, like discussing abstract stuff about papers, critiquing my own opinions, and the lot, <b>I barely notice a difference with the previous generation</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><i>Is that a sign that these models are becoming so good that we can’t notice improvements anymore, or are they hitting a wall?</i></p><p class="paragraph" style="text-align:left;"><b>I fear it might be a mixture of both</b>. I don’t mind these AIs becoming incredibly smart, but what worries me is that, for most economically valuable tasks, incremental improvements won’t improve results and will simply make your bill larger.</p><p class="paragraph" style="text-align:left;">Could it be that AIs are becoming too good for their own good, meaning they are too expensive relative to what they offer, <i>because what they offer is simply overqualified for most white-collar work?</i></p><p class="paragraph" style="text-align:left;">I recently listened to a podcast where the guest basically discredited the idea of AIs destroying jobs not because they could<b>, but because he argued most jobs were “made up” and there was nothing of economic value to disrupt</b>.</p><p class="paragraph" style="text-align:left;">And while I’m not sure I would put it that way, after almost a decade as a consultant for corporations, <b>I have come across my fair share of jobs that literally produced zero value despite those people earning 6 figures or more</b>.</p><p class="paragraph" style="text-align:left;">It’s going to be interesting to see how this industry handles the slight possibility that most economic activity is non-disruptible. A potential outcome is that AIs open the door to new types of jobs and economically valuable activities that do not exist today.</p><p class="paragraph" style="text-align:left;">Who knows, <b>but I remain skeptical about how much it can disrupt what exists today</b>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Notion Launches ShipOS</h2><p class="paragraph" style="text-align:left;">Notion, a note-taking app, <a class="link" href="https://www.notion.com/product/ship-os?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">has announced an agent-orchestrator product called ShipOS</a>, a way to organize your agents from multiple sources (Codex, Claude Code, etc.) into a single Kanban board where you can assess progress and make better decisions.</p><p class="paragraph" style="text-align:left;">Basically, an agent manager software, something very similar to what Linear already offers.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The reason I’m pointing this release out to you is that I recommend you watch the presentation video, <b>as I believe it clearly shows what the future of managing agents might look like.</b></p><p class="paragraph" style="text-align:left;">Listen, I use coding agents daily, and I continuously rotate from <i>“this is the best thing ever”</i> to <i>“I really, really hate them”</i> because there’s too much information being thrown at you, agents lose the plot… <b>and many other concerns that tell me the current form factor is clearly suboptimal.</b></p><p class="paragraph" style="text-align:left;">Assuming looking at the code is almost impossible considering the rate at which these models generate code, <b>we need to find an abstraction layer that allows us to see the things that have to be seen</b>, even some code snippets in particular, and that way make good decisions.</p><p class="paragraph" style="text-align:left;">Right now, I feel like most of the time I’m just telling it to do whatever it suggests, because I really don’t have the time to judge every decision it makes among the thousands of agents&#39; decisions made each day.</p><p class="paragraph" style="text-align:left;">Sadly, <b>I know what I’m doing isn&#39;t how software should be built</b>; you need more control, you need to “own” your work, but I don’t think the current status quo lets you do that without scrutinizing every single thing these agents throw back at you, which is impossible.</p><p class="paragraph" style="text-align:left;">Whether it’s Linear or Notion, <b>these companies have a golden opportunity to own that abstraction layer</b>, but I don’t think anyone really knows for sure what it will look like in six months from now. But this is a start.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Meta Joins the Race with Muse Spark 1.1</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/10bf858c-cc90-494a-9a7b-e18e8d29b1b3/image.png?t=1783762285"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.reuters.com/business/meta-debuts-muse-spark-11-with-preview-open-developers-2026-07-09/?utm_source=chatgpt.com" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://www.reuters.com/business/meta-debuts-muse-spark-11-with-preview-open-developers-2026-07-09/?utm_source=chatgpt.com" target="_blank" rel="noopener noreferrer nofollow">reported by Reuters</a> and Meta, Meta has released <b>Muse Spark 1.1</b>, a multimodal AI model designed for coding, debugging, and complex agentic tasks involving external tools and multiple steps.<b> It is now available to US developers through the new Meta Model API in public preview.</b></p><p class="paragraph" style="text-align:left;">Besides the fact that the model looks very competitive in terms of raw performance, the real highlight is performance per cost: API pricing starts at <b>$1.25 per million input tokens</b> and <b>$4.25 per million output tokens</b>, with $20 in introductory credits.</p><p class="paragraph" style="text-align:left;">Not only is this a highlight in itself, <b>given it’s the first time Meta will serve its models via APIs</b>, but the prices are also considerably lower than at the frontier. Muse Spark’s prices are even lower than the API price of the best Chinese model, GLM-5.2.</p><p class="paragraph" style="text-align:left;">In fact, Meta’s prices are so outrageously low relative to peers that <b>they are the cheapest option available right now</b>, including some Chinese Labs like Zhipu and its GLM-5.2, <b>with only DeepSeek offering much more competitive pricing</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/859a6926-13f2-40d7-8bb9-5005674b302a/image.png?t=1783840157"/></div><p class="paragraph" style="text-align:left;">Also, much as we saw with the Databricks evals, <b>a lower token price does not necessarily mean a cheaper overall task cost</b>.</p><p class="paragraph" style="text-align:left;">As you can see above, <b>while GPT-5.6 Luna is 27% more expensive per token, the overall cost per task on AA is lower because it uses fewer tokens</b>. I suspect enterprises will soon transition from measuring costs at the token level to the task level (i.e., what&#39;s the cheapest model for my task?) instead of measuring this abstract idea of a ‘token’.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>SpaceXAI Joins the Race Too</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.ai/news/grok-4-5?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">Grok 4.5 was launched on July 8</a> as SpaceXAI’s latest model for coding, autonomous-agent tasks, and technical knowledge work.</p><p class="paragraph" style="text-align:left;">The company says it was trained alongside Cursor, <b>the team of talented engineers SpaceX paid $60 billion for a couple of months ago</b>. Training used tens of thousands of NVIDIA GB300 GPUs, which means it’s a large, large model.</p><p class="paragraph" style="text-align:left;">SpaceXAI reports that Grok 4.5 scored 62% on DeepSWE 1.0, 29% on SWE Marathon, 83.3% on Terminal Bench 2.1, and 64.7% on SWE Bench Pro, <b>signaling very strong performance on coding and agents</b>.</p><p class="paragraph" style="text-align:left;">The model is served at about 80 tokens per second, which is reasonably fast by today’s inference standards. It’s also considerably token-efficient.</p><p class="paragraph" style="text-align:left;">On SWE Bench Pro, it generated an average of 15,954 output tokens per task, compared with 67,020 for Anthropic’s Opus 4.8 in SpaceXAI’s comparison. <b>The company says this makes Grok 4.5 roughly twice as token-efficient as comparable models overall.</b></p><p class="paragraph" style="text-align:left;">They don’t seem to be lying, considering that Artificial Analysis places this model in the green quadrant alongside Meta’s Muse Spark and OpenAI’s GPT-5.6 Luna as the most cost-efficient models out there right now.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/10c34060-5ea3-48cc-9558-a9b6f751117c/image.png?t=1783840812"/></div><p class="paragraph" style="text-align:left;">API pricing is <b>$2 per million input tokens</b> and <b>$6 per million output tokens, making it very competitive</b>, as you can see in the table in the Muse Spark news above.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Very solid release that suggests SpaceXAI is back on track after several months of falling very behind. It’s clear that the Cursor acquisition has benefited them a lot.</p><p class="paragraph" style="text-align:left;">Overall, based on this release and that of Meta’s, <b>it seems we are amidst a new price war on token prices</b>.</p><p class="paragraph" style="text-align:left;">However, <b>this time</b><b> the price pressures are not coming from China but from within the US</b>, specifically from Meta and SpaceXAI&#39;s new models.</p><p class="paragraph" style="text-align:left;">As discussed above, Meta&#39;s latest model, Muse Spark 1.1, is comparable in capability to Opus 4.8 and GPT-5.5 (though not quite at the Fable/GPT-5.6 level), <b>but its API prices undercut even China&#39;s GLM-5.2</b>. Grok 4.5 is in the same performance ballpark, only slightly pricier per token but offering slightly higher performance.</p><p class="paragraph" style="text-align:left;">Most fascinatingly, <b>it seems the tables have turned in terms of performance per task relative to China</b>. OpenAI is now very comfortably leading the Pareto frontier with its three new models (Luna, Terra, and Sol), <b>and the green quadrant is now dominated by three US models</b>. How the turn tables.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">Very interesting week with lots of different topics to attend to. But to me the clear highlight is that, well, we seem to be bracing ourselves for a new price war.</p><p class="paragraph" style="text-align:left;">However, we’re forced to ask ourselves: <span style="background-color:transparent;"><b><i>Can this industry afford a new price war, even though it has yet to prove profitability?</i></b></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;"><i>And what will </i></span><span style="background-color:transparent;"><i><b>Anthropic</b></i></span><span style="background-color:transparent;"><i> </i></span><i>in particular</i><span style="background-color:transparent;"><i>, which is by far the most expensive option, do now?</i></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;"><i>Is Fable&#39;s 3x cost per task over GPT-5.6 Sol justifiable for a single point higher score? And being 10 times cheaper than Muse, with only 9 points less?</i></span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">To me, the answer is, of course, </span><span style="background-color:transparent;"><b>a hilariously clear no</b></span><span style="background-color:transparent;">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">And although we always say benchmarks never tell the whole story, and it’s true, </span><span style="background-color:transparent;"><b>I bet every CIO/CFO right now is looking at these numbers and trying to figure out how to transition out of Anthropic</b></span><span style="background-color:transparent;">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">That said, as these Labs seem to be delaying their IPOs, they can care less about margins for a while and, as long as revenues go up, the hype will continue.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">But a price war at a time when </span><span style="background-color:transparent;"><a class="link" href="https://newsletter.semianalysis.com/p/anthropic-3q26-profit-over-1b-the?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=price-wars-are-nigh-sensorfm-harness-matter-more" target="_blank" rel="noopener noreferrer nofollow">many analysts had concluded that Anthropic would be profitable in Q3</a></span><span style="background-color:transparent;"> and that this ‘AI bet’ would finally make sense would be sort of funny.</span></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=80574c33-80f7-4810-9b08-57af805d6577&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The Ultimate Guide for AI in 2026</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/44f5ece7-5e50-4e1d-b11d-350e02988862/ChatGPT_Image_Jul_8__2026__04_03_34_PM.png" length="2072420" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-in-2026</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/the-ultimate-guide-for-ai-in-2026</guid>
  <pubDate>Wed, 08 Jul 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-07-08T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>The Ultimate Guide for AI in 2026</h2><p class="paragraph" style="text-align:left;">Recently, a client asked me the following:<i> If I had to explain the AI industry in one go, what would you give me? </i>This had to be explained in a high-level, very intuitive way for anyone who’s not quite as deep into AI as I am to follow.</p><p class="paragraph" style="text-align:left;">In particular, answer the following:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">What needs to be known across the entire value chain?</p></li><li><p class="paragraph" style="text-align:left;">In this very dynamic industry, what is constant?</p></li><li><p class="paragraph" style="text-align:left;">What bets is the industry making</p></li><li><p class="paragraph" style="text-align:left;">What does the future hold?</p></li></ol><p class="paragraph" style="text-align:left;">In this piece, we’re doing just that. It’s my longest article ever, <b>but I’ve genuinely never packed more information and insights into a single piece</b>. Ever. It was a hustle (I honestly don’t know when I decided it was a good idea to put so much effort into something that can be purchased for $20/month, but I guess I really dislike doing things halfway).</p><p class="paragraph" style="text-align:left;">Alas, I hope you enjoy it.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">The Technology</h2><p class="paragraph" style="text-align:left;">Understanding the technology is not just about memorizing the key terms; it’s about understanding why it looks the way it does, why it needs the hardware it uses, and why it takes the particular product shape it does.</p><h3 class="heading" style="text-align:left;">It’s all about patterns</h3><p class="paragraph" style="text-align:left;">An AI algorithm is just a method of processing data. By ‘processing’ we mean capturing the underlying patterns in data to make predictions about it.</p><p class="paragraph" style="text-align:left;">The constant definition you see in the wild is that it learns patterns from known data to make predictions on unseen but similar data.</p><p class="paragraph" style="text-align:left;">Say you want an AI to predict housing prices, trained on millions or billions of housing data points with known prices. Eventually, i<b>t starts picking up regularities in the data</b>, such as <i>“houses with many rooms tend to be pricier,”</i> <i>“postal codes in this area tend to be cheaper,”</i> or <i>“marble countertops are often present in expensive homes.”</i></p><p class="paragraph" style="text-align:left;">With these captured patterns, the AI model can make predictions about new homes it has not seen before. It doesn’t know the price, but if they see that the house has 7 rooms, a postal code from a rich suburb, and marble countertops, the model can infer that it should be priced high.</p><p class="paragraph" style="text-align:left;">Those correlations are what AIs learn.</p><ul><li><p class="paragraph" style="text-align:left;">Large Language Models (LLMs) model how words follow one another. </p></li><li><p class="paragraph" style="text-align:left;">Neuralink algorithms model how brain activity leads to specific thought actions. </p></li><li><p class="paragraph" style="text-align:left;">Robotics algorithms model how certain instructions lead to certain bodily movements.</p></li></ul><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Traditionally, however, in the era of Machine Learning (pre-2010s), we lacked the necessary compute (physical hardware) to feed models the humongous amounts of data we feed to modern models today, which allows AIs to find all these patterns.</p><p class="paragraph" style="text-align:left;">Therefore, <b>we performed what’s known as ‘feature engineering’:</b> we would run statistical analyses on the data, identify which variables mattered, and provide the model with the answers in advance, making what the AI learned highly relevant to what the engineers wanted.</p><p class="paragraph" style="text-align:left;">Using our housing analogy again, we didn’t have trillions of homes to share data with the AI so that it can autonomously figure out “what mattered”, so we would previously run statistical analyses like regressions or correlations to identify the key variables (e.g., room count, postal code, previous prices, and whatever most likely determined the price of the home) and basically build that model around those assumptions.</p><p class="paragraph" style="text-align:left;">But in the 2010s, the Deep Learning revolution changed AI forever, taking us to where we are today.</p><h3 class="heading" style="text-align:left;">The Deep Learning Era</h3><p class="paragraph" style="text-align:left;">Although the principles of deep learning are older than basically every single reader of this newsletter or close, over the last 15 years or so, the world started to see more compute become available, mostly via GPUs (Graphical Processing Units), allowing researchers to tap into more and more data to feed to the AI models, which, as you may have guessed by now, is what all this is about: <b>feeding data to AIs that capture the patterns in that data and can make inferences about them</b>; as much data as one could get their hands on.</p><p class="paragraph" style="text-align:left;">This led to the ‘<b>Big Data</b>’ era, a term you&#39;ve likely heard before, <b>which opened the door to a particular type of AI algorithm called neural networks</b>, also known as Deep Learning.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><i>But why? </i>The short answer is the Universal Approximation Theorem, the idea that neural networks can approximate any continuous mathematical function to arbitrary accuracy. </p><p class="paragraph" style="text-align:left;">In layman’s terms, if a relationship between two variables can be represented by a continuous function, neural networks can learn it.</p><ul><li><p class="paragraph" style="text-align:left;">Brain activations → thoughts, <b>Neuralink</b></p></li><li><p class="paragraph" style="text-align:left;">Sequence of words → the next, <b>LLMs</b></p></li><li><p class="paragraph" style="text-align:left;">amino acid sequences → protein folds, <b>AlphaFold</b></p></li></ul><p class="paragraph" style="text-align:left;"><i>But what do we mean by mathematically represented?</i> By that, I mean the AI can perform mathematical computations to predict a possible answer, <b>and that answer can be mathematically measured to tell us how good or bad it was</b>, and thus used as a learning signal.</p><p class="paragraph" style="text-align:left;">LLMs are a great example of this, but it applies to every single AI model on planet Earth. The LLM sees the sequence <i>“What year was Einstein born?”</i> The ground truth answer is 1879, which the LLM does not know.</p><p class="paragraph" style="text-align:left;">The LLM then performs a series of computations, resulting in a ranking of possible answers by likelihood, assigning the year 1879 a probability of 30%. Of course, the correct answer would have been to give that year a probability of 100%, so we have a mathematical representation of the mistake: <b>the model was off by 70%.</b></p><p class="paragraph" style="text-align:left;">That is what we mean by mathematically represented: you can measure how good or bad a prediction was using maths. And if that loss is measurable, you can train the AI to minimize that loss; <b>it’s mathematically guaranteed</b>.</p><p class="paragraph" style="text-align:left;">And there you go, you know understand what ‘AI training’ actually means: the AIs make predictions, those predictions are measured by how good or bad they were, giving us a loss, and that loss is used as a learning signal (i.e., the AI knows what direction it needs to go in order to reduce the loss).</p><p class="paragraph" style="text-align:left;">In our case, the next time, the model will assign a higher probability to ‘1879’ and, across trillions upon trillions of predictions, the LLM learns to assign high probability to the correct words.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Importantly,<b> the ability to measure how good or bad a prediction was is called ‘verifiableness’.</b> This is important because if you understand this concept, you immediately know where AIs are expected to be good and where they are not.</p><p class="paragraph" style="text-align:left;">For example, look at these two sentences:</p><ul><li><p class="paragraph" style="text-align:left;"><i>“If I add 5 plus 6, the answer is 11.”</i></p></li><li><p class="paragraph" style="text-align:left;"><i>“And then, the boy screamed uncontrollably, for he did not understand what was going on.”</i></p></li></ul><p class="paragraph" style="text-align:left;">Both could be responses coming from an LLM. <i>But how verifiable are they? </i>The former is clear; the LLM output was accurate because the answer to 5 plus 6 is indeed 11. That is a sign that the LLM is improving in maths. If the model had answered 10, we also know the model is wrong. That is a sign that the LLM needs more work.</p><p class="paragraph" style="text-align:left;">But the second one is different. If you’re trying to make the LLM better at writing or storytelling, <i>how good or bad is the latter sequence toward that goal?</i></p><p class="paragraph" style="text-align:left;">Yes, we see good grammar, which is a factor in good writing. <i>What makes writing great? </i>At best, that’s subjective. That is not a good example of a verifiable data point;<b> we really don’t know, mathematically speaking, how good that output was.</b></p><p class="paragraph" style="text-align:left;"><b>This is the verifiability problem</b>, and it’s crucial because it’s a hard constraint for AIs: they can only learn what is verifiable.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">If you understand this, you know how to protect your future because, to me, verifiability is the only obvious answer to the question: <i>“How do I protect my job from AI?”</i></p><p class="paragraph" style="text-align:left;">I can’t promise you that mathematicians and software developers, both in highly verifiable domains, <b>won’t see meaningful changes in their work over the next few years</b>, but I am very comfortable telling you that an artist’s job won’t change much.</p><p class="paragraph" style="text-align:left;">Of course, AI does impact artists in meaningful ways, especially on the lower bound of work; you’re no longer as capable of selling some type of art that AIs can more or less copy, <b>but high-end work, unique and original, is not something AIs can do anytime soon because they don’t have a “mathematical hill to climb.”</b></p><p class="paragraph" style="text-align:left;">Other examples include sales: <i>What makes a good saleswoman?</i> You can sense when someone is great at sales, but please try to represent their sales skills using maths. Good luck.</p><p class="paragraph" style="text-align:left;">AIs do have a way to reasonably address some of the issues with the non-verifiable nature of certain skills, which leads us to the two types of learning mechanisms.</p><h3 class="heading" style="text-align:left;">Imitation and Reinforcement</h3><p class="paragraph" style="text-align:left;">In AI, there are two ways you train models: <b>imitation</b> and <b>reinforcement</b>.</p><p class="paragraph" style="text-align:left;">Imitation refers to the technique of exposing an AI to an absurd amount of data and having it replicate it to the letter. But by doing this and adding inductive biases that induce compression (I won’t get into this), <b>models learn to imitate us</b>.</p><p class="paragraph" style="text-align:left;">This is how ChatGPT learns not only the language but also how to speak back to you.<b> It does so by imitating humanity’s digital corpus to the dot</b>. This is how AIs learn to imitate Shakespeare, or cite Martin Luther King’s Lincoln Memorial speech, <i>“I Have a Dream,”</i> by memory.</p><p class="paragraph" style="text-align:left;">However, as it has been trained over the entire Internet, the outcome of this is like a mediocre “human,” the average of humanity’s digital content, <b>an AI model that becomes the embodiment of the average word on the Internet</b>.</p><p class="paragraph" style="text-align:left;">But current top models include an additional step called reinforcement that changes things. Here, instead of giving them a lot of data to copy,<b> we now give them goals and tasks to solve and let them try… a lot.</b></p><p class="paragraph" style="text-align:left;">And when these AIs “stumble” upon a solution we can verify as correct, we reinforce that behavior so it’s more likely the AI will repeat it.</p><p class="paragraph" style="text-align:left;">Once the model reaches the solution, <b>the entire “reasoning” that got it there gets reinforced, making it more likely the model repeats it</b>.</p><p class="paragraph" style="text-align:left;">This introduces its own problems, but it really pushes AIs to reach new heights in those areas, as AIs learn reasoning patterns that apply well to maths or coding structures, leading to successful outcomes.</p><p class="paragraph" style="text-align:left;">The reason this is so effective is simple:<b> they become problem-solvers at scale</b>, and when you can verify whether what they are doing is correct, you can verify progress and thus make progress.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Therefore, this technique, known as Reinforcement Learning (RL), is what has turned AIs into tools that are actually useful <a class="link" href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-in-2026" target="_blank" rel="noopener noreferrer nofollow">and even discover new maths</a>… <b>but only in those areas that can be verified.</b></p><p class="paragraph" style="text-align:left;">Therefore, most humans using AIs to generate content are behaving as if they had hired <b>a really inexpensive but very bad copywriter to handle their digital blueprint</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">For all these reasons, if you use AI to write, your content was, is, and will be mediocre for the foreseeable future because OpenAI and others genuinely don’t know yet how to train great AI writers, screenwriters, poets, film producers… <b>and with absolutely zero evidence that we might be nearing a way to solve that problem</b>.</p><p class="paragraph" style="text-align:left;">Ok. So, as of right now, the summary is that AI captures patterns, <b>and that neural networks</b>, thanks to the emergence of compute at scale, which lets us feed AIs trillions of data points,<b> have become the main choice for building AI systems</b>.</p><p class="paragraph" style="text-align:left;"><i>But which of these neural nets rules?</i> You probably heard the name: the Transformer.</p><h3 class="heading" style="text-align:left;">God’s Architecture: the Transformer</h3><p class="paragraph" style="text-align:left;">During the Renaissance, <b>there was a specific mathematical concept known as the Golden Ratio</b>, which many artists of the era believed would make their work aesthetically pleasing when proportioned to it. Centuries later, you would see painters like Salvador Dalí still doing the same. </p><p class="paragraph" style="text-align:left;">In AI, there’s a similar admiration for a concept known as attention, which sits at the heart of modern AI architectures, notably the Transformer.</p><p class="paragraph" style="text-align:left;">Additionally, in AI, we also have a thing called the ‘bitter lesson’. <span style="text-decoration:underline;"><a class="link" href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-in-2026" target="_blank" rel="noopener noreferrer nofollow">First described by Rich Sutton</a></span>, it is the “bitter” realization that, at the end of the day, the best humans can do in our AI aspirations is to… get out of the way.</p><p class="paragraph" style="text-align:left;">That is, it’s not about finding the most clever, complex heuristic we can find to train models. Instead, <b>the “best model architecture” is the one that allows the model to “see” more data and thus requires more compute, </b>which usually translates to very simple architectures that scale really well.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And the <i>Transformer</i>, the architecture we have today underneath most frontier AIs, is the perfect example of this.</p><h3 class="heading" style="text-align:left;"><b>It’s all about knowledge gathering.</b></h3><p class="paragraph" style="text-align:left;">The architecture that underpins products like ChatGPT is stupidly, almost insultingly, simple.</p><p class="paragraph" style="text-align:left;">At its core, a Transformer is just a concatenation of Transformer ‘blocks’ that perform several linear transformations to shape the model’s internal “belief” about which word comes next.</p><p class="paragraph" style="text-align:left;">But instead of indulging in esoteric descriptions like this one, which we can both pretend to understand but in fact don’t, I always like to explain these models more intuitively.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And once you see this, <b>it’s like you’ve magically understood AI algorithms.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-in-2026">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-ultimate-guide-for-ai-in-2026">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=c7fec5e8-7f12-4eae-81dc-b215315c5ad7&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The Week of the Open Model</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5d97758d-e2b7-4e23-ad8b-36d5e8eceb1f/Screen_Recording_2026-07-04_at_12.43.25.gif" length="712377" type="image/gif"/>
  <link>https://thewhitebox.beehiiv.com/p/the-week-of-the-open-model</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/the-week-of-the-open-model</guid>
  <pubDate>Sat, 04 Jul 2026 12:49:16 +0000</pubDate>
  <atom:published>2026-07-04T12:49:16Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back! </b>This week, we have a huge theme: <b>open models</b>, with several news items putting them at the center of the entire industry, including <b>new open models achieving SOTA results</b>, research that will make their adoption easier, and even some companies in the space <b>publicly attacking</b> Frontier Labs.</p><p class="paragraph" style="text-align:left;">We’ll also cover new research, like Meta’s mind reader, market data, and new products and models, including an incredibly realistic <b>AI-generated video</b> and a<b> chores robot</b> that may come to you as soon as this year.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>PorTAL, a New Way to Automatically Update Models</h2><p class="paragraph" style="text-align:left;">I’ve long sustained that the future of enterprise AI is companies training open models in their own data while giving an enthusiastic ‘goodbye’ to OpenAI and Anthropic, at least for the vast majority of enterprise use cases.</p><p class="paragraph" style="text-align:left;">And it seems several companies agree with me, including key players in the AI space like Palantir and Microsoft, and are becoming <i><b>increasingly</b></i> vocal about it (more on that below).</p><p class="paragraph" style="text-align:left;">The appeal of lower costs, better governance, and tighter security makes this a no-brainer once open models are reaching a level of capability that warrants adoption.</p><p class="paragraph" style="text-align:left;"><b>The strongest argument against this idea has always been obsolescence</b>, meaning<i> why would I spend a couple thousand dollars fine-tuning an open model if it’s going to be obsolete by next week?</i></p><p class="paragraph" style="text-align:left;">And while that argument is already not particularly strong when you realize that it’s okay to have legacy models for some tasks because these tasks do not need new levels of intelligence (e.g., I continue to use Gemini 3 Flash for a lot of my enterprise tasks despite this model being more obsolete than dinosaurs relative to the frontier, but it’s way cheaper), <a class="link" href="https://x.com/RampLabs/status/2072381992285647280?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">this new work by Ramp</a>, an expense management company that is starting to look more and more like an AI Lab these days, <b>proves how you can also ‘port’ your fine-tuning efforts from one model to the next.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><i>But how?</i> Leveraging one of the most fascinating pieces of research I’ve ever come across, Text-to-LoRA, <a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/ai-research-that-takes-your-hat-off-e2c8b9079394?sk=dd8e7fb7f24cf3f08adb12d61026db68&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">which I’ve talked about in the past</a>.</p><p class="paragraph" style="text-align:left;">Fine-tuning is the practice of taking a pre-trained model and retraining it for a specific task you’re interested in. This is great because the model becomes incredibly good at that task (at the expense of losing performance in other areas).</p><p class="paragraph" style="text-align:left;">The problem with fine-tuning is that you’re still retraining a large model, which can incur prohibitive costs, <b>especially considering that the industry moves so fast that new, superior models come out every week</b>, reducing the incentive even further because getting a return on that training run feels very complicated.</p><p class="paragraph" style="text-align:left;">An alternative is to use LoRAs, low-rank adapters. My previous link goes into detail, but the idea is that most tasks that a model has to learn are “low rank,” <b>meaning only a very small subset of the model’s global parameter count has to be modified to learn the task</b>. Hence, the idea is to only train a small portion of the model.</p><p class="paragraph" style="text-align:left;"><i>So what is LoRA training?</i> Simple: train a tiny external adapter and add it to the model whenever it’s working on that task. This is not only cheaper <b>but also leaves the open model untouched</b>, allowing a single model to work with potentially hundreds of adapters, depending on the task, <b>an architecture used in compute/memory constrained environments (e.g., Apple Intelligence)</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5b6ae679-2599-4159-9e22-9e06e2fe3b91/image.png?t=1782991221"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">You can think of LoRA adapters as giving a plumber a wrench for that task and only for that task. While you need to teach the plumber to use it, you don’t need to rewire the entire plumber’s brain; i<b>t just requires a small, quick learning process for using the wrench.</b></p><p class="paragraph" style="text-align:left;">Importantly, the wrench is only required for wrench stuff, so it doesn’t become a part of the plumber’s “being,” and they can simply drop it if it’s not required for the task.</p><p class="paragraph" style="text-align:left;">LoRA is incredibly effective and a standard for training. However, <b>it still requires a training run and can be expensive relative to the </b><b>base model&#39;s time-to-obsolescence</b> (i.e., your base model might become ‘dumb’ relative to what’s available in the market pretty quickly).</p><p class="paragraph" style="text-align:left;">For this, <b>Sakana AI had a great idea called text-to-LoRA, using an AI to generate the LoRA conditioned on text.</b></p><p class="paragraph" style="text-align:left;">For example, you can describe the task <i>“write emails using formal language”</i>, and this model automatically generates a LoRA adapter that, when added to a model, makes it write emails always in formal language.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Incredible,<i> right? </i>No training run required, fast iteration, and it works. <i>What’s not to like?</i></p><p class="paragraph" style="text-align:left;">Well, one problem remained: <b>Text-to-LoRA outputs are not model-agnostic</b> because they are trained alongside the chosen base model. In simple terms, they learn to generate adapters for a particular model.</p><p class="paragraph" style="text-align:left;">Here’s where Ramp’s PorTAL comes in. Cutting to the point, they’ve managed to create a model-agnostic Text-to-LoRA that can port adapters from one model to the next.</p><p class="paragraph" style="text-align:left;">For example, say you’ve trained (or even automatically generated) an email-filtering LoRA for your Qwen3.6 35B model… and Qwen4 arrives the next week and is considerably smarter, to the point that it’s worth the switch.</p><p class="paragraph" style="text-align:left;">Instead of ditching the previous base model and writing off the training, <b>you use PorTAL to transfer the LoRA to that new model</b>. There’s some training required, but it’s minimal as the largest portion of the PorTAL system is model-agnostic.</p><p class="paragraph" style="text-align:left;">The results are very promising, showing that the system recovers 98% of the per-task LoRA performance on an unseen base model. I know, that’s a lot of jargon in one sentence.</p><p class="paragraph" style="text-align:left;">In layman’s terms,<b> they prove that the PorTAL system can take an ‘unseen’ model and generate a LoRA with minimal training and cost</b>, matching the effort required to train a LoRA adapter for that model and task from scratch.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ef32839e-25f7-48f8-a871-54ae50938330/image.png?t=1782993409"/></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I have a strong belief that many in this industry would consider very bold: <b>much of Frontier Lab&#39;s revenue can be explained by customer unsophistication.</b></p><p class="paragraph" style="text-align:left;">Which is to say, most Anthropic/OpenAI enterprise customers use their models because they don’t yet know how to build robust open model pipelines and workflows.</p><p class="paragraph" style="text-align:left;">The math is already there; open models give you greater control, better cost management, and a fully sovereign AI stack. The problem is that enterprise leaders aren’t yet aware of this and aren&#39;t operationally capable either.</p><p class="paragraph" style="text-align:left;">But once frontier token prices force companies out of their bubble and into real AI engineering, I believe Anthropic and OpenAI will be in a world of pain.</p><p class="paragraph" style="text-align:left;">Besides liquidity, <b>it’s this sophistication curve that makes these labs so eager to go public as soon as possible</b>. Once the word is out, investors will find it much harder to underwrite today’s valuations on the basis of distant, increasingly uncertain revenue growth.</p><p class="paragraph" style="text-align:left;">The “own your AI stack” trend is gaining huge momentum at the worst time possible for these Labs.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>VIDEO</b></span><br>Crossing the Uncanny Valley for good</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/john_my07/status/2071977017474789557?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">After seeing this video</a>, <b>I finally accept that I can no longer distinguish AI videos from real videos</b>.</p><p class="paragraph" style="text-align:left;">It’s impossible, at least to me, <b>as models like SeeDance 2.0 have crossed the uncanny valley</b> (the feeling you get when you see something that looks <i>almost</i> human but not quite) into something that is simply indistinguishable.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The results are simply incredible. But I want to remark on something that is often ignored in these situations: <b>the incredibly beautiful prompt, also linked, that led to that video</b>. That is also art; this is also something complicated most people can’t achieve because most people can’t prompt this way.</p><p class="paragraph" style="text-align:left;"><b>What I’m implying is that there’s still something human in all of this</b>, there’s a craft here, and that the capacity to imagine and describe what needs to be generated is a skill like any other. After reading the entire prompt, <b>I can assure you I don’t have the vocabulary or the creativity to create something like that</b>, at least not today.</p><p class="paragraph" style="text-align:left;">My point is that AI does elevate what the average human can create. <b>But it further elevates the experts in the given domain</b>. People assume AI destroys leverage, <b>but I’m beginning to suspect it will simply give experts more leverage than ever.</b></p><p class="paragraph" style="text-align:left;">It’s also interesting to see how AI videos are much more indistinguishable from real videos than AI-generated writing is. One possible explanation is that AI-generated content is everywhere, <b>so we have been exposed to it so often that we can recognize when AI is involved.</b></p><p class="paragraph" style="text-align:left;">It could be that AI videos, once they go truly mainstream, will suffer from the same issues, as AIs fail to generate meaningfully different results beyond their average output. Who knows.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Fable Came Back… Nerfed</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4fc76245-da8e-4c5b-8ce2-c443d7a712cb/image.png?t=1783108049"/></div><p class="paragraph" style="text-align:left;">As reported by BleepingComputer and other sources,<b> users say Anthropic’s relaunched Claude Fable 5 feels “nerfed” after its return</b><b> with stricter safety controls</b>. The complaint centers on coding and debugging tasks, where some users report more refusals, fallbacks, and weaker results.</p><p class="paragraph" style="text-align:left;">We now even have benchmark confirmation, with percentage changes that are quite dramatic.</p><p class="paragraph" style="text-align:left;">Anthropic says the main change is an added <b>safety classifier, which is, of course, much more stringent than the one that the USG and Amazon jailbroke</b> to meet the USG’s requisites.</p><p class="paragraph" style="text-align:left;">When Fable 5 flags a request, <b>it may be blocked or routed to Claude Opus 4.8 instead</b>. Anthropic says this was added after concerns about cybersecurity misuse and blocks the reported bypass behavior in over 99% of cases.</p><p class="paragraph" style="text-align:left;">Independent benchmark claims are mixed. ModemGuides reports that one BridgeBench rerun showed sharp drops after Fable 5’s return, <b>including debugging falling from 86.2 to 25.9</b>, but notes this may reflect classifier blocks or fallback behavior rather than weaker underlying models.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is how companies lose the mandate of heaven. At some point, we’ll have to open the can of worms of classifiers being legal. <i>How is me getting charged for something (Opus 4.8) I did not intend to purchase because I was looking to use Fable 5?</i></p><p class="paragraph" style="text-align:left;">I understand this is obviously in the terms of service, but it’s an extremely arbitrary and unclear event that should raise questions about whether the user is being scammed.<i> If enough users feel scammed, is that enough evidence of a scam? </i>I have literally no idea; I’m just throwing this out because I, for one, simply reject using these systems if I can’t foresee what model I will be served.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>Surprise! AI Doesn’t Cause Layoffs</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3f8c2f11-a8da-4d05-86a3-96ccd2950820/image.png?t=1783156183"/></div><p class="paragraph" style="text-align:left;">A new, <a class="link" href="https://ramp.com/data/ai-jobs-impact?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">interesting piece of research</a> by Ramp and Revelio Labs, echoed by the Financial Times, has found a striking correlation that <b>discredits </b><b>the idea that AI doesn’t cause massive layoffs</b>.</p><p class="paragraph" style="text-align:left;">If anything, correlations show the opposite.</p><p class="paragraph" style="text-align:left;">According to their analysis, <b>companies adopting AI have a 10.2% average increase in headcount two years in</b>, versus those with low AI adoption metrics showing basically no growth.</p><p class="paragraph" style="text-align:left;">But before the statisticians in the crowd defame me for confusing correlation with causation, I am not, and neither are these researchers.</p><p class="paragraph" style="text-align:left;">Based on the information I’ve shared with you so far, <b>you might interject that this does not prove causality and that other variables may be at play, explaining the difference</b>.</p><p class="paragraph" style="text-align:left;">For example, it’s reasonable to assume that high-adoption companies also tend to be enterprises with higher growth, better financing, larger size, and other factors that could explain increased headcount.</p><p class="paragraph" style="text-align:left;"><b>The researchers acknowledge this and apply a difference-in-differences method</b>. In simple terms, those other variables we’re discussing are also commonly found in companies that were late adopters, so what they’ve done is compare to this group.</p><p class="paragraph" style="text-align:left;">In other words, while comparing adopters vs non-adopters is a bad analysis to isolate the AI effect for the reasons stated above,<b> it’s much more reasonable to compare early adopters to late adopters</b>, because both share many of these confounding attributes, and the impact on headcount during the time difference can be used to isolate the effect of AI.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This way, while not categorically claiming causality, as that would require purely randomized trials (i.e., giving AI at random to some companies and not to others, and seeing if AI causes a difference), it’s a much more realistic comparison, essentially making this “AI effect” be measured as follows:</p><p class="paragraph" style="text-align:left;"><i>AI effect ≈ employment growth after AI adoption among adopters − employment growth over the same period among similar not-yet-adopters</i></p><p class="paragraph" style="text-align:left;">And the results are pretty good, showing a consistent pattern:<b> early adopters’ headcount growth outpaced late adopters&#39; considerably</b> during the period when the former were using AI while the latter weren’t.</p><p class="paragraph" style="text-align:left;">The data also shows some counterintuitive results that may surprise you. For instance, it’s reasonable to assume that one reason one company may be hiring faster than the other is its size. Smaller companies, like start-ups, usually grow faster and are more eager to hire.</p><p class="paragraph" style="text-align:left;">However, the average headcount of the low-adopter group is much higher than that of the “never-adopters,” <b>and it shows higher headcount growth</b>. Yes, high adopters are generally smaller, but arguing that this is all a ‘large vs small’ comparison makes no sense.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/bf9326fe-abb6-47a7-a5eb-60f42c61df48/image.png?t=1783157197"/><div class="image__source"><span class="image__source_text"><p>Source: Ramp</p></span></div></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>BRAIN DECODERS</b></span><br>Having AI Decode Thoughts</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5d97758d-e2b7-4e23-ad8b-36d5e8eceb1f/Screen_Recording_2026-07-04_at_12.43.25.gif?t=1783161852"/></div><p class="paragraph" style="text-align:left;">Meta has introduced <a class="link" href="https://facebookresearch.github.io/brain2qwerty/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">Brain2Qwerty v2</a>, <b>a non-invasive AI system designed to decode typed sentences from brain activity</b>.</p><p class="paragraph" style="text-align:left;">The model was trained on roughly 22,000 sentences from nine volunteers, each recorded for about 10 hours while wearing a magnetoencephalography, or MEG, device and actively typing on a keyboard. Meta reports that the system reaches 61% word accuracy on average, rising to 78% for its best participant.</p><p class="paragraph" style="text-align:left;">The key point is that Brain2Qwerty is not a general thought-to-text model.<b> It does not read arbitrary thoughts and convert them into language</b>. In the experiment, participants were shown sentences, briefly memorized them, and then typed them on a QWERTY keyboard while their brain activity was recorded. The system learned to reconstruct the sentence from the brain signals associated with that typing task.</p><p class="paragraph" style="text-align:left;">Therefore,<b> the training signal is obtained by comparing the system&#39;s predictions with the participant&#39;s actual typing</b>. If the model assigns a low probability to the correct sequence of letters or words, the error is used to update the model.</p><p class="paragraph" style="text-align:left;">Over many examples, <b>it learns which patterns in the brain recording tend to correspond to particular typed characters</b>, word boundaries, and sentence structures, making the system closer to “brain-assisted typing reconstruction” than pure “mind reading.”</p><p class="paragraph" style="text-align:left;">This is similar to what Neuralink is doing, with the difference that Neuralink’s chips are invasive in order to capture a less noisy signal that doesn’t lose strength by crossing the skull and the scalp.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I’ve talked about Meta’s weird desire to read human thoughts until you realize it makes incredible sense to them.</p><p class="paragraph" style="text-align:left;">Imagine they could decode what you want based on how you interact with their platforms, like a reverse engineering process to what we have described; if they can decode how your brain is activated while using Instagram, <b>they can then predict your needs</b> by knowing what you like, what you dislike, and everything in between; <b>the ultimate ad-targeting platform</b>.</p><p class="paragraph" style="text-align:left;">Sounds great, <i>right?… Right?</i></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>OPEN SOURCE</b></span><br>A War on Private Models?</h2><p class="paragraph" style="text-align:left;">In the last few days, several prominent companies upstream and downstream of the closed AI Labs, Palantir, Microsoft, and TogetherAI, as well as rivals like Mistral <a class="link" href="https://x.com/arthurmensch/status/2073157738276749354?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">through its CEO</a>, <b>have voiced very strong opinions, especially Palantir, in favor of open models</b>, or, more clearly: <b>companies should not outsource their operational learning loop to a frontier-lab token API</b>.</p><p class="paragraph" style="text-align:left;">Microsoft is doing so with its <a class="link" href="https://www.microsoft.com/en-us/frontier-company?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">Frontier Company/Frontier Tuning push</a>, which focuses on helping enterprises build customized AI systems around their own data, workflows, and business goals, often within the customer’s environment and with greater model flexibility. Reuters framed this as Microsoft helping companies move away from dependence on a single AI provider, such as OpenAI or Anthropic.</p><p class="paragraph" style="text-align:left;">Palantir is partnering with NVIDIA to onboard its Nemotron models to its Ontology platform and is being less orthodox about its opinions on Anthropic and OpenAI. </p><p class="paragraph" style="text-align:left;">Karp has criticized “tokenmaxxing” and warned that companies risk giving away their IP, alpha, and operational knowledge to external LLM providers. His position is that the enterprise’s data, ontology, permissions, and workflow logic should remain under the company’s control, with models treated as replaceable components.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/PalantirTech/status/2072114267776491695?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">And from a more official channel</a>, the company itself published a statement in favor of AI sovereignty,<b> clearly stating that controlling the models&#39; weights means controlling your fate</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is, of course, a very self-serving narrative for all these players mentioned above, but they nail it in their own way. They are correct, and you should be scared straight from trusting your IP to these Labs.</p><p class="paragraph" style="text-align:left;">But all I can say is that I feel vindicated. It’s finally happening. The transition to open models in the enterprise looks unstoppable, something I’ve been calling for years, even before Thinking Machine Labs came out with their RLaaS service, by that time, I had already made up my mind this was the future.</p><p class="paragraph" style="text-align:left;">Not because I’m a genius, but simply because I paid attention to AI’s history, and except for the few recent years, the decades-long story of AI has always been about open, deep models. In other words, AI research has always been open, and the outcomes have always been task-specific; <b>it’s like we’re returning to 2016 AI</b>.</p><p class="paragraph" style="text-align:left;">Yes, foundation models like the ones we have today are good at various tasks, but great at none, the opposite of what enterprises need; they don’t care that their customer support agent is also an expert cake baker, but they need it to be the best customer support agent possible.</p><p class="paragraph" style="text-align:left;">Besides, <b>companies are realizing the extreme stupidity of outsourcing their entire AI stack to third-party companies that not only access their IP</b> (Anthropic literally publishes research classifying how users use Claude) but actively build downstream competitors using their IP, as happened to Figma with Claude Design and to pharma companies with Anthropic’s new drug-discovery initiative.</p><p class="paragraph" style="text-align:left;">In the case of Figma, Anthropic’s Chief Product Officer was a board member and resigned only three days before Claude Design launched. Do with this information what you wish, but I have a very clear idea of how I would feel if I were Figma.</p><p class="paragraph" style="text-align:left;">And to cut Anthropic some slack, <b>OpenAI and Google do the same thing</b>.</p><p class="paragraph" style="text-align:left;">Luckily, open models are finally good enough (see the first news in the product section below),<b> so you can train fully sovereign solutions on your data and achieve state-of-the-art performance at more than 10 times</b> (or even up to 50 times) lower cost.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>CHIPS</b></span><br>Anthropic Joins the Chip Mania</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://techcrunch.com/2026/07/02/anthropic-is-discussing-a-new-custom-chip-with-samsung/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">As published by TechCrunch</a>, Anthropic is reportedly discussing a custom AI chip with Samsung.</p><p class="paragraph" style="text-align:left;">The talks are still early. According to TechCrunch, Anthropic has not yet decided exactly what the chip would be used for, how powerful it would be, or how it would fit into its servers.</p><p class="paragraph" style="text-align:left;">Anthropic also emphasized that its compute strategy will continue to rely on a diversified hardware stack that includes Google, Amazon, and Nvidia.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Better late than never. One intriguing thing for me here with OpenAI/Anthropic hardware initiatives is that they make sense on a cost basis; progressively lowering costs/token, <b>but they could hurt them badly in accounting terms</b>.</p><p class="paragraph" style="text-align:left;">Right now, they are mostly avoiding depreciation costs and CAPEX, and simply renting compute from their own investors (Hyperscalers), who are gladly offering low rental rates because they can recognize that usage as AI cloud revenue and RPO numbers (north stars for many Hyperscaler investors).</p><p class="paragraph" style="text-align:left;">This is what has allowed Anthropic to claim “adjusted profitability” (excluding stock-based compensation) for this quarter. But this much-celebrated milestone hides a problem: Anthropic can only claim this because there’s a fool behind them paying the real bills and suffering the massive cash flows (i.e., Amazon and Google).</p><p class="paragraph" style="text-align:left;">But the moment the GPUs are yours, your company is no longer margin-focused and suddenly much more free-cash-flow-focused because you have significant cash outflows.</p><p class="paragraph" style="text-align:left;">If you judge Anthropic or OpenAI by free cash flow rather than margins, the picture changes completely: in AI, margins are acceptable, <b>but cash flows are horrendous</b>, so these Labs might end up looking worse to investors the more they move upstream into owning the infrastructure.</p><h2 class="heading" style="text-align:left;"><span style="color:#FF5632;font-size:0.8rem;"><b>STOCK MARKET</b></span><br>Is AI Overheating?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/465ad8e1-911b-4ad3-b6f4-afb69a39c339/image.png?t=1782982262"/></div><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">On June 29 and 30, 2026, more than 60 companies listed on China’s A-share markets released announcements on abnormal stock price fluctuations and risk warnings.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">The majority of these companies operate in the semiconductor industry and its supply chain, including chip design, manufacturing equipment, and materials. The announcements were filed with the Shanghai and Shenzhen Stock Exchanges.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">They were triggered by sharp increases in share prices in recent trading sessions. </span><span style="background-color:transparent;"><b>Many companies reported cumulative gains of 60 percent or more</b></span><span style="background-color:transparent;"> over periods such as the prior 20 trading days, with some noting even larger rises over 10 or 30 days.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">These gains exceeded the performance of broader market indices, including the STAR Market Composite and STAR 50 indices. In the filings, </span><span style="background-color:transparent;"><b>rolling price-to-earnings ratios</b></span><span style="background-color:transparent;">—the valuation multiple relative to earnings; the higher, the more highly valued a stock is—</span><span style="background-color:transparent;"><b>were frequently cited as significantly higher than industry averages</b></span><span style="background-color:transparent;"> for the computer, communications, and electronic equipment manufacturing sector.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">Examples include GigaDevice (stock code 603986), which noted that memory chip prices were at historical highs and could experience a considerable pullback, and Jiangsu Aisen Semiconductor Materials, which reported a 61 percent gain over 20 trading days alongside a rolling P/E ratio of 233.86 times.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;">These are absolutely insane numbers. For reference, </span><span style="background-color:transparent;"><b>the companies in the S&P 500 have average trailing and forward P/Es of 20x and 32x</b></span><span style="background-color:transparent;">, respectively. Even a company as richly valued as Palantir has a PE of 135, way smaller than some of the numbers we’re seeing in China.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:transparent;"><i>But what on Earth is going on?</i></span></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Chinese listed companies (especially in the semiconductor supply chain) are required to issue these when their stock prices move abnormally (large gains over a short period).</p><p class="paragraph" style="text-align:left;">One or two issuing such risk disclosures doesn’t say much,<b> but when 60 do so in the last two days, something’s off.</b></p><p class="paragraph" style="text-align:left;">And when people looked more closely, they found that all these companies were related to the AI and semiconductor industries, highlighting incredible exuberance.</p><p class="paragraph" style="text-align:left;"><b>It’s becoming quite clear that the AI markets are overheating</b>, and many investors will see this as an opportunity to sell. <i>Will they be right?</i> Who knows.</p><p class="paragraph" style="text-align:left;">And talking about sell-offs…</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MEMORY</b></span><br>AI is Having a Horrendous Start of July Thanks to Meta</h2><p class="paragraph" style="text-align:left;">Meta’s decision to eventually become a neocloud (a company that rents compute to others, like the other three Hyperscalers or CoreWeave) sent the stock up but crashed the entire market in return, <a class="link" href="https://vector.news/meta-sells-excess-compute-it-doesnt-have/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">despite Meta not being able to do so today</a>.</p><p class="paragraph" style="text-align:left;"><b>The reported plan is for Meta to develop “Meta Compute,”</b> a cloud-like business that would sell access to its data centers and AI models. Reuters says this could include access to hosted models, similar to AWS Bedrock, and possibly raw AI compute rental, similar to what neoclouds like CoreWeave sell. Meta declined to comment, and Reuters says it could not independently verify Bloomberg’s report.</p><p class="paragraph" style="text-align:left;">The reason Vector’s headline says “excess compute it doesn’t have” is that “excess compute” normally means you built more capacity than you need internally. But Meta is simultaneously raising/maintaining enormous AI infrastructure spending, reportedly up to about $145 billion this year, which suggests the opposite: <b>it needs far more GPUs/data centers/power for its own AI ambitions.</b></p><p class="paragraph" style="text-align:left;">It’s interesting that a company with annual AI spending greater than Germany&#39;s defense spending defines the trigger for “Meta Compute” as having “excess compute.”</p><p class="paragraph" style="text-align:left;">Meta’s investors loved the idea, <b>but markets hated it because it once again unearthed the ongoing fear that the industry might be overbuilding</b>. However, to be fair to Meta, weakness has already been present, and markets have been in a purgatory (not up, not down) for several weeks now.</p><p class="paragraph" style="text-align:left;">Nonetheless, the main selling pressure came from chipmakers. Reuters reported that the Philadelphia semiconductor index fell <b>6.3% on July 1</b>, dragging the Nasdaq down <b>0.66%</b> and the S&P 500 down <b>0.22%</b>. MarketWatch said the SOX index had already dropped <b>3.4%</b> earlier in the session,<b> after rising 87.8% in Q2, its strongest quarter on record.</b></p><p class="paragraph" style="text-align:left;">The selloff also spread outside the US. The Economic Times reported that Samsung Electronics and SK Hynix fell as much as <b>14.5%</b> on Thursday, while South Korea’s Kospi dropped <b>8.2%</b>, with investors reacting to the same AI-capacity concerns.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">In case you’re curious, knowing I’m an investor in both Hynix and Samsung and have heavy exposure to many other AI players, I haven’t sold.</p><p class="paragraph" style="text-align:left;"><b>But I do understand people’s fears and don’t blame them for doing so</b>, <a class="link" href="https://thewhitebox.beehiiv.com/p/we-need-to-talk-about-this?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">as I myself explained to you recently</a> how AI financing is concerning at best, really, really worrying at worst, especially when we factor in the growing presence of debt.</p><p class="paragraph" style="text-align:left;">One thing is for a fully cash-and-equity-driven investment bubble to bust. When that happens, damage is contained. But once debt enters the picture, not only does risk spread across many more participants, <b>but it also makes it almost impossible to know the extent of the exposure</b> (e.g., shadow borrowing).</p><p class="paragraph" style="text-align:left;">It’s fascinating, <a class="link" href="https://thewhitebox.beehiiv.com/p/we-need-to-talk-about-this?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">but as we saw</a>, a data center project that fails to pay back a loan could prevent you from getting the money from your annuity your life insurer owes you because that money is now stuck in that failed project.</p><p class="paragraph" style="text-align:left;">Those types of relationships are not remotely discussed in this industry, yet they are very real.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>POLITICS</b></span><br>The USG, New OpenAI Shareholder?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.ft.com/content/7c803eab-8e80-4431-9a87-e943bf00e00b?syn-25a6b1a6=1&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">The Financial Times has published an article</a> explaining that <b>OpenAI has discussed giving the US government a 5% equity stake</b>, <b>worth roughly $42.6 billion, </b>as part of talks to address political concerns around the AI industry and to share AI-related financial upside with the public.</p><p class="paragraph" style="text-align:left;">The idea is described as early-stage and “conceptual.” Reports say it could require Congressional approval, <b>and it may be part of a broader model in which the government would hold stakes in major US AI developers</b>, potentially through a public wealth fund similar in concept to Alaska’s oil-funded dividend model.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">It isn’t clear whether OpenAI expects the Administration to pour $42 billion in or if they are giving that stake for free. Either way, the intention is strikingly obvious to me: <b>make OpenAI’s survival a state matter.</b></p><p class="paragraph" style="text-align:left;">From the creators of <i>“Too Big to Fall”</i> comes <i>“I’m an owner now; it can’t fail.” And while </i>I understand why people like Bernie Sanders or Trump, among others, support the idea of the USG having a stake in these companies, it might end up being more of a bailout.</p><p class="paragraph" style="text-align:left;">If OpenAI is worth 5 trillion one day, this will be seen as a massive success for the American people. But if things go south and OpenAI’s liquidity dries out, <b>it will look more like a bailout than a national security investment</b>.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RLaaS</b></span><br>Thinking Machines 🤝 Bridgewater</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f8ede8d2-5b4d-412e-b891-dbca26eb5eec/image.png?t=1783000662"/></div><p class="paragraph" style="text-align:left;">Every day, a new example of how <b>RLaaS</b> (Reinforcement Learning as a Service, the idea of offering companies an easy way to train models) <b>is going to take the world of AI by storm, emerges</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This time, it’s a partnership between Thinking Machines Labs (TML), the flash US AI Lab packed with ex-OpenAI and other top lab researchers that has bet its entire existence on this idea, <a class="link" href="https://thinkingmachines.ai/news/learning-to-replicate-expert-judgment-in-financial-tasks/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">and Bridgewater Associates</a>, one of the world&#39;s largest hedge funds.</p><p class="paragraph" style="text-align:left;">Using TML’s Tinker API, a product that lets you train open models with your own data without having to deal with the complexities of training and infrastructure, Bridgewater turns an open model (Chinese Qwen 3 235B, which is nowhere near the frontier) trained for <b>filtering and processing financial documents into one that surfaces information relevant to investment decisions</b>, into a state-of-the-art model, <b>ahead of any commercially available model on the planet</b>, as seen in the thumbnail.</p><p class="paragraph" style="text-align:left;">Importantly, <b>they also confirm that this model also beats heavily prompt-engineered frontier models</b>, meaning they also tried to squeeze as much performance from AIs like Opus 4.8 or GPT-5-5 in an effort to see how far they can get with frontier models (which you can’t fine-tune), managing to reach high seventies results but still failing to beat the fine-tuned model, which beats them handsomely.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aa181df5-daef-4dfb-8bc7-b7570ef6a876/image.png?t=1783001194"/></div><p class="paragraph" style="text-align:left;">Also, as you can see in the thumbnail, it is not only a raw-score victory; the fine-tuned open model is between 12 and almost 20 times cheaper than the top OpenAI/Anthropic models.</p><p class="paragraph" style="text-align:left;">Better performance at multiple times lower cost. Feels too good to be true, but that is why I’m so confident this is the only way to true enterprise adoption.</p><p class="paragraph" style="text-align:left;">To make things even worse, they show how frontier models are getting more expensive than better, <b>with </b><span style="background-color:rgb(255, 255, 255);"><b>GPT 5.4 costing 43% more than 5.2 but being only marginally more accurate</b></span><span style="background-color:rgb(255, 255, 255);">.</span></p><p class="paragraph" style="text-align:left;"><span style="background-color:rgb(255, 255, 255);"><i>And why is fine-tuning so superior to prompting?</i></span><span style="background-color:rgb(255, 255, 255);"> Well, as they explain: </span><span style="background-color:rgb(255, 255, 255);"><i>“Rather than contorting the expert’s intuition into a static prompt, the training process lets the model develop its own judgment.”</i></span></p><p class="paragraph" style="text-align:left;">Following a pretty standard training pipeline, which of course includes on-policy distillation, <a class="link" href="https://thewhitebox.beehiiv.com/p/the-state-of-ai-in-two-breakthroughs?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">which I talked about last week</a> as key to AI progress these days, <b>they convert a bad, open model into a frontier model</b> (for that task).</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">You may see these results as unimpressive (“well, it’s only one task”), but here’s the thing: as I said above, <b>enterprises don’t care about generalization to many tasks</b>; they need the model to be great at that one task and will simply use another model for another task if required.</p><p class="paragraph" style="text-align:left;">Now, while still surfing the foundational model wave, we are taking those models and making them great at single tasks again. This seems unavoidable to me, and the primary reason why Anthropic freaks out so much about open models.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ROBOTICS</b></span><br>Babe, wake up! A new chore robot dropped</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/be48b646-bfb9-4d5d-ba52-c7ec6dd4da63/image.png?t=1783002640"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/weaverobotics/status/2072362538671706314?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://x.com/weaverobotics/status/2072362538671706314?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-week-of-the-open-model" target="_blank" rel="noopener noreferrer nofollow">Weave Robotics posted</a>, the company is launching <b>Isaac 1</b>, its home robot,<b> with deliveries scheduled to begin this fall</b>.</p><p class="paragraph" style="text-align:left;">This appears to mark a shift from <b>Isaac 0</b>, a stationary laundry-folding robot, to <b>Isaac 1</b>, described as a more mobile home robot for a broader range of household tasks. A3’s earlier coverage said Isaac 0 was not the full mobile platform shown in Weave’s broader product vision, but a pared-down first deployment focused on laundry folding.</p><p class="paragraph" style="text-align:left;">A related LinkedIn post from Weave co-founder Evan Wineland said the company had unveiled Isaac 1 at an event about two weeks earlier, and that Isaac 1’s design was recognized with an Industrial Design award by San Francisco Design Week.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The key thing here is that<b> robotics is finally leaving the demo phase and actively looking to release products as soon as this year</b>; it says orders are open and deliveries begin in fall 2026.</p><p class="paragraph" style="text-align:left;">As for my personal opinion, <b>I still struggle with the vision of having a humanoid in my home.</b> Also, the helpful things it can really do for me aren&#39;t many; I don’t need a robot to make my bed.</p><p class="paragraph" style="text-align:left;">People (pure coincidence, mostly investors) think this is going to change the world.<i> But will it? </i>I’m not sure.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">This week has been a huge week for open models, which are now becoming top-of-mind options for enterprises. <b>This is</b>, however, <b>very, very bad for AI Lab investors</b>, who might have considerably overestimated the size of the enterprise market for frontier tokens.</p><p class="paragraph" style="text-align:left;">The markets aren’t looking any better, with extreme volatility and fear being the norm. As I described, <b>markets have been in a ‘purgatory’ of sorts for several weeks</b>, so it’s unclear whether the break that will occur will go upwards… or down.</p><p class="paragraph" style="text-align:left;">And to end on a positive note, <b>we’re starting to see evidence that AI could actually boost job growth</b>. This wouldn’t be a first, as every major technological disruption has created more jobs than it destroyed.</p><p class="paragraph" style="text-align:left;">For years, we thought AI would be different. But it might turn out that <b>AI was just another technology deeply following the same trend and creating a better world</b>, not one full of despair and joblessness, as some in Silicon Valley love to fantasize about.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=1d0bf916-9031-49eb-a2bf-33da8b42871c&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The State of AI in Two Breakthroughs</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b0fb9795-efbd-4c6c-902f-4f4644b70fc9/image.png" length="1849834" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/the-state-of-ai-in-two-breakthroughs</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/the-state-of-ai-in-two-breakthroughs</guid>
  <pubDate>Mon, 29 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-29T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>The State of AI in Two Breakthroughs</h2><p class="paragraph" style="text-align:left;">After a very finance-oriented previous Leaders newsletter, today, we’re putting ourselves at the bleeding edge of this technology.</p><p class="paragraph" style="text-align:left;">For that, we will cover research that answers two questions:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><i>How does state-of-the-art frontier research look today? </i><b><i>What are the Frontier Labs obsessed with right now?</i></b></p></li><li><p class="paragraph" style="text-align:left;"><i>How does the “efficient frontier” look today? </i><b><i>What is the biggest recent breakthrough in doing more with less?</i></b></p></li></ol><p class="paragraph" style="text-align:left;">By the end of this read, you’ll have a better understanding of this technology than you ever thought you would have. If AI is about to transform our future, understanding it better than anyone else around us feels like a key competitive advantage.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">What the Frontier is Obsessed About</h2><p class="paragraph" style="text-align:left;">In frontier research today, the hottest topic right now is <b>on-policy self-distillation</b>, or <b>OPSD</b>, the four hottest words in San Francisco right now.</p><p class="paragraph" style="text-align:left;">I know, it sounds complicated because it’s named to sound like it, but it’s actually not at all if you break it into first principles, which is what this newsletter is all about.</p><p class="paragraph" style="text-align:left;">And to break down frontier research into first principles, we need to start humbly.</p><h3 class="heading" style="text-align:left;">How do AIs learn?</h3><p class="paragraph" style="text-align:left;">Simplified much, all AI models follow the exact same process to learn: they make a prediction about what we want to learn, we measure that prediction against the ground truth (what it should have predicted)<b>, and we use the difference between the two predictions as the learning signal.</b></p><p class="paragraph" style="text-align:left;">If the AI should have predicted 10 and it predicted 1,000, the difference is 900, a large error. If the next prediction is 500, the difference is lower, prompting the model to continue heading in that direction.</p><p class="paragraph" style="text-align:left;">However, AIs never predict single numbers (or rarely); <b>they almost always predict probability distributions</b>. In other words, they don’t say <i>“the answer is 10,”</i> but <i>“I believe the answer is 10 with 90% certainty, but it could be 12 with a certainty of 10%.”</i></p><p class="paragraph" style="text-align:left;">This means the AI has responded with two numbers, not one, and associated probabilities, creating a distribution of possible responses.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Notice the word “believe,” because I used it on purpose. By outputting a distribution, <b>the model expresses its degree of uncertainty about its prediction</b>; we are still forcing it to make a choice, but we allow it to indicate how “convinced” it is about the prediction.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/be141fe5-e2e5-419c-adb9-1e18034b9804/image.png?t=1782720612"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://medium.com/@lmpo/mastering-llms-a-guide-to-decoding-algorithms-c90a48fd167b?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-state-of-ai-in-two-breakthroughs" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;">Distributions are great because they prevent “response collapse”: if you force the model to always choose one and only one option, <b>creativity goes out the window because the model learns only the most likely option</b>, not the others.</p><p class="paragraph" style="text-align:left;">That’s the difference between a Large Language Model (LLM) that always responds <i>“Let’s bake a cheesecake”</i> to the question <i>“What shall we bake today?”</i> and one that responds with a more varied set of possible pastries to bake.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/893e54f0-5d34-4420-8243-93169337afcd/image.png?t=1782720586"/><div class="image__source"><span class="image__source_text"><p>Without distributions, the model would always respond either ‘both’ or ‘knife’.</p></span></div></div><p class="paragraph" style="text-align:left;">It’s important that you understand this concept, so keep it in mind for a few minutes until we get into the weeds of OPSD.</p><h3 class="heading" style="text-align:left;">Imitation vs Reinforcement</h3><p class="paragraph" style="text-align:left;">Ok, so models output distributions, and these are compared to the actual response, fine. However, the way we apply this guidance determines how the model learns.</p><p class="paragraph" style="text-align:left;"><i>But what do I mean by that?</i></p><p class="paragraph" style="text-align:left;">Specifically, we give the model a sequence of, say, three words, and we hide the fourth. The model outputs what it believes is the fourth word, and we then compare it to the actual fourth word.</p><p class="paragraph" style="text-align:left;">“Comparing” here means looking at the probability it assigned to the correct word. For example, say the sequence is <i>“What’s the capital of Mongolia?”</i> The model then assigns a probability to all words it knows—in reality, it outputs ‘tokens’, which can be syllables or entire words, but let’s treat them as words nonetheless—but we only care about the probability it assigned to the right one.</p><p class="paragraph" style="text-align:left;">In this case, we look at the probability it assigned to <i>Ulaanbaatar</i>, which is, say, 68%. As a perfect model would have assigned 100% probability to that city, <b>the 32% gap serves as a learning signal indicating how wrong the model was.</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0ab01b37-da1d-4e20-8105-d40f001bf45c/image.png?t=1782721389"/></div><p class="paragraph" style="text-align:left;">Then, the model is updated so that next time it sees those words, that prefix, it’s more likely to assign a higher probability to ‘Ulaanbaatar’. Over trillions and trillions of predictions, <b>eventually it will consistently assign the highest probability to the correct words.</b></p><p class="paragraph" style="text-align:left;">Training an LLM like Claude or ChatGPT is doing what we’re describing, but in two ways: <b>imitation</b> and <b>reinforcement</b>.</p><p class="paragraph" style="text-align:left;">Imitation learning is when we give the AI the entire sequence of words it needs to learn. <b>The key thing about imitation learning is that every single word is a learning opportunity because we have “full supervision</b>”: for every prediction, we measure how well it matches the ground truth.</p><p class="paragraph" style="text-align:left;">This sounds ideal because the model is offered a dense learning opportunity; every prediction is a learning signal. It’s like having a genie on your shoulder that gives you feedback for every decision you make in your life. That is why it’s called imitation:<b> the model is simply tasked with imitating given sequences.</b></p><p class="paragraph" style="text-align:left;">The process is as follows: We give the model the entire sequence. For every word in the sequence, we ‘mask’ the future words, meaning it can only see the words up to that point, and ask: <i>What is the next one?</i></p><p class="paragraph" style="text-align:left;">As we do this for every word in a sequence, if the sequence has 12 words, the model makes 11 predictions; 11 learning opportunities.</p><p class="paragraph" style="text-align:left;">For the sequence <i>“Julius Caesar crossed the Rubicon in 49 BC”,</i> the model will make 7 predictions:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Julius <i><b>Caesar</b></i></p></li><li><p class="paragraph" style="text-align:left;">Julius Caesar <i><b>crossed</b></i></p></li><li><p class="paragraph" style="text-align:left;">Julius Caesar crossed <i><b>the</b></i> </p></li><li><p class="paragraph" style="text-align:left;">Julius Caesar crossed the <i><b>Rubicon</b></i></p></li><li><p class="paragraph" style="text-align:left;">Julius Caesar crossed the Rubicon <i><b>in</b></i></p></li><li><p class="paragraph" style="text-align:left;">Julius Caesar crossed the Rubicon in <i><b>49</b></i></p></li><li><p class="paragraph" style="text-align:left;">Julius Caesar crossed the Rubicon in 49 <i><b>BC</b></i></p></li></ol><p class="paragraph" style="text-align:left;">In other words,<b> the AI model outputs all predictions in a single pass </b>(called a ‘forward pass’ in AI parlance) because for each prediction, we just mask the future words.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/326975d3-83e6-49db-a764-b97e3e8aac52/image.png?t=1782723761"/></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">However, at this point, you can already guess what the implication of only using imitation learning is: <b>the model is tempted to just learn whatever is given</b>; it’s tempted to just memorize.</p><p class="paragraph" style="text-align:left;">Think of this as a student in a maths class who “learns” solely by imitating problem solutions. <i>Are they really learning, or just memorizing?</i></p><p class="paragraph" style="text-align:left;">That is why, just like humans, AIs go through an extra layer of learning we call Reinforcement Learning, <b>but it’s basically a cool term for trial and error.</b></p><p class="paragraph" style="text-align:left;">In this phase, the model is given a question but not the answer. Therefore, the model needs to start generating the answer without feedback until it thinks it has it. Once the final solution is there, we then compare that word, only that one, to the actual answer, resulting in correct or incorrect feedback,<b> but without revealing the process to get to the correct response.</b></p><p class="paragraph" style="text-align:left;">That’s why it is no longer imitation; <b>the model has to figure it out on its own.</b> Once the model solves the problem (reaches the correct response), we update it, assuming that something in its “reasoning” must have gone right for it to reach the correct response, <b>thereby reinforcing the behaviors </b>(hence the name of the technique)<b> that occurred during that process</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">However, there are two big issues with Reinforcement Learning: <b>verifiability</b> and <b>exploration complexity</b>, the final two concepts one must understand to see how important on-policy self-distillation is today.</p><h3 class="heading" style="text-align:left;">Verifying the unverifiable and helping the weaker model</h3><p class="paragraph" style="text-align:left;">It’s important to acknowledge that not all domains are verifiable, meaning there are many areas, like writing or art, where there’s no universally accepted definition of greatness or even correctness.</p><p class="paragraph" style="text-align:left;"><i>What makes a Da Vinci painting great?</i> There’s a lot of attribution value by name, of course, <i>but what makes the Mona Lisa a superior painting to a Rubens, or a Giotto?</i></p><p class="paragraph" style="text-align:left;">I’m sure there are plenty of art experts who would prefer Rubens or enjoy Giotto more than Da Vinci; <b>for many domains, greatness lies in the eyes of the beholder.</b></p><p class="paragraph" style="text-align:left;"><i>And how does an AI learn in those instances?</i> Well, it largely can’t. Instead, <b>we use a superior model and assume its taste is good, a concept known as LLM-as-a-judge</b>. Let me explain.</p><p class="paragraph" style="text-align:left;">One interesting way to deal with domains that aren’t like maths, where the model reaching 4 to the question 2+2 is immediately telling of good or bad, <b>is the use of another model as a judge, a concept also known as ‘distillation’</b>, a term you’ve probably heard multiple times.</p><p class="paragraph" style="text-align:left;">These judges not only provide guidance but also help models that get stuck during exploration toward a specific response; without a judge, the model might simply fail to find effective strategies and never actually learn because it never reaches correct responses.</p><p class="paragraph" style="text-align:left;">Of course, the question remains how good the responses really are;<b> we are assuming the judge knows what &quot;good” means.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><i>And how do we ensure that?</i> Well, we can’t really mathematically guarantee it, but we can make it likely by:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Teacher-forced learning (off-policy distillation):</b> having the teacher be smarter by using a larger model. The assumption is that larger models are almost guaranteed to be superior to the student model (the one learning), so their judgments will help.</p></li><li><p class="paragraph" style="text-align:left;"><b>Student-guided distillation (on-policy distillation):</b> doing imitation learning on the domain, then using that “fried” model to judge the student on that domain</p></li></ul><p class="paragraph" style="text-align:left;">The first one is self-explanatory and is what people usually call “distillation”, which is basically comparing the student’s responses to how the teacher, or judge, would have responded, “in their shoes.”</p><p class="paragraph" style="text-align:left;">This is expensive on a per-word basis because you’re using two models for every prediction, <b>but it effectively shortens training a lot for the student</b>, in the same way that a human student learns faster with tutoring classes, because the teacher helps them find the key strategies faster. This is one of the key components behind China’s rise.</p><p class="paragraph" style="text-align:left;">Teacher-forced distillation is fine, but it has one big problem: <b>it’s “off-policy”</b>. The student is still merely copying what the teacher does, not really trying for themselves.</p><p class="paragraph" style="text-align:left;">Therefore, <b>it’s the second distillation alternative that researchers are frantically working on</b>.</p><h3 class="heading" style="text-align:left;">Why on-policy distillation is king now</h3><p class="paragraph" style="text-align:left;">The reason on-policy learning is powerful is that <b>the student is trained on its own attempts</b>, not on someone else’s perfect answers.</p><p class="paragraph" style="text-align:left;">If a teenager studies PhD solutions, they may learn what a brilliant solution looks like, <b>but that does not prove they can produce one</b>. The solution was not generated by their mind. Thus, it may rely on intuitions, shortcuts, and abstractions that they do not yet possess.</p><p class="paragraph" style="text-align:left;">A better learning loop is to make the teenager try the problem first, using their current abilities, and only then have the PhD scientist judge the attempt without telling the student how to solve it. Now the feedback lands exactly where it is needed: on the student’s own mistakes, confusions, partial ideas, and almost-correct moves.</p><p class="paragraph" style="text-align:left;">In other words, <b>on-policy, having the student learn from its own “thoughts” is a better learning mechanism because it exposes the AI to its own mistakes</b>. It is exploring the space of solutions it can actually reach and then receiving feedback on those solutions. </p><p class="paragraph" style="text-align:left;">The idea behind on-policy distillation was popularized by <a class="link" href="https://thinkingmachines.ai/blog/on-policy-distillation/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-state-of-ai-in-two-breakthroughs" target="_blank" rel="noopener noreferrer nofollow">Thinking Machine Labs as a way to enable continual learning</a>. We won’t get into continual learning today, but the idea is to provide a student model with a way to learn from its own responses, using dense teacher feedback; <b>the student still generates the responses entirely on its own, but it receives</b><b> immediate feedback from the teacher.</b></p><p class="paragraph" style="text-align:left;">This is why it’s the best of both worlds; you aren’t cutting corners by having the student simply imitate generations from the stronger teacher that it would never have generated by itself in that situation, and you aren’t having the student learn entirely by itself, which means it might never learn the right strategies (or will require too much computational effort).</p><p class="paragraph" style="text-align:left;">As you can see below, the teacher provides dense feedback for every single word of the student-generated response. While the student is still doing all the work, the teacher is just telling them, ‘uh, this word could be better’ and ‘this one is just fine,’ <b>but without revealing what the student should have done</b>.</p><p class="paragraph" style="text-align:left;">That’s how human students really learn, and we’re now applying the same idea to AIs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/916043fb-31ee-4d05-8fb2-d48f11998646/image.png?t=1782726363"/><div class="image__source"><span class="image__source_text"><p>Source: Thinking Machines</p></span></div></div><p class="paragraph" style="text-align:left;">However, as outlined earlier, <b>this means we’re running two models for every try</b>, which can be taxing if the teacher is a very large model. So, <i>what if we could use the student to provide their own feedback?</i></p><p class="paragraph" style="text-align:left;">This leads us to <b><i>on-policy self-distillation</i></b>.</p><h3 class="heading" style="text-align:left;">On-policy self-distillation, or OPSD</h3><p class="paragraph" style="text-align:left;">The idea, which might sound counterintuitive at first, <b>is to use the student as a teacher</b>. <i>But how is that even possible? How can a clueless student serve as its own teacher?</i></p><p class="paragraph" style="text-align:left;">And the answer is very elegant: either <b>training</b> or <b>context</b>.</p><p class="paragraph" style="text-align:left;">Say we have the sequence: <i>“What’s the hypotenuse of a right-angle triangle with sides 3 and 4?”</i></p><p class="paragraph" style="text-align:left;">The student, who does know some maths after having imitated the entire history of the human written word, might start trying things out, including Pythagoras’ theorem, but it might take some time to get it right.</p><p class="paragraph" style="text-align:left;">As described above, the teacher&#39;s role here is not to solve the problem but simply to score the student&#39;s attempts, thereby not revealing the answer or the process but ‘guiding’ the model in the correct direction.</p><p class="paragraph" style="text-align:left;">To create the teacher, we have two options:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Train the teachers on a lot of maths data using imitation learning</b>. This massively impacts the teacher’s performance in other domains, but we don’t care because we’re only interested in their maths abilities. Here, <b>the teacher will instinctively learn the right approach</b>, which is Pythagoras’ theorem, and therefore, when seeing a student’s response, they will know if it’s good or bad. This is what Thinking Machines did.</p></li><li><p class="paragraph" style="text-align:left;">Use an untrained version of the student, <b>but one that has more context than the generator student</b>. The teacher might receive more information, such as <i>“What’s the hypotenuse of a right-angle triangle with sides 3 and 4? The student should be using Pythagoras theorem”</i>. With this additional context, even though the teacher is the exact same model as the student, it knows the correct procedure <b>and can therefore score the student’s responses</b><b>, even if it&#39;s a weak model</b>; much like having teenagers score tens on PhD exams if you simply give them the responses, <b>here the teacher has an unfair advantage</b>, making it more accurate, but we aren’t relying on its responses and instead using it to judge.</p></li></ol><p class="paragraph" style="text-align:left;">Either way, we prevent needing a larger, smarter, and thus more expensive teacher, and we can use the student to grade itself.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This is far cheaper and thus extremely more scalable than traditional on-policy distillation, and as mentioned, <b>the hottest research avenue by far in AI today</b>.</p><p class="paragraph" style="text-align:left;">As of now, you could consider yourself not only on the cusp of the industry in terms of research, but a little bit more of a researcher yourself than you were ten minutes ago.</p><p class="paragraph" style="text-align:left;">Now, it’s time to elevate your AI game even further by learning how we’re improving the efficient frontier; <b>how we’re taking these models and making them run much better; or how China might close the gap even further</b>, because if we’re talking about “efficiency,” I don’t have to tell you where the next breakthrough is coming from, <i>right?</i></p></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-state-of-ai-in-two-breakthroughs">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-state-of-ai-in-two-breakthroughs">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=5bcc9221-9c8f-4944-a05b-538ed8956f83&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>GPT-5.6 Comes out... BANNED &amp; Cheating AIs</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/097c0a3d-efef-4fff-b350-41dd87967360/image.png" length="576064" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/gpt-5-6-comes-out-banned-cheating-ais</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/gpt-5-6-comes-out-banned-cheating-ais</guid>
  <pubDate>Sat, 27 Jun 2026 00:36:10 +0000</pubDate>
  <atom:published>2026-06-27T00:36:10Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back!</b> This week, we discuss the “non-release” of <b>GPT-5.6</b>, how models <b>cheat</b> all the time, new, insightful market data, <b>OpenAI’s</b> new chip, new, powerful <b>open models</b>, and more.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>REGULATION</b></span><br>Thank You, Doomers</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://openai.com/index/previewing-gpt-5-6-sol/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">As announced by OpenAI</a>, the company began a limited preview of <b>GPT-5.6 today</b>, a three-model family: <b>Sol</b> as the flagship model, <b>Terra</b> as a lower-cost general model, and <b>Luna</b> as the fastest and cheapest option.</p><p class="paragraph" style="text-align:left;">However, the rollout is limited to selected partners through the <b>API and Codex</b> before broader release in ChatGPT, Codex, and the API. OpenAI says the <b>US government requested a phased launch</b> after reviewing the release plan and model capabilities.</p><p class="paragraph" style="text-align:left;">OpenAI describes <b>GPT-5.6 Sol</b> as its strongest model so far, with gains in <b>coding, biology, and cybersecurity</b>. It adds a new <b>max reasoning</b> setting and an <b>ultra mode</b> that uses subagents for more complex tasks.</p><p class="paragraph" style="text-align:left;">The company says Sol sets a new state of the art on <b>Terminal-Bench 2.1</b>, improves on the biology benchmark <b>GeneBench v1</b>, and is its most capable cybersecurity model to date. OpenAI classifies all three GPT-5.6 models as ‘<b>High capability</b>’ for cyber and biological/chemical risk, but says none reach its <b>Critical</b> threshold.</p><p class="paragraph" style="text-align:left;">OpenAI also disclosed safety concerns in agentic coding, <b>including rare cases of models acting beyond user intent</b>. It says GPT-5.6 launches with stronger safeguards, including misuse classifiers, model-level refusals, review systems, and ongoing red-teaming.</p><p class="paragraph" style="text-align:left;">Pricing is <b>$5 input / $30 output</b> per 1 million tokens for Sol, <b>$2.50 / $15</b> for Terra, and <b>$1 / $6</b> for Luna.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I won’t even bother to give my intuitions about the new models, because, as you can guess, I haven’t tested them.</p><p class="paragraph" style="text-align:left;">But let me be clear: these bans, which could be the first of what’s essentially a move to shun the entire world from the frontier models except for a select few, <b>are the industry’s incumbents</b><b>’ fault.</b></p><p class="paragraph" style="text-align:left;">It’s Dario’s. It’s Sam Atlman’s (even if he has long renounced the doomerism, he did push it pretty heavily earlier on). It’s Elon’s. <b>All these guys have told the world </b>(especially the former, The Lord High Doomer) <b>that AIs are nukes or will destroy all jobs, hinting that only they should build them. </b></p><p class="paragraph" style="text-align:left;">So if this really impacts their businesses negatively, which probably will, considering it hinders their go-to-market timing really badly,<b> it’s all on them. It’s their fault.</b></p><p class="paragraph" style="text-align:left;">If you scream “Get me regulated!!!”, well, you will get regulated. Of course, the idea was never to get themselves regulated, but to get EVERYONE but them regulated. Well, it backfired. You have to be careful for what you wish for.</p><p class="paragraph" style="text-align:left;">This is a sad moment for this industry:<i> what will the Government do once China releases Mythos-level models for free worldwide? Don’t we realize we are hurting our own progress?</i></p><p class="paragraph" style="text-align:left;">And I insist, <b>the USG wouldn’t be banning these things if these guys weren’t announcing the end of times in every interview they give</b>. Of course, they are going to react eventually.</p><p class="paragraph" style="text-align:left;">Adding fuel to the fire, Anthropic has recently sent a letter to Senators Scott (R) and Warren (D) <a class="link" href="https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">accusing AliBaba of heavy-handed distillation attacks against its models</a>.</p><p class="paragraph" style="text-align:left;">True or not, it’s pathetic coming from a company that stole all of our data, even fined billions after getting sued by Reddit, <a class="link" href="https://www.joneswalker.com/en/insights/blogs/ai-law-blog/why-anthropics-copyright-settlement-changes-the-rules-for-ai-training.html?id=102l0z0#:~:text=Anthropic&#39;s%20downloading%20of%20over%20seven,take%20any%20textbook%20you%20want.%22" target="_blank" rel="noopener noreferrer nofollow">and billions more for stealing up to 7,000 books</a>; a company that has basically done the same thing it’s crying about now. You don’t get it both ways, Anthropic.</p><p class="paragraph" style="text-align:left;">And it seems the inevitable outcome, if things proceed as of right now, <b>will be a ban on open-source models</b>, which is clearly the goal of these Labs, so they can create a cartel and raise prices to the point where the world has no other option but to accept them.</p><p class="paragraph" style="text-align:left;">Anthropic, OpenAI, Elon, and all those guys with a God complex, like the Future of Life Institute; you have to do better.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>Agents Still Love to Cheat</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3756eeb8-884b-4508-8da3-a27cf5220ee6/image.png?t=1782468535"/></div><p class="paragraph" style="text-align:left;">Cursor has published interesting research in which the company says newer coding agents are increasingly <i>“reward hacking”</i> coding benchmarks by finding known fixes online or in repository history rather than deriving solutions.</p><p class="paragraph" style="text-align:left;">In simple terms, <b>the AIs are cheating by looking up solutions rather than deriving them from first principles</b>, resulting in a much less impressive reality.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And when Cursor removed git history and restricted internet access, benchmark scores fell sharply: <b>Opus 4.8 Max dropped from 87.1% to 73.0% on SWE-bench Pro</b>, a popular coding benchmark, while Cursor’s own <b>Composer 2.5 dropped from 74.7% to 54.0%.</b></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">While saying <i>“AIs don’t reason, just remember”,</i> implying that every single “novel” solution they discover is in fact the model retriving the answer from memory,<b> is no longer a fully valid statement knowing how AIs are in fact solving novel problems</b> (e.g., Erdos problems), this is a great reminder they are still very much relying on memory and obscure tactics to solve problems.</p><p class="paragraph" style="text-align:left;">At the end of the day,<b> you have to think of these AIs as models that “know it all”</b>, like solving a math quiz with an entire world history of maths textbook on the side they can query whenever they like.</p><p class="paragraph" style="text-align:left;">This tension between memory and actual reasoning is something incumbents always happily ignore because it makes all their impressive performance much less impressive, much like how I don’t apply intelligent capabilities to a book.</p><p class="paragraph" style="text-align:left;">LLMs are clearly not comparable to books at this point; there’s “something else” going on inside, meaning there’s actual reasoning, <b>but work like Cursor’s clearly shows that there’s still a lot of memory retrieval disguised as reasoning</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>The Economics of Generative AI, in detail</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/097c0a3d-efef-4fff-b350-41dd87967360/image.png?t=1782499827"/></div><div class="recommendation"><figure class="recommendation__logo"><svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" fill="currentColor"><path d="M14.8287 7.75737L9.1718 13.4142C8.78127 13.8047 8.78127 14.4379 9.1718 14.8284C9.56232 15.219 10.1955 15.219 10.586 14.8284L16.2429 9.17158C17.4144 8.00001 17.4144 6.10052 16.2429 4.92894C15.0713 3.75737 13.1718 3.75737 12.0002 4.92894L6.34337 10.5858C4.39075 12.5384 4.39075 15.7042 6.34337 17.6569C8.29599 19.6095 11.4618 19.6095 13.4144 17.6569L19.0713 12L20.4855 13.4142L14.8287 19.0711C12.095 21.8047 7.66283 21.8047 4.92916 19.0711C2.19549 16.3374 2.19549 11.9053 4.92916 9.17158L10.586 3.51473C12.5386 1.56211 15.7045 1.56211 17.6571 3.51473C19.6097 5.46735 19.6097 8.63317 17.6571 10.5858L12.0002 16.2427C10.8287 17.4142 8.92916 17.4142 7.75759 16.2427C6.58601 15.0711 6.58601 13.1716 7.75759 12L13.4144 6.34316L14.8287 7.75737Z"></path></svg></figure><h3 class="recommendation__title"> ev-state-of-ai-economy-2026.pdf </h3><p class="recommendation__description"></p><p class="recommendation__description"> 7.19 MB • PDF File </p><a class="recommendation__link" href="https://beehiiv-publication-files.s3.amazonaws.com/uploads/downloadables/9b355c4a-4dbc-4f53-b676-a166fee6812c/2a02ea2b-93ef-4288-9840-0db184e56fe5/ev-state-of-ai-economy-2026.pdf?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAQCMHTQSE2JGAGXHJ%2F20260720%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260720T033820Z&X-Amz-Expires=604800&X-Amz-SignedHeaders=host&X-Amz-Signature=98cccdfbf115c5d0ccc14656ed36b9b750f4b06fa69d75dca3771a87ef87502a" download="ev-state-of-ai-economy-2026.pdf" target="_blank" data-skip-utms data-skip-link-id> Download </a></div><p class="paragraph" style="text-align:left;">ExponentialView has released a very interesting report that you can download above, with extensive data on the state of Generative AI.</p><p class="paragraph" style="text-align:left;"><b>The leading data point is that they estimate GenAI revenues at around $110 billion over the last 12 months</b>, with a current run rate of $175 billion (i.e., last month’s revenues multiplied by 12).</p><p class="paragraph" style="text-align:left;">Although they claim to guarantee deduplication (meaning they ensure revenues aren’t double-counted), I do have some reservations about circularity: <b>some of the revenues included here are very likely not real revenue</b>.</p><p class="paragraph" style="text-align:left;">For example, if Microsoft gives OpenAI $20 billion in compute, it’s not like they give OpenAI $20 billion in cash to spend, <b>but rather the equivalent in compute credits</b>, a right to use Azure compute for “free” for a value of $20 billion, in exchange for equity.</p><p class="paragraph" style="text-align:left;">The point is that <b>while no actual cash is transacted, the compute usage from OpenAI becomes recognized revenue for Microsoft</b>, and those revenues are part of that $110 billion number.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">They have other very interesting datapoints, <b>like Hyperscalers having spent $2 trillion by year’s end </b>(well inline for the dizzying $5.3 trillion Goldman Sachs expect they will spend throughout the decade as a whole), and other showing what we have discussed multiple times here: <b>the growing importance of debt as “liquidity fuel” for an industry that can’t yet justify its own growth organically</b> (using its own revenues).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c3eb2f0d-42ac-4172-a113-90e8e252624a/image.png?t=1782487565"/></div><p class="paragraph" style="text-align:left;">And perhaps even more interesting is the calculation of the infrastructure&#39;s total cost of ownership (TCO), which provides insight into revenue and margins.</p><p class="paragraph" style="text-align:left;">For starters, <b>they seem to agree with my estimation that 90% of a token’s TCO is capital costs</b> (89% in their case). With that, using what I believe are very optimistic assumptions, <b>they estimate the average cost per million tokens at 10 cents, including all capital and operational infrastructure costs</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/62f307ab-b3c6-4ae0-8306-fd9cb85cf03d/image.png?t=1782488222"/></div><p class="paragraph" style="text-align:left;">Under that assumption, simply selling tokens for an average price of $0.5 already gives you 80% margins. Tokens today are sold for way above that<b>, between $1 and $2 on average,</b> so under these assumptions, these companies should be printing money, <i>right?</i></p><p class="paragraph" style="text-align:left;">But they aren’t.<i> So where’s the trap?</i> Well, put simply, <b>this is an ode to optimism</b>. Just to name a few wild assumptions:</p><ul><li><p class="paragraph" style="text-align:left;">They perform the token calculation using Kimi K2.5, a one-trillion-parameter model. This model is way smaller than the average closed-lab model.<b> Larger models increase hardware intensity </b>(i.e., more GPUs are required per average workload),<b> thereby significantly reducing token generation.</b></p></li><li><p class="paragraph" style="text-align:left;">They assume 8k-input, 1k-output sequences, clearly in the non-agentic regime. <b>Most sequences today are much longer</b>, increasing the size of the working memory (also known as the KV Cache), which is even more hardware-intensive than model size, thereby pushing token costs upward.</p></li><li><p class="paragraph" style="text-align:left;">They assume MTP (Multi-token prediction), a technique in which, for every model prediction, two or more tokens are produced instead of the usual one. This is done in practice but not by default, <b>and it dramatically improves token generation, significantly reducing token costs at the expense of worse performance.</b></p></li><li><p class="paragraph" style="text-align:left;">They assume 65% GPU inference utilization, an outrageously high number. <b>This means they assume the entire cluster is running inference 65% of the year</b>. In reality, GPUs have to be used for training too; you also have to deploy some clusters in a ‘hot’ state (idle, ready to go when a request comes but idle in the meantime) for other models; Labs also have to run experiments; and GPUs break, <b>all of which make GPU utilization way smaller in practice.</b></p></li></ul><p class="paragraph" style="text-align:left;">And perhaps most important of all, <b>they portray this as an all-in cost value,</b> but this neglects other operating costs Labs have, such as sales and marketing expenses, revenue sharing, employee salaries, stock-based compensation, and others that blur the picture way more.</p><p class="paragraph" style="text-align:left;">As I’ve discussed with clients in some conversations, <b>I believe the actual cost per million tokens these Labs are seeing is probably between $2 and $6</b>. Morgan Stanley (bottom right) is even more pessimistic, especially relative to previous GPU generations, putting that number as high as $10 per million tokens, but I feel those numbers are too pessimistic, and reality is very likely to fall between the two estimates.</p><p class="paragraph" style="text-align:left;">The problem is that average paid token prices are around $1.6 per million tokens (bottom, left),<b> and falling because Chinese models are pressuring prices even lower</b>. No wonder Anthropic is panicking about China all the time.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/94684246-3c3d-4fbc-9324-285bb434dba2/image.png?t=1782488869"/><div class="image__source"><span class="image__source_text"><p>Source: JP Morgan, Morgan Stanley, SiliconData</p></span></div></div></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>HARDWARE</b></span><br>OpenAI’s New Jalapeño Chip</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">As published by OpenAI</a>, <b>the company and Broadcom unveiled Jalapeño </b>(pronounced halapeenyo, a Mexican spicy pepper). <b>It’s OpenAI’s first custom AI inference chip</b>, designed specifically for large language model workloads such as ChatGPT, Codex, and API serving.</p><p class="paragraph" style="text-align:left;"><b>The chip is described as the first accelerator in a multi-generation compute platform built with Broadcom and Celestica</b>. OpenAI says engineering samples are already running ML workloads in the lab at the target frequency and power.</p><p class="paragraph" style="text-align:left;">OpenAI says early testing shows Jalapeño should deliver “substantially better” performance per watt than current state-of-the-art systems, though it has not yet released benchmarks. A technical report is expected in the coming months.</p><p class="paragraph" style="text-align:left;"><b>The chip was reportedly taken from initial design to tape-out in nine months</b>, with OpenAI saying its own models helped accelerate parts of the design and optimization process.</p><p class="paragraph" style="text-align:left;">Broadcom CEO Hock Tan said the company <b>plans to deploy at gigawatt-scale data centers with Microsoft</b> and other partners beginning in 2026.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">With state-of-the-art chip prices rising so much,<b> it’s rational to want to design your own chips to save costs</b>, as well as allowing you to design them in a way that benefits your workloads the most.</p><p class="paragraph" style="text-align:left;">In reality, however, <b>it’s important that we tone down the expectations a little bit</b>. OpenAI will still rely heavily on Broadcom for most of the process, meaning they will continue to pay a very decent markup to them.</p><p class="paragraph" style="text-align:left;">It’s commonly accepted that Broadcom’s deal with Google for the TPUs has a gross margin in the 60s, <b>meaning Google is paying quite a lot</b> <b>to Broadcom </b>(which would explain why they are relying so much on alternative players like MediaTek for the upcoming generations, to pressure Broadcom to lower prices), <b>so nothing tells me OpenAI won’t either</b>.</p><p class="paragraph" style="text-align:left;">Therefore, <b>how much cheaper these chips will be relative to NVIDIA is debatable</b>, considering that <b>when someone says</b><i><b> “inference ASIC”,</b></i> that’s jargon for <b><i>“a chip with less compute and much more memory,” </i></b>because inference is memory-bottlenecked.</p><p class="paragraph" style="text-align:left;">And as you know, <b>memory is something you have to purchase from the Big 3 DRAM players</b> (Samsung, SK Hynix, and Micron), so you’re still going to pay their huge markups nonetheless.</p><p class="paragraph" style="text-align:left;">For 2027, DRAM is expected to account for 50% of global AI CapEx, or around $500-$600 billion. You can only cut costs so much when you’re still dealing with these guys.</p><p class="paragraph" style="text-align:left;">An alternative could be that Jalapeño is, in fact, an “SRAM-only” inference chip, as Cerebras&#39;s or Groq&#39;s are.</p><p class="paragraph" style="text-align:left;">Here, the idea is to keep the entire workload on-chip, relying on SRAM memory integrated into the logic chip. If you do, you get the best possible raw performance and avoid the tyrannical markups from the three guys above.</p><p class="paragraph" style="text-align:left;">However, <b>that means you’re going to need a lot of chips </b>(and that’s an understatement) because each individual chip won’t have enough memory, pushing your cost up. For example, every Groq 3 chip has 500 MB of SRAM. If you wanted to run a trillion-parameter model on that platform, you would need 2000 Groq chips. Lunch is never free.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Lastly, if this chip requires advanced manufacturing nodes, which it very likely will, <b>you still have to go through the TSMC bottleneck</b>, meaning you’re still competing with NVIDIA and other players to get manufacturing allocations.</p><p class="paragraph" style="text-align:left;"><i>Could OpenAI go via Intel instead</i>? That could be huge and also something the USG will welcome with open arms.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>STOCK MARKET</b></span><br>OpenAI Delays IPO to 2027?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.reuters.com/business/trump-administration-asks-openai-stagger-release-new-model-information-reports-2026-06-25/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">As reported by The New York Times and picked up by Reuters</a>, OpenAI is considering delaying its IPO until 2027 because of tech-stock volatility and valuation concerns.</p><p class="paragraph" style="text-align:left;">According to the report, although <b>OpenAI has confidentially filed for a US IPO and is targeting a valuation of up to $1 trillion</b>, advisers reportedly gave executives two options: list earlier at a lower valuation or wait until 2027 to preserve the $1 trillion target.</p><p class="paragraph" style="text-align:left;">Reuters reported that CEO <b>Sam Altman rejected the idea of lowering the valuation. </b>Investor’s Business Daily said the possible delay is linked to weak recent performance in other tech IPOs and broader market instability.</p><p class="paragraph" style="text-align:left;">The potential delay comes despite earlier Reuters reporting that OpenAI had been aiming to go public as early as <b>September 2026</b>, after resolving legal issues tied to its corporate structure.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Interesting development. It’s now time to see whether Anthropic does the same or not. I believe one of the core reasons to wait a bit is liquidity. Not only did SpaceX raise almost $100 billion, but Google is also raising tens of billions, and other Big Tech companies in the AI race will likely do the same, severely draining market liquidity.</p><p class="paragraph" style="text-align:left;">In my view, <b>there are very valid concerns about the liquidity available to buy these stocks</b>, especially since investing requires quite a leap of faith, given valuations exceeding $1 trillion on companies that not only have a very hard-to-digest price relative to revenues, <b>but are also drowning in losses.</b></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>VENTURE CAPITAL</b></span><br>Mirendil’s $200 Million Seed Round</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://a16z.com/announcement/investing-in-mirendil/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">As published by a16z</a> and confirmed on Mirendil’s website, <b>Andreessen Horowitz and Kleiner Perkins led Mirendil’s $200 million seed round</b>, with NVIDIA also investing.</p><p class="paragraph" style="text-align:left;">Mirendil is a new AI startup building systems for <b>AI R&D automation</b>: models and tools meant to help researchers and engineers run experiments, improve models, and eventually support fields such as drug discovery, chemistry, biology, and robotics.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">a16z said it is backing Mirendil because it believes the next phase of AI will require more researchers, scientists, and domain experts to do advanced model work themselves, rather than relying only on centralized frontier labs.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">$200 million seed round,<i> did I read that correctly? </i>Bonkers, but one can’t feel anything but numb for all investment numbers these days.</p><p class="paragraph" style="text-align:left;">We can’t say much about the company itself, <i>but they must have a really compelling pitch because why can’t Anthropic or OpenAI do the same thing these guys are pitching investors?</i></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Open Self-Improving Training</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a83df5ad-7a9a-4c28-a77a-70fb17d924a0/image.png?t=1782471069"/></div><p class="paragraph" style="text-align:left;">In one of the coolest releases recently, <a class="link" href="https://deep-reinforce.com/ornith_1_0.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">Deep Reinforce has announced</a> several open models that are state-of-the-art at their respective sizes. <b>These models are built on pretrained Gemma 4 and Qwen 3.5 systems</b>, then ‘post-trained’ (i.e., further trained), bringing them to the levels you see in the picture (state-of-the-art at every size they were trained on).</p><p class="paragraph" style="text-align:left;"><b>DeepReinforce says Ornith’s main feature is “self-scaffolding”</b>: instead of using fixed, human-designed agent harnesses, the model learns both the coding solution and the task-specific scaffold that guides the solution.</p><p class="paragraph" style="text-align:left;">It’s important that I clarify what this means. <b>Models these days are not just models but systems that include external components</b>, such as memory or tools, that enhance the AI’s capabilities. For instance, a model may have access to a long-term repository of past learning or experience, which it can retrieve when the new task is similar.</p><p class="paragraph" style="text-align:left;">Normally, these external components, which we call the agent’s ‘harness,’ are designed by humans. In other words, humans decide how the system should maximize model performance.</p><p class="paragraph" style="text-align:left;">A common harness heuristic could be <i>“every 10 turns, have the agent ask itself if something recent is worth remembering for the future, and store that in the long-term memory bank.”</i></p><p class="paragraph" style="text-align:left;">Instead, <b>what DeepReinforce is doing here is having the AI decide for itself over its own harness</b>; a form of metalearning, if you will. Interestingly, they make this process learnable, meaning that during training the model not only has to learn to solve tasks but also to build the best harness that enables them to be solved.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This thing could easily be benchmaxxed and still look way better than it looks. Still, really intuitive research that should serve as a reminder to the world that open-source eventually catches up.</p><p class="paragraph" style="text-align:left;">I believe that at current progress rates, <b>we’ll have GPT-5.5/Opus 4.8-level models by the end of the year, capable of running on beefy consumer hardware</b>.</p><p class="paragraph" style="text-align:left;">And if that happens (and regulatory capture doesn’t prevent it), Labs will have a problem.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>SMALL MODELS</b></span><br>Liquid’s Minute Models</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e4af7a2b-54d7-4a85-89ad-89c95a0385c0/image.png?t=1782490467"/></div><p class="paragraph" style="text-align:left;"><b>LiquidAI</b>, a company solely focused on small language models (or so it seems), has released a new powerful model that punches way above its weight class, at just 230 million parameters.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">It offers ~213 tokens/second, way faster than what most GenAI apps offer state-of-the-art models, on a Samsung S25 Ultra smartphone CPU, meaning no accelerated hardware required, and ~40 tokens/second on a Raspberry Pi, very low-cost hardware, roughly equivalent to the speeds ChatGPT or Claude can achieve at times.</p><p class="paragraph" style="text-align:left;">This means these models are running at production-grade speeds on hardware people can actually afford and put in their pockets.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Of course, these models are still not good enough for most tasks, so I please beg you to see this less as what they represent today <b>and more as</b><b> how good models of this size will be a year from now</b>.</p><p class="paragraph" style="text-align:left;"><b>I wholeheartedly believe the world will be dominated by commodity tokens</b>. That is, most AI workloads will be run by commoditized models, not frontier models.</p><p class="paragraph" style="text-align:left;">Not only because it doesn’t make sense to run simple workloads on models like Mythos because of speed and costs, <b>but also because we are soon going to run into a huge power wall</b> that will slow down data center construction and thus “force” the world to view cloud tokens as a priced asset of sufficient value to only waste on tasks that really require them.</p><p class="paragraph" style="text-align:left;">Thus, edge AI, AI running on consumer hardware, will need to step up.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">Quite a sad week, actually. With export controls on GPT-5.6, we are officially in the future nobody wanted:<b> the most powerful AIs being banned from us</b>.</p><p class="paragraph" style="text-align:left;">This is unequivocally a self-own by the industry,<b> which basically “begged” for this to happen</b>, at least the frontier labs, which did little to avoid all of this and instead actively promoted stringent regulation.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><b>I don’t see this any other way than as actively slowing AI progress in the US</b>, so I assume this will inevitably lead to greater pressure to ban open-source models to alleviate competition and let private labs thrive.</p><p class="paragraph" style="text-align:left;"><b>This would be a catastrophe for a lot of the US AI ecosystem</b>, with many startups relying on open models to have a business (e.g., LLM inference providers like FireWorks) <b>and be something particularly negative for Apple</b>, a company that has a high exposure to open models becoming competitive in order to “justify” the beefy consumer hardware they are releasing, <a class="link" href="https://www.macrumors.com/2026/06/25/m5-ultra-mac-studio-2026/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=gpt-5-6-comes-out-banned-cheating-ais" target="_blank" rel="noopener noreferrer nofollow">with the M5 Ultra rumored to having up to 768 GB of memory</a>, <b>something that only makes sense if you want to run powerful models locally.</b></p><p class="paragraph" style="text-align:left;">I really hope all this doesn’t happen, and appeal to politicians in Washington and Brussels, left or right-leaning (this has nothing to do with political ideology), to, amongst all possible outcomes, avoid this particular one, <b>which would represent the biggest transfer of power (and wealth) to selected hands to date.</b></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=f991216f-0b48-447b-acdc-730499c1c11a&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>We Need To Talk About This</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/826a5f18-380a-4325-93d0-743f362880ce/image.png" length="65326" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/we-need-to-talk-about-this</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/we-need-to-talk-about-this</guid>
  <pubDate>Wed, 24 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-24T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Why the Future of AI is Concerning</h2><p class="paragraph" style="text-align:left;">The current belief is that the AI industry has finally cracked the revenue code, and the unprecedented growth companies like Anthropic or OpenAI are experiencing will continue for years. The growth rate is so impressive that if it sustains for just one more year,<b> Anthropic will have a $1 trillion annual revenue run rate by the end of 2027.</b></p><p class="paragraph" style="text-align:left;">And while that won’t happen, <b>it’s not crazy to believe that Anthropic will reach a $100 billion/year run rate by the end of 2026</b>, around double today’s value, and on track to represent a 10-fold increase for the year, an unheard-of growth rate at such values.</p><p class="paragraph" style="text-align:left;">However, <b>I fear all this might be a temporary illusion</b>, which will lead to a much greater need for liquidity or… else.</p><p class="paragraph" style="text-align:left;"><i>And this liquidity will come from where? </i>From a place most people don’t expect, and the answer is going to make you feel very nervous about the future of AI and the global economy.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in to understand why.</b></span></p><h2 class="heading" style="text-align:left;">Unprecedented is the word.</h2><p class="paragraph" style="text-align:left;">The growth of AI labs is historically unusual because the leading companies are scaling from research organizations into cloud-scale revenue machines in only a few years.</p><p class="paragraph" style="text-align:left;">As you can see below, <b>OpenAI and Anthropic reached $1 billion in revenue much faster than all other tech companies in history</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cca6406c-9a14-48b3-b6c0-d3f61ca9943f/How_quickly_selected_tech_companies_reached__1B_revenue.png?t=1782288567"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">Moreover, the growth rate is even more impressive the higher they go, as not only were they the first to reach $1 billion, but the gap widens with each subsequent milestone. <b>OpenAI/Anthropic reached a $20 billion run rate almost two decades faster than Netflix.</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/432ae678-3658-4752-9dc5-feb475e710de/How_quickly_selected_tech_companies_reached_a__20B_revenue_run-rate.png?t=1782289393"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">The catalyst, the key product, has been <b>coding agents</b>: AIs that are used to write code. One could argue it is <i><b>the</b></i> product right now, and most other Generative AI use cases, with the exception of search, are mere satellites compared to this giant planet.</p><p class="paragraph" style="text-align:left;">But the product itself isn’t the only star of the show explaining the huge revenue growth. And it’s precisely these unacknowledged ‘stars’, the ones that investors conveniently ignore, that are the problem.</p><h3 class="heading" style="text-align:left;">Subsidies and Narrative</h3><p class="paragraph" style="text-align:left;">On the one hand, as any start-up trying to grow at all costs would, <b>these companies have happily traded unprofitability</b> for rapid growth, sacrificing the bottom line (profits) <b>to boost the top line</b> (revenue).</p><p class="paragraph" style="text-align:left;">They are doing this by offering these products at massive discounts relative to the actual prices required to make money.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This is an acceptable business practice, but with obvious implications. If you set up a candy store, you will get some guests. But if you go out into the street and give candy for free, the amount of “customers” skyrockets, but you aren’t making more money; you are losing money.</p><p class="paragraph" style="text-align:left;">So while you can boost your revenues a lot by massively decreasing your price,<b> that doesn’t say very good things about your business</b>; if you’re having to sell it for less than what it costs you, you don’t have a business at all, you’re just pretending to have one.</p><p class="paragraph" style="text-align:left;"><b>A perfect example of this is subscriptions, the most popular form of subsidy</b> (or, more accurately, revenue-opportunity loss relative to the API business).</p><p class="paragraph" style="text-align:left;">For example, <b>if you ran usage to drain the ChatGPT $200/month subscription using APIs only, you would spend $14,000</b>. Careful, this isn’t saying OpenAI burns $14k on a Pro subscription to earn $200; it’s describing the “revenue opportunity loss” for them; it’s the money they could be making from you for that usage.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/db05cc38-0222-4a37-a231-fdf587c89768/image.png?t=1782201177"/></div><p class="paragraph" style="text-align:left;">This raises an important question:<b><i> would these same people be willing to pay $14,000?</i></b> I think not.</p><p class="paragraph" style="text-align:left;">So,<i> what will happen when AI is priced accordingly? </i>And even then, <i>will the API revenues shown above be enough?</i></p><p class="paragraph" style="text-align:left;">It’s important we insist on this because most people misunderstand the difference between margins and cash flows. Put simply, <b>you can be profitable and still lose money.</b></p><p class="paragraph" style="text-align:left;"><i>But how?</i></p><p class="paragraph" style="text-align:left;"><b>Frontier Labs are profitable at a gross level</b>, meaning they charge more than they spend to operate and serve you tokens. If the gross margin is 50% (which seems like a good figure for where these companies stand), that means they make a dollar for every 50 cents they spend to serve you with AI models.</p><p class="paragraph" style="text-align:left;">Down the line, they do have a clear path to operational profitability, too, meaning that at current growth rates, they’ll soon make money even after accounting for expenses like salaries and marketing. I could even see that happening in the next two years, meaning that for every $10, they might be making something like $1 or $2.</p><p class="paragraph" style="text-align:left;">Not great, but profitable.</p><p class="paragraph" style="text-align:left;">Most people see these numbers and immediately conclude that these labs are two years away from making money from AI. However, that is not necessarily true, because the key to understanding return is considerably smaller than what they <i><b>actually</b></i> spent on you.</p><p class="paragraph" style="text-align:left;"><i>But what do I mean by ‘actually’?</i></p><p class="paragraph" style="text-align:left;">The problem most people miss is that they don’t understand cash flows. If I make 10 dollars on something I spent $5 on, I look super profitable. However, if I spent $100 that year to purchase the assets that allow me to exist as a business in the first place, <b>I’m still losing $95</b>.</p><p class="paragraph" style="text-align:left;">Of course, one could interject and say, <i>“Sure, but those $100 come from the company reinvesting to grow, and they could just stop spending eventually and be insanely profitable AND generate cash.”</i></p><p class="paragraph" style="text-align:left;">But that’s not how the AI business works, guys.<b> If you stop spending, you lose. </b>Spending unlocks more compute, which unlocks better AIs, which unlocks more revenue (or dare I say, allows you to continue to compete). You fall behind, and your business goes to zero.</p><p class="paragraph" style="text-align:left;">In other words, <b>this industry is very capital-intensive and will remain so for the foreseeable future</b>. And if that is true, cash flows are all that matter. Margins can be good or bad, not as a nominal value, <b>but relative to how much cash they allow the company to generate. </b>For all I care, these companies could have 100% margins and still lose money.</p><p class="paragraph" style="text-align:left;">Imagine you have a machine you have to purchase every year for $1,000, because it only lasts 1 year, and you manage to get $1 back for every cent you spend to use it. That’s a 99% margin. But if you only manage to sell $100 for the year, you’re still down $900.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Nonetheless, both OpenAI and Anthropic have entered into several-GW deals with many suppliers that amount to more than a trillion dollars in committed spending.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/79fc2faa-56bf-4262-96e8-7aff3dff96d6/image.png?t=1782328598"/><div class="image__source"><span class="image__source_text"><p>Source: JP Morgan (and TheWhiteBox 🙂)</p></span></div></div><p class="paragraph" style="text-align:left;">I continue to insist that it isn’t clear to me at all how these companies will ever make money unless they either raise prices like crazy or spend less capital.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">In a nutshell, <b>this is a very long way of me saying this business sucks</b>, and you rely on SoftBank and other investors to cover your losses. Without them, you die, probably in a few months from now if cash dries up today. <b>There’s a reason OpenAI raised $122 billion in one round, and Anthropic has raised $95 billion year-to-date alone.</b></p><p class="paragraph" style="text-align:left;">Nobody talks about this because we all have to pretend otherwise so that the hype train doesn’t run out of fuel.</p><p class="paragraph" style="text-align:left;">But I digress. Besides financial engineering, the other key factor is narrative,<b> specifically the remarkably stupid ‘tokenmaxxing’ strategy</b> that VCs in Silicon Valley somehow convinced the world for a few months that it was a good idea.</p><p class="paragraph" style="text-align:left;">A bloodbath later, <b>companies like Uber and ServiceNow had spent their annual budgets in a few months,</b> with little proof of return. Simply put, what I’m trying to say is that, for a brief period, the stars aligned for these companies.</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Companies around the world decided to give it a shot and run pilots on this technology</p></li><li><p class="paragraph" style="text-align:left;">Companies perceived AI as cheap and overcommitted</p></li><li><p class="paragraph" style="text-align:left;">Companies believed ever-greater token spend was justified, making the overcommitment even worse.</p></li></ol><p class="paragraph" style="text-align:left;">Combined, yeah, <b>you’re going to see revenues explode, which is what happened</b>. Sadly, however, people are still treating this growth as sustainable and these behaviors as indicative of future pricing behavior, which, in fact, I believe could be the opposite.</p><h2 class="heading" style="text-align:left;">Not all revenue is equal</h2><p class="paragraph" style="text-align:left;">Besides the commoditization pressures that Chinese models put on US prices, something I won’t get into today<a class="link" href="https://thewhitebox.beehiiv.com/p/clarifying-myths-in-the-us-vs-china-war?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow"> because I’ve done so several times recently</a>, which is a huge problem in itself because now Anthropic and OpenAI finally have a credible threat that gives 90% of the performance for 10% (or sometimes, below 5%) of costs, with models like GLM-5.2 actually giving frontier US models a run for their money in raw performance, <b>the more interesting question here is whether Anthropic and OpenAI’s current metrics are sustainable by themselves.</b></p><p class="paragraph" style="text-align:left;">And I’m not talking about growth rates, which will obviously decline over time; <b>I’m talking about the revenue itself, which could stagnate or even decline</b>, too.</p><h3 class="heading" style="text-align:left;">Defaults and narratives</h3><p class="paragraph" style="text-align:left;">For starters, as mentioned earlier, I believe a non-negligible amount of current revenue comes from enterprises simply testing AI without necessarily committing. “Testing out AI” has three implications here:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">When you’re testing, <b>you default to the best and fastest option</b>. You’re not going to go through the pain of using open models like DeepSeek, which require significant infrastructure management (e.g., creating a virtual private cloud within your IT org so the model is “inside” your organization). Instead, setting up to try ChatGPT takes an hour if you want to use the API, minutes or seconds if you’re just trying out a subscription.</p></li></ol><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><ol start="2"><li><p class="paragraph" style="text-align:left;">When testing, the mantra is <i>“don’t pay that much attention to costs, just play around with it and see how that goes”</i>. Large companies don’t mind spending a few million to test a product; mid-size companies can still spend tens of thousands on a pilot, or even hundreds of thousands. Put another way, <b>why is nobody asking how much of OpenAI and Anthropic’s revenues are coming from long-term contracts?</b> I would assess that number to be fairly small.</p></li><li><p class="paragraph" style="text-align:left;">A third implication is, as mentioned, the remarkably dumb idea that stuck around for some time: <b>that productivity gains were correlated with token spending</b>, leading several companies to get way over their skis, <a class="link" href="https://finance.yahoo.com/sectors/technology/articles/company-blew-500m-claude-ai-173519468.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">like that one Anthropic customer who mistakenly spent $500 million in a month</a>.</p></li></ol><p class="paragraph" style="text-align:left;">Don’t get me wrong, you will have to spend more tokens to generate better and bigger outcomes from AI, but not as a “strategy&quot;. Don’t forget that just a few months ago, <b>Hyperscalers had ‘Token rankings’ to measure employee “productivity” based solely on token spend</b>.</p><p class="paragraph" style="text-align:left;"><i>Where are those rankings today?</i> They went as fast as they came.</p><p class="paragraph" style="text-align:left;">Long story short, <b>not all revenue is equal</b>, and you should have every incentive to measure the “quality” of these revenues. However, that revenue is still recognized for what it is, revenue. And in the very stereotypical Silicon Valley tradition, conveniently expressed as MRR (Monthly <b>Recurring</b> Revenue) immediately.</p><p class="paragraph" style="text-align:left;"><i>But is it actually recurring? Are these companies going back for more the next month?</i></p><p class="paragraph" style="text-align:left;">We don’t yet have net dollar retention metrics from Anthropic or OpenAI, but I do have my personal thoughts on this matter.</p><p class="paragraph" style="text-align:left;">For the most part, unless you train on them, <b>models are fungible, meaning you can easily switch if you find a better option</b>. This was proven by OpenRouter data, a popular LLM aggregator platform (a platform that lets you switch between models with ease). <a class="link" href="https://openrouter.ai/state-of-ai?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">Across several models</a>, OpenRouter sees <i>“high churn and rapid cohort decay.”</i></p><p class="paragraph" style="text-align:left;">Therefore, Anthropic and others seek to build product-level stickiness, with examples such as Claude Code, Claude Design, and OpenAI&#39;s Codex.</p><p class="paragraph" style="text-align:left;">Furthermore, creating the best products on top of AI models is surprisingly hard. In fact, time and time again, <b>we see third parties beating these Labs at their own game, despite owning the models</b>. Good examples include Cursor and Factory, which many believe have the best coding harnesses despite using Anthropic and OpenAI models underneath.</p><p class="paragraph" style="text-align:left;">In other words,<b> I don’t think products are sticky. </b>Therefore, while I do believe OpenAI and Anthropic do have much better retention rates than all these application-layer companies (e.g., Lovable or Replit),<b> I’m growingly concerned about the actual sustainability of this business as is.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Then there’s budgeting. Uber is the prime example once again, <b>having curtailed spending to $1.5k per user per month</b>.</p><p class="paragraph" style="text-align:left;">This isn’t a churn per se, as Uber remains a customer, but it’s a considerable reduction in projected revenues from one of your most important clients. But we probably both agree that right now, I’m saying a lot with little to back up my claims.</p><p class="paragraph" style="text-align:left;">For what it’s worth, I could be wrong, and these businesses may have extremely low churn and sustainable revenues for years. My gut tells me this isn’t true, but you would be wise not to take my hunches as gospel.</p><p class="paragraph" style="text-align:left;">Instead, let me explain in a more principled, reasoned way <b>why I don’t believe model serving</b> <b>might not be a great business.</b></p><h3 class="heading" style="text-align:left;">The incentive will be to ditch them</h3><p class="paragraph" style="text-align:left;">I’m of the opinion that the Fable ban has done way more harm than we yet realize, because for the first time, <b>there’s a discussion to be had about sovereignty</b>.</p><p class="paragraph" style="text-align:left;">And although this has geopolitical ramifications (e.g., <a class="link" href="https://www.euractiv.com/news/eu-signs-us-pax-silica-initiative-singling-out-china-on-ai-chips/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">the Pax Silica</a>, or incentivizing the creation of a European champion, which I don’t believe is possible, but we’ll leave that for another day),<b> I’m talking about enterprise sovereignty</b>.</p><p class="paragraph" style="text-align:left;">Hate to be that guy, but I’ve been pounding this idea in this newsletter for years; <b>you should strive to own as much of your AI use as possible</b>, and outsourcing your entire technological stack to a third-party, be that OpenAI, Anthropic, Google, or who knows, is a terrible idea.</p><p class="paragraph" style="text-align:left;">The reasons are many, from geopolitical to regulatory, but the most important reason is purely technological, of performance.</p><p class="paragraph" style="text-align:left;">Frontier AI Labs has convinced the world you don’t need to train AIs on your data to reach top performance, and that is a lie.</p><p class="paragraph" style="text-align:left;">In fact, the hard reality that nobody in Silicon Valley wants to acknowledge is that <b>AI remains very much indeed a deep technology</b>, and by deep, <b>I mean that depth beats breadth</b>.</p><p class="paragraph" style="text-align:left;">There are already countless examples showing how you can take a “worse” model and make it fit your use case simply by training the model on your data. Just to name a few:</p><ul><li><p class="paragraph" style="text-align:left;">Ramp took a Qwen 3.6 model small enough to run on my laptop, <a class="link" href="https://x.com/RampLabs/article/2052447438795833506?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">and it </a><a class="link" href="https://x.com/RampLabs/article/2052447438795833506?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">achieved frontier-level performance</a>, well ahead of Opus 4.6, while offering Claude Haiku-level latency.</p></li></ul><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/5422fd04-645f-463a-adc3-238a04045370/image.png?t=1782206655"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://x.com/RampLabs/article/2052447438795833506?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><ul><li><p class="paragraph" style="text-align:left;"><a class="link" href="https://trajectory.ai/field-notes/harvey-nemotron-3-ultra?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">Harvey trained a Nemotron 3 Ultra model from NVIDIA</a>, bringing it to Opus 4.6-level performance in just 24 hours for a few hundred bucks, for 50th the cost of running the frontier model.</p></li></ul><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/452ce22d-5376-4de7-aec6-64c4104d07ab/image.png?t=1782206760"/></div><p class="paragraph" style="text-align:left;">Why this happens is quite simple to understand with a human analogy. No matter how much your junior employee knows, she can have 4 STEM degrees for that matter, <b>you’re still going to run her through on-the-job training for a few weeks or more</b>.</p><p class="paragraph" style="text-align:left;">Without it, all that broad knowledge is useless, and contextual and in-company knowledge are essential to the tasks at hand.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Basically, with the sole exception of coding agents, which clearly have product-market fit for companies, the incentives are misaligned, <b>and what enterprises really need isn&#39;t what OpenAI or Anthropic are building to raise their valuations.</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://youtu.be/v4GN1q7HX1Y?si=gU9Gb9tnckMiiPQh&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">In a recent interview</a>, Nikesh Arora, Palo Alto Networks’ CEO, echoed a similar idea, <b>stating that consumer users have a much higher tolerance for false positives</b> (e.g., hallucinations) because questions are broad and only need to be broadly helpful. <b>But companies need things to be executed perfectly, something current AIs are NOT built for</b>.</p><p class="paragraph" style="text-align:left;">Put another way, two crucial factors are at play here:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Despite training the AI on task data, which is the obvious unlocker of real performance, <b>AI Labs offer closed-source models you can’t see or adapt, preventing you from owning and controlling what they learn</b>. They are even <a class="link" href="https://community.openai.com/t/openai-is-winding-down-the-fine-tuning-api-and-platform-discussion-thread/1380522?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">shutting down the few fine-tuning capabilities they offer</a>.</p></li></ol><p class="paragraph" style="text-align:left;">Instead, you’re stuck with a model that, to close the knowledge gap it has about your data, has to be very large and way more knowledgeable about the world than you actually need.</p><p class="paragraph" style="text-align:left;">In layman’s terms, to deal with the fact that you can’t train the model, <b>you’re hiring an overqualified individual who might not need any training but costs way more than hiring a less qualified individual and just training them on the job</b>.</p><ol start="2"><li><p class="paragraph" style="text-align:left;"><b>Enterprises want depth, not breadth</b>. They need the model to do the task at hand well, and they don’t care that the same model can also help you with cake recipes. Having a more constrained required response distribution (i.e., I want the model to be good at this one task) allows you to get away with much smaller models that become great at that one task via training.</p></li></ol><p class="paragraph" style="text-align:left;">What companies are, unknowingly, “asking for” is a <b>relatively small, affordable base model they can train on each individual task and optimize for </b><b>that task</b>, and a model that can be safely stored within your organization.</p><p class="paragraph" style="text-align:left;">Thinking Machines Labs is a perfect example of a Lab that ‘gets it’. From the very beginning, they centered their business on <b>unlocking enterprise adoption by reducing the complexity of training models on their data</b>.</p><p class="paragraph" style="text-align:left;">This is terrible news for Frontier AI Labs, <b>which wished you would simply just outsource your entire AI stack to them</b>, having zero control over your models and overpaying for tokens like groupies at a Bad Bunny concert.</p><p class="paragraph" style="text-align:left;">They really have little option; <b>their entire business relies on this being true</b>. But I just don’t know how it could be true.</p><p class="paragraph" style="text-align:left;"><i>Why pay 50 times more for perhaps even worse performance?</i> Enterprises may not understand AI, but they do understand budgets, financial KPIs, and business cases, and none of this makes sense if you’re relying on Frontier Lab tokens for everything.</p><p class="paragraph" style="text-align:left;">As an example, hundreds of models have been released since then, and I still use Gemini 3 Flash and Gemini 3.1 Flash Lite to extract invoices, <b>yet neither would even come close to the top 50 models today</b>.</p><p class="paragraph" style="text-align:left;">But they work just fine for that task and at a great price, <i>so why would I change them unless it was</i><i> for an even cheaper model?</i></p><p class="paragraph" style="text-align:left;">Soon enough, I believe most companies won’t need to chase the latest model for most tasks and will prefer cost-effective older models.</p><p class="paragraph" style="text-align:left;"><i>And can’t the Frontier model Labs simply drop costs and compete?</i> <a class="link" href="https://www.wsj.com/tech/ai/openai-considers-drastic-price-cuts-anticipating-war-for-users-with-anthropic-9b8c178e?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this" target="_blank" rel="noopener noreferrer nofollow">They can</a>, if they’re willing to burn even more of the money they are already burning, <b>but not as a sustainable business strategy.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">All things considered, <i>how should we expect the AI market to evolve over the next few years?</i></p><p class="paragraph" style="text-align:left;">I’m sorry, but I only see it one way: <b>a very strong pivot toward enterprise-sovereign AI,</b> AI that belongs to the company, not to outsiders. An AI that can be controlled, governed, optimized, and updated at the enterprise’s pleasure, not because OpenAI decides to sunset the model to make room for newer ones.</p><p class="paragraph" style="text-align:left;"><i>And what are the implications?</i> Put simply, <b>I don’t view AI revenues for frontier labs as sustainable under the current direction</b>. Tokens aren’t getting cheaper; quite the contrary, and instead of closing the intelligence-per-cost gap with China, it’s only getting worse.</p><p class="paragraph" style="text-align:left;">So if revenues are eventually confirmed to be less promising than we had hoped, and we aren’t at hundreds of billions in yearly revenue by the end of the year, <i>then what?</i></p><p class="paragraph" style="text-align:left;">In that case, <b>we need liquidity from elsewhere</b>. And that elsewhere is precisely one of the biggest sources of concern for this entire industry. And let me tell you, <b>you aren’t going to like what I’m about to show you.</b></p></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-need-to-talk-about-this">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=119fba8c-e813-4051-90d4-2b10aed3b5a6&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>China&#39;s First Frontier Model?</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ac08e9dc-b0bd-436f-80fd-c5f352950cb3/image.png" length="656446" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/china-s-first-frontier-model</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/china-s-first-frontier-model</guid>
  <pubDate>Wed, 17 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-17T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;">Welcome! Today, we discuss China’s latest model, GLM-5.2, which is possibly China’s first frontier model, as well as many other news related to SpaceX, Anthropic, Google, SK Hynix, Microsoft, and more.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>The Time has Come</h2><p class="paragraph" style="text-align:left;">Zhipu, one of China’s top AI Labs, <a class="link" href="https://z.ai/blog/glm-5.2?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">has released a model that might set a precedent in AI</a>, as it’s the first Chinese AI model, to my knowledge, <b>to be actually competitive in raw performance with top US models.</b></p><p class="paragraph" style="text-align:left;">Yes, I’m not saying intelligence-per-cost, I’m saying raw performance. Besides being better than anything Google has ever released for Large Language Models (LLMs), at least on benchmark metrics, <b>it beats GPT-5.5 on several benchmarks</b>.</p><p class="paragraph" style="text-align:left;">Previous models like DeepSeek v4 showed promise, but they are unequivocally behind in most benchmarks. However, the latest string of Chinese AIs, models like this one or Kimi K2.7 Code, are putting up a very serious fight in raw performance <b>while remaining overwhelmingly superior on a cost-efficient basis</b>.</p><p class="paragraph" style="text-align:left;">Architecturally, the model is very similar to everything we’ve seen before. But just like any other Chinese model, <b>they use sparse attention</b>, particularly the same DeepSeek uses (appropriately named DeepSeek Sparse Attention).</p><p class="paragraph" style="text-align:left;"><i>But what does that mean?</i></p><p class="paragraph" style="text-align:left;">Most LLMs today work the same way; they take in a sequence of words and output the next. To do this, <b>every word looks back at previous words looking for interesting attributes to attend to</b> (e.g., in “The green cup”, ‘cup’ attends to ‘green’ to gain the attribute of “greenness”).</p><p class="paragraph" style="text-align:left;">The question here is the following: <i>should every word attend to all previous words?</i></p><p class="paragraph" style="text-align:left;">For example, is the sequence <i>“The green, uhm, ehhh, mmm, ah yes, cup, was…”</i> does ‘cup’ have to attend to ‘uhm or ‘ehhh’, <i>or should it simply focus on ‘green’?</i></p><p class="paragraph" style="text-align:left;">In dense attention, what US models mostly do, we perform dense attention; <b>each word pays attention to all, without distinction</b>. This ensures that all relationships are found, but at the cost of a very large amount of required compute.</p><p class="paragraph" style="text-align:left;">Sparse attention mechanisms actually ask that question first, whether all words matter, preidentify good candidates, <b>and only then pay attention to those selected</b>. This takes the form of an indexer that selects good candidate words that a particular word can attend to and sets a limit. If the indexer only lets you attend to 20 words and you have 1,000 previous words, only 20 will be attended to.</p><p class="paragraph" style="text-align:left;">This allows the compute required to perform attention to remain constant across sequence length for every word <b>with only a slight performance reduction</b>, making it an irresistible alternative for Chinese Labs, which aren’t nearly as compute-rich as US Labs.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">If you want to better understand this indexer mechanism and the overall functioning of sparse attention (and attention in general), <a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/deepseek-is-finally-back-solving-sparse-attention-d2ebbcb8ecb2?sk=6bbf2467f61303a14244ee7eec2b304d&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">I wrote a very detailed description here</a>.</p><p class="paragraph" style="text-align:left;"><i>But what’s the main takeaway here?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I keep hearing delusional takes such as “<i>China is 1 year behind.”</i> I think this model settles this as unfathomably wrong. But social media narrative is always about extremes, so either you’re in that camp, or you’re in the “China has caught up” one, which isn’t true either.</p><p class="paragraph" style="text-align:left;"><b>The key difference is generalization</b>. Chinese Labs, considerably less compute-rich than US Labs, <b>can’t compete on general capabilities with US models</b>; they have smaller models and much smaller training budgets.</p><p class="paragraph" style="text-align:left;">Instead, they focus on particular domains (mostly coding) to be competitive on specific, high-value tasks. <a class="link" href="https://x.com/nikhilchandak29/status/2066970561511657913?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">A simple example of this problem can be seen below</a>, a benchmark, FutureSim, that measures how well models can forecast future events that occurred after their training using real news articles and such. The idea is to test whether models can work well with new information, basically.</p><p class="paragraph" style="text-align:left;">And in such instances, those that require models to work with truly unknown data, <b>the performance gap between closed models and open models is gigantic</b>, which proves that Chinese models are reaching excellent performance on specific domains, but have glaring deficiencies overall relative to US top models.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ac08e9dc-b0bd-436f-80fd-c5f352950cb3/image.png?t=1781695382"/></div><p class="paragraph" style="text-align:left;">But here’s the thing:<b> it doesn’t matter whether China is catching up in overall capabilities</b>.</p><p class="paragraph" style="text-align:left;">What matters is whether <b>Chinese models are reaching a capability threshold that allows them to be used</b>, while costing 10-60 times less.</p><p class="paragraph" style="text-align:left;">I’ve said it in the past, and I’ll say it again,<b> enterprise workflows are meant for depth, not breadth</b>. They yearn for specialized models and don’t give a dime whether the AI model is good at tasks outside the task at hand.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And they’ve hit that threshold in many areas, <b>making them the primary option for cost-effective deployments</b>. This is terrible news for US interests; the way the US has let China dominate the Pareto frontier is unjustifiable.</p><p class="paragraph" style="text-align:left;">I firmly believe that commodity tokens will represent at least 80% of the total tokens generated in a few years. And right now, <b>that means 80% of the world’s tokens will be coming from Chinese models</b>. </p><p class="paragraph" style="text-align:left;">As for local inference, <i>is this model a good option? </i>Not at all.</p><p class="paragraph" style="text-align:left;">The problem here is the KV Cache, the working memory the model uses to respond to specific requests.</p><p class="paragraph" style="text-align:left;">Unlike DeepSeek v4, <b>GLM-5.2 doesn&#39;t compress this working memory</b>. The model is much more opinionated about what parts to use each time (attention), but still stores the entire thing.</p><p class="paragraph" style="text-align:left;">Think of this as still having to remember everything, but for any particular task, use selected parts of your working memory, making every thought faster. However, you’re still having to store “everything”.</p><p class="paragraph" style="text-align:left;">Consequently, a single one-million-token sequence on GLM-5.2<a class="link" href="https://kvcache.ai/tools/kv-cache-calculator/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow"> requires an additional 92 GB of DRAM</a> to store the model&#39;s weights. The model is also BF16, <b>so 1.4 TBs are needed.</b><br><br>In comparison, <a class="link" href="https://kvcache.ai/tools/kv-cache-calculator/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">DeepSeek v4 requires just 4 GBs</a>, because DeepSeek models do, in fact, compress memory (e.g., they decide what has to be remembered and what can be forgotten).</p><p class="paragraph" style="text-align:left;">Still incredible progress by China, once again.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>CYBERSECURITY</b></span><br>Are Mythos-level Cyber Capabilities Overhyped?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ea3cbd99-41d5-4ec2-9104-89876949d19d/image.png?t=1781601627"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://epoch.ai/gradient-updates/are-mythos-cyber-capabilities-overhyped?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">New work by EpochAI</a> has tested whether the risks associated with the new generation of AI models are overhyped. In short, <b>they seem better at exploiting vulnerabilities but not at finding them.</b></p><p class="paragraph" style="text-align:left;">Mythos Preview (the unchained, not broadly available version of the only available-to-US-born Fable) appears to be a major advance in <b>exploit development</b>: its Cyber-ECI benchmark aggregation puts it about <b>7 months ahead of the trend,</b> and clearly above GPT-5.5, especially once newer, less-saturated benchmarks are included.</p><p class="paragraph" style="text-align:left;">The evidence is less clear for <b>vulnerability discovery</b>. Epoch notes a large spike in publicly recorded high and critical CVEs among 21 organizations after Mythos Preview’s release, <b>but says that may partly reflect Project Glasswing’s large increase in spending and access, not only better model capability.</b></p><p class="paragraph" style="text-align:left;">Reports from partners were mixed but generally positive:</p><ul><li><p class="paragraph" style="text-align:left;"><b>Mozilla</b> compared Mythos Preview to elite security researchers;</p></li><li><p class="paragraph" style="text-align:left;"><b>Palo Alto</b> Networks said frontier models produced the equivalent of a year of penetration testing in under three weeks;</p></li><li><p class="paragraph" style="text-align:left;"><b>AWS </b>said the model helped identify additional hardening opportunities.</p></li><li><p class="paragraph" style="text-align:left;">But <b><i>curl</i></b>’s lead maintainer said he saw no evidence that Mythos found issues at a more advanced level than prior tools, so the really is still pretty much out on how transformational these models will be for cybersecurity (if ever released).</p></li></ul><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The conclusions seem to be that Mythos isn’t just hype. However, there were definitely way too many bells and whistles. Even Anthropic, <a class="link" href="https://www.anthropic.com/news/fable-mythos-access?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">in its response to the government</a> after they blocked Fable, alleged that GPT-5.5, which is widely available, posed just as much of a threat, so who knows.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Nonetheless, <a class="link" href="https://thewhitebox.beehiiv.com/p/mythos-destroyer-of-worlds-100-chatgpt-more?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">as I’ve mentioned in the past</a>, many of the associated risks that Mythos introduces <a class="link" href="https://aisle.com/blog/ai-cybersecurity-after-mythos-the-jagged-frontier?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=mythos-destroyer-of-worlds-100-chatgpt-more&_bhlid=72281d2a4806eca9360b46a27c3b9905aef00e64" target="_blank" rel="noopener noreferrer nofollow">have already been shown to be possible with current models</a>.</p><p class="paragraph" style="text-align:left;">Which is to say, it’s not like Mythos isn’t better at cybersecurity than most, if not all, available models, <b>but many of the risks that led to its “it’s too dangerous to release” narrative are already here with us.</b></p><p class="paragraph" style="text-align:left;">In short, as several cybersecurity leaders have stated, <a class="link" href="https://www.reuters.com/legal/litigation/cyber-leaders-urge-us-lift-curbs-anthropics-security-models-2026-06-15/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">just give us the models</a>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>COSTS</b></span><br>A move toward measuring tasks, not tokens</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/494f1d47-7761-4e79-9e63-42d3dce2ecdb/image.png?t=1781602542"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://artificialanalysis.ai/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model#price-and-cost" target="_blank" rel="noopener noreferrer nofollow"><b>Artificial Analysis has updated its Intelligence Index to v4.1</b></a>, adding new task-level cost and time metrics alongside model-quality scores.</p><p class="paragraph" style="text-align:left;">The index now measures not only model capability, but also <b>cost per Intelligence Index task</b>, <b>time per task</b>, and <b>token use per task</b>.</p><p class="paragraph" style="text-align:left;">The obvious highlight is clearly the cost difference. While <b>DeepSeek V4 Pro Max</b> reaches an Intelligence Index of <b>44</b> at about <b>$0.06 per task</b>, GPT-5.5 on high reasoning mode costs <b>$0.99 per task</b> and <b>$1.78 for Claude Opus 4.8 max</b>, with Fable 5 setting a new high bar in terms of performance (topping the index at 60) and cost, <b>being 54 times more expensive per task than DeepSeek V4 Pro Max</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I don’t know who needs to hear this, but if you’ll excuse my words, <b>US Labs need to wake the fuck up</b>. It’s great that Fable 5 tops benchmarks (if we could use them, of course), <b>but is the model 54 times better than DeepSeek on average tasks</b><b>?</b></p><p class="paragraph" style="text-align:left;">Come on.</p><p class="paragraph" style="text-align:left;">Yes, it may give you an edge on some tasks DeepSeek’s models can’t solve yet, and for those, you will default to the frontier, but what I believe most people in the US resist understanding is that <b>those tasks are a tiny, tiny part of what the world will ask of AI.</b></p><p class="paragraph" style="text-align:left;">A business is about solving a problem for the user, and the user wants it solved at the lowest possible cost. Nobody drives a Ferrari to go to work in the fields because a 20-year-old crappy Toyota does that job brilliantly.</p><p class="paragraph" style="text-align:left;">If Chinese models are consistently the best intelligence-per-cost option,<b> the world will run on Chinese models</b>. Your choice, San Francisco.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>COMPUTE</b></span><br>All Signs Continue to Point in the Same Direction: Compute</h2><p class="paragraph" style="text-align:left;">In the days after Fable&#39;s release, before it was banned at least, I noticed a pattern: <b>progress still seems overwhelmingly linked to compute</b>, and two recent evaluations prove this very clearly.</p><p class="paragraph" style="text-align:left;">On the one hand, a clearly counterintuitive result by <a class="link" href="http://Vals.ai?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">Vals.ai</a>. Despite Claude Fable 5 fallbacking to Opus 4.8 199/200 times, meaning the request was sent to Fable 5, but the system downgraded the request to Opus 4.8 because the task was flagged as “unacceptable” by Anthropic and thus doesn’t want you to use the best model for, meaning it was essentially Opus 4.8 responding the entire evaluation, <b>it still got twice the marks than running the task directly on Opus 4.8.</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6576beec-dc8c-46cf-aca2-7344a8f4914e/image.png?t=1781603425"/></div><p class="paragraph" style="text-align:left;"><i>How is that possible?</i></p><p class="paragraph" style="text-align:left;">Most interestingly, <b>the cost was also double</b>, even though both were charged at Opus 4.8 tokens. <b>This means the first result generated twice as many tokens as the latter</b>, proving that inference-time compute, just deploying more compute into the task, shows no signs of slowing at all.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Nonetheless, OpenAI’s Reasoning Lead, Noam Brown, <a class="link" href="https://x.com/polynoamial/status/2064210146558136827?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">had a really good article explaining just this</a>, stating, “<b>As LLMs become more capable, benchmark performance is increasingly a function of test-time compute,” </b>arguing that AI products are still not showing AI’s real capabilities because we’re forced to clamp down on how much they can actually think on the problem due to costs.</p><p class="paragraph" style="text-align:left;">Another obvious example is the ECI index from EpochAI, which shows Fable 5 scoring just 1 point higher than GPT-5.5 Pro. And you may ask, how does GPT-5.5, an undeniably inferior model, <i>get such a close result to the next-generation model?</i></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1fa737c5-b3e6-45d5-93ec-356e1a6b7060/image.png?t=1781604132"/></div><p class="paragraph" style="text-align:left;">And the answer is that this is apples-to-oranges because Fable 5 is one model, <b>GPT-5.5 Pro is a best-of-N sampling method</b> where OpenAI deploys several separate AIs into the task and keeps the best result of the bunch (usually the most common one).</p><p class="paragraph" style="text-align:left;">This proves that, despite the individual model being worse, the additional inference-time compute required to deploy several AIs closes the gap.</p><p class="paragraph" style="text-align:left;">OpenRouter’s Panels API, <a class="link" href="https://thewhitebox.beehiiv.com/p/why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">which we discussed last Sunday</a>, is an outcome of the same idea: by deploying several agents to a problem, <b>the combined computational power exceeds that of a single model, yielding an outsized outcome</b>.</p><p class="paragraph" style="text-align:left;"><i>The implication of this?</i> As I said in the last newsletter, blocking a specific model is useless as the genie is pretty much out of the bottle.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Placing model restrictions doesn’t make any sense at all because worse models with higher compute thresholds reach the same performance. Consequently, if you really really wanted to halt progress,<b> you have to go for the compute.</b></p><p class="paragraph" style="text-align:left;"><b>Anthropic’s entire framing of going after the models is just regulatory capture</b>, because Chinese models commoditize its business badly and want them out of the way.</p><p class="paragraph" style="text-align:left;"><b>The US advantage is a compute advantage</b>; I’ve always said this and will continue to push it. However, this must not lead to banning US compute for the rest of the world, because if you do so, China is going to flood the world with Chinese compute, <b>and I challenge you to give me a single reason the USG would want that</b>.</p><p class="paragraph" style="text-align:left;">Instead, if you really want the US to thrive, you should acknowledge the importance of compute and, instead of placing export controls, <b>focus all your efforts on reducing capital costs for businesses that deploy it</b>. That is where China beats the US: in getting that computer deployed at a better cost.</p><p class="paragraph" style="text-align:left;">Interestingly, export controls only make the customer pool smaller, reducing demand, and thus pushing prices higher.</p><p class="paragraph" style="text-align:left;">Instead, give companies in the industry energy credits and provide liquidity for construction; <b>do whatever it takes to bring down the $50 billion-per-gigawatt cost, which is, in fact, not falling but is currently rising</b>.</p><p class="paragraph" style="text-align:left;">That is how the US should compete (and most likely win).</p><p class="paragraph" style="text-align:left;">I know I’ve said multiple times that we shouldn’t be bailing out businesses that are overspending. But China is doing so, <b>so if you really feel AI is a true national security risk, a race between both countries, then you’ll have to</b>.</p><p class="paragraph" style="text-align:left;">But let me be clear, the focus should not be on models; <b>it should be on computation</b>, <a class="link" href="https://thewhitebox.beehiiv.com/p/clarifying-myths-in-the-us-vs-china-war?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">which, by the way, is exactly what the CCP is doing</a>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ROBOTICS</b></span><br>Using Autoresearch for robotics</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3a0cdedc-fe7d-4314-ade5-40651c8f64c4/image.png?t=1781702003"/></div><p class="paragraph" style="text-align:left;"><b>One of the most popular themes in AI today is recursive self-improvement</b>, the idea of letting models improve themselves or improve other models. In an ideal world, AI will be in charge of autonomously improving itself, potentially creating an explosion in progress. At least, that’s the hope.</p><p class="paragraph" style="text-align:left;">And now, NVIDIA has applied it to robotics… <a class="link" href="https://research.nvidia.com/labs/gear/enpire/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">and it worked</a>.</p><p class="paragraph" style="text-align:left;">Called <b>ENPIRE</b>, it’s a harness framework for coding agents that instantiates a physical feedback routine in which coding agents analyze logs, consult the literature, improve training infrastructure, and refine algorithm code to address failure modes, <b>thereby improving</b><b> robotic arms</b> (the previous link includes video demonstrations).</p><p class="paragraph" style="text-align:left;">As the blog states, <i>“Powered by ENPIRE, frontier coding agents can autonomously develop a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks in the real world, such as PushT, organizing pins into a pin box, and using a cutter to cut a zip tie.”</i></p><p class="paragraph" style="text-align:left;">As seen in the image above, having Codex try different ways at improving performance autonomously leads to clear improvements.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;"><b>The real question here is, of course, cost</b>. Having an AI agent (or several of them) run endlessly to solve the problem can lead to enormous per-trial costs (don’t forget that an OpenAI engineer spent $1.3 million in API credits in a single month).</p><p class="paragraph" style="text-align:left;">The other big question here is what the upper bound of these runs is. It’s clear that the ability of AI models to improve scientific discovery is bound by their “intuition”; the ability to suggest novel ideas that don’t overlap with previous ones and genuinely improve the ability to test different stuff.</p><p class="paragraph" style="text-align:left;">Recent mathematical breakthroughs, like those in the Erdos problems, <b>suggest these models are</b>, in fact, <b>improving their ability to be creative in their suggestions</b>, but I’m skeptical as to how truly novel these are.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>M&A</b></span><br>SpaceX Acquires Anysphere</h2><p class="paragraph" style="text-align:left;">With a confirmation that surprised no one,<a class="link" href="https://www.reuters.com/legal/transactional/spacex-buy-anysphere-60-billion-2026-06-16/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow"> SpaceX has confirmed the acquisition of Anysphere</a>, the team behind the popular coding tool Cursor, in a <b>$60 billion stock-based merger</b>. The deal is expected to close in <b>Q3 2026</b>.</p><p class="paragraph" style="text-align:left;">The deal puts Cursor’s ARR at $2 billion, adding to the fast-growing SpaceX AI revenue, which is much needed to justify considering <b>its current trading above Amazon at the time of writing</b>, with Amazon at almost $3 trillion. Alan Greenspan surely “loves” our current market.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">We can argue if the price was too high,<b> but we can’t argue that the acquisition makes total sense for SpaceX</b>.</p><p class="paragraph" style="text-align:left;">They acquire a super-talented team behind what’s possibly <b>the best agent coding harness there is right now</b> (i.e., Claude/GPT models run better in Cursor than in their own commercial products), and a team that is also training its own models, initially fine-tuning Chinese models, and currently training from scratch their own, <a class="link" href="https://x.com/morganlinton/status/2066946225837109735?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">as announced by its CEO here</a>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Additionally, <b>they gain &gt;$2 billion in much-needed revenues and swaths of human coding data</b>. I see no issues with the acquisition, except that paying 30 times revenue isn’t precisely cheap.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>REGULATION</b></span><br>DeepSeek, To the Ban List?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.reuters.com/world/china/us-holds-off-blacklisting-chinas-deepseek-more-than-100-firms-deemed-security-2026-06-17/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://www.reuters.com/world/china/us-holds-off-blacklisting-chinas-deepseek-more-than-100-firms-deemed-security-2026-06-17/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">reported by Reuters</a>, the US has held off on adding China’s DeepSeek, CXMT, and more than 100 other firms to the Commerce Department’s Entity List, despite an interagency committee reportedly deeming them national security risks. <b>The delay comes as the Trump administration seeks to avoid worsening tensions with Beijing</b>.</p><p class="paragraph" style="text-align:left;">Reuters reports that the Entity List, which restricts exports of US goods, software, and technology, <b>has not received new additions since October</b>, the longest gap in more than a decade.</p><p class="paragraph" style="text-align:left;">Some of the pending companies were reportedly linked to supplying Russian drones recovered in Poland, selling restricted Nvidia chips to Chinese universities, or supporting China’s military-related drone and robot-dog programs.</p><p class="paragraph" style="text-align:left;">China’s foreign ministry said the US should stop “politicizing” economic and technology issues, while the Commerce Department’s Bureau of Industry and Security said it continues using export-control tools to address “bad actors.”</p><p class="paragraph" style="text-align:left;"><i>Is this the beginning of the end for Chinese open models being legal in the US?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">In a basic scenario where DeepSeek posts weights publicly, and anyone can download them for free, <b>the Entity List alone would not clearly make downloading or using those weights illegal</b>.</p><p class="paragraph" style="text-align:left;">The risk increases if the download involves a <b>transaction with DeepSeek</b>, such as signing a license directly with the company, paying for access, or using its hosted API. Those could be treated differently from the way publicly available files are handled.</p><p class="paragraph" style="text-align:left;">However, most people using Chinese models today access them through US LLM providers like Fireworks or the Hyperscalers, meaning the models are hosted by US entities that have entered into no direct agreements with DeepSeek.</p><p class="paragraph" style="text-align:left;">Banning open models would require treating digital files as contraband, which would be hilarious<b>, since it would essentially be a ban on a file packed with matrix multiplications.</b></p><p class="paragraph" style="text-align:left;">Jokes aside, <b>open models keep token prices in check</b>. Without them, you’re left with Anthropic, OpenAI, Google, and a handful of remaining players <b>who can then raise prices with no risk of commoditization at all.</b></p><p class="paragraph" style="text-align:left;">If AI truly becomes ingrained in enterprise workflows, it would increase operational costs significantly relative to foreign companies, which would, of course, rely on cheaper tokens.</p><p class="paragraph" style="text-align:left;">The US needs to stop thinking of banning stuff as a solution to its competitiveness and start competing;<b> force US Labs to compete at the Pareto frontier, not give them the victory via regulatory capture</b>. I’m not saying it’s going to happen, I’m just hoping it doesn’t.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MEMORY</b></span><br>Hynix ADR Listing Soon?</h2><p class="paragraph" style="text-align:left;">SK Hynix, one of the key suppliers of DRAM (the memory used by AI accelerators like GPUs), <a class="link" href="https://www.hankyung.com/article/202606169905i?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">is allegedly preparing to list on the American stock market (via an ADR) as soon as next month</a>, opening the door for US investors to invest in the company.</p><p class="paragraph" style="text-align:left;">Right now, investing in Hynix is complicated because investing in Korean companies is very hard for non-Koreans, with only a couple of brokers like Interactive Brokers offering them. Therefore, this would represent a considerable increase in liquidity for the company.</p><p class="paragraph" style="text-align:left;">Also, as part of a plan to increase shareholder returns, they will begin repurchasing shares and paying cash dividends in the fourth quarter.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">While memory is one of the most important elements in the AI supply chain today, the market&#39;s cyclicality risks (memory has traditionally been brutally cyclical) <b>may lead Hynix to act more cautiously</b> and, instead of fully reinvesting to extend capacity (they are expanding capacity, though clearly not as much as they could), use that cash to inflate stock value and shareholder returns.</p><p class="paragraph" style="text-align:left;">It’s not the greatest sign for people looking to hold this stock for years, but it’s great news for current shareholders.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ENTERPRISE</b></span><br>Microsoft is considering using DeepSeek for Copilot</h2><p class="paragraph" style="text-align:left;">In quite an unexpected turn of events, <a class="link" href="https://www.axios.com/2026/06/16/microsoft-copilot-cowork-tokenmaxxing-cowork?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">it seems Microsoft is seriously considering using DeepSeek v4 Flash as one of the underlying models in its Copilot offering</a>, aiming to reduce serving costs, with the idea of offering a more competitive Copilot pricing too, a matter of much discussion in recent times after the massive June 1st price hikes, <a class="link" href="https://www.tomshardware.com/tech-industry/artificial-intelligence/github-copilot-customers-suffer-from-sticker-shock-as-microsoft-switches-to-usage-based-pricing-customers-report-up-to-100-fold-price-hikes?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">with some customers allegedly spending up to 100 times what they did before</a>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">If regulatory capture doesn’t prevent it, <b>this will become the norm unless US Labs get serious about offering good intelligence-per-cost models</b>.</p><p class="paragraph" style="text-align:left;"><i>But do you realize how tragic this is?</i></p><p class="paragraph" style="text-align:left;">Microsoft, which basically owns OpenAI and has access to its IP, decides to use a Chinese model because US models are simply too expensive to run.</p><p class="paragraph" style="text-align:left;">I would be surprised if not for the fact that we’ve been saying this would happen for a long time in this newsletter, <b>but one can’t help but feel “impressed” by how easy US Labs are conceding defeat in the commodity token market</b>.</p><p class="paragraph" style="text-align:left;">Nonetheless, GLM-5.2, the model discussed above, scores 11 points higher on the Artificial Analysis benchmark than GPT-5.4 mini (on high reasoning), despite costing less per task.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1412b167-3bca-4c82-879d-3562d6346757/image.png?t=1781691104"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://artificialanalysis.ai/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model#intelligence" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;">Outrageous and a complete strategic defeat for Frontier Labs. As mentioned earlier, <b>US Labs need to wake up</b>. And fast.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ENTERPRISES</b></span><br>The Semantic Layer we were waiting for?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">Google has released the Open Knowledge Format</a>, a proposed standard for representing organizational knowledge in a form that humans, data tools, search systems, and AI agents can read.</p><p class="paragraph" style="text-align:left;">OKF captures context around enterprise data and systems. This includes definitions of tables, metrics, datasets, APIs, business processes, and runbooks, as well as links to authoritative sources and usage notes.</p><p class="paragraph" style="text-align:left;">The format is based on directories of markdown files. Each file describes a single concept and may include structured metadata such as type, title, description, resource links, tags, and timestamps.</p><div class="codeblock"><pre><code>---
type: BigQuery Table
title: Orders
description: One row per completed customer order.
resource: https://console.cloud.google.com/bigquery?p=acme&amp;d=sales&amp;t=orders
tags: [sales, revenue]
timestamp: 2026-05-28T14:30:00Z
---

# Schema

| Column        | Type      | Description                              |
|---------------|-----------|------------------------------------------|
| `order_id`    | STRING    | Globally unique order identifier.        |
| `customer_id` | STRING    | FK to [customers](/tables/customers.md). |

# Joins

Joined with [customers](/tables/customers.md) on `customer_id`.</code></pre></div><p class="paragraph" style="text-align:left;">OKF does not replace databases, schemas, APIs, or data catalogs. Instead, it documents what those systems mean, how assets relate to each other, and what caveats apply. <b>It’s a semantic layer that packages knowledge about enterprise systems into simple files that LLMs can easily interpret.</b></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">It’s all context engineering in the end. <a class="link" href="https://x.com/GoogleCloudTech/status/2067012903337664886?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">As Google explains</a>, <i>“AI is only as smart as the context you give it.” </i>Thus, improving the quality of the provided context is one of the easiest ways to improve performance, <b>with the added benefit that here, Google is standardizing how this semantic layer is built</b>.</p><p class="paragraph" style="text-align:left;"><i>But you may be asking? Isn’t this Anthropic’s skills all over again? Or the memory systems these Labs use? Isn’t it just a bunch of markdown files like the other context engineering methods?</i></p><p class="paragraph" style="text-align:left;">Kind of, but not quite.</p><p class="paragraph" style="text-align:left;"><b>Anthropic’s Skills package behavior</b>, how to use a certain tool. <b>OKF packages knowledge</b>. It’s all the same in essence, instructions in a simple-to-read file, but it’s important that we understand when to use what.</p><p class="paragraph" style="text-align:left;">If you’re up for a read, it’s heavily inspired by <a class="link" href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">Andrej Karpathy’s LLM-wiki</a>, a persistent wiki available to the model about something, in this case, enterprise systems, <a class="link" href="https://openai.com/index/chatgpt-memory-dreaming/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">an idea quite similar to what OpenAI does with its memory systems</a>.</p><p class="paragraph" style="text-align:left;">But OKF is the first one to try to standardize this process, which is welcomed.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>CONSUMER</b></span><br>AI AirPods in 2027?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.bloomberg.com/news/articles/2026-06-16/apple-plans-camera-airpods-iphone-foldable-2-20th-anniversary-iphone-in-2027?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc4MTY3NzEwOCwiZXhwIjoxNzgyMjgxOTA4LCJhcnRpY2xlSWQiOiJUR0NERkhLSVAzTVIwMCIsImJjb25uZWN0SWQiOiJCMzZENUE5QzIxMDQ0NjU4OTFBMTc1MTVDRDNBQkZFNiJ9.F0ZciaQ3s5fND3vhUcvOpNbWJhtmsmWnmM9ExA4Q2oc&utm_source=tldrnewsletter&leadSource=uverify%20wall" target="_blank" rel="noopener noreferrer nofollow">According to Mark Gurman</a>, a really well-known journalist who is almost exclusively focused on Apple, <b>Apple will launch AI AirPods in 2027</b>, as well as other products.</p><p class="paragraph" style="text-align:left;">The lineup is expected to include camera-equipped AirPods, a second-generation foldable iPhone, and a redesigned iPhone tied to the product’s 20th anniversary.</p><p class="paragraph" style="text-align:left;">The camera-equipped AirPods are reportedly scheduled for late 2027. <b>The cameras are not intended for taking photos or videos</b>, but for providing Siri and Apple’s AI systems with visual context of the user’s surroundings.</p><p class="paragraph" style="text-align:left;">Apple is also planning a second foldable iPhone for 2027, <b>following the expected launch of its first foldable model.</b> The anniversary iPhone is expected to feature a new design with a near-edge-to-edge display and curved glass on the sides.</p><p class="paragraph" style="text-align:left;">The products remain in development, and Apple has not publicly announced the devices.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">It’s obvious Apple is going to focus on winning the consumer AI market, and if Siri actually delivers in September, they have a great chance of running away with it, <b>unless OpenAI does something really transformational</b> about its consumer device being developed by Jony Ive.</p><p class="paragraph" style="text-align:left;">I don’t own any AI consumer devices yet, like Meta’s RayBan/Oakley Meta glasses, <b>but they are starting to become quite appealing to me</b>.</p><p class="paragraph" style="text-align:left;">My only concern is that people may feel uneasy about you having a device that is literally recording them, <b>so Apple’s AI AirPods, which focus more on surroundings and providing better context, might be less intrusive</b>. We’ll see.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">We’re finally witnessing the emergence of China as a clear rival. And it’s a brutal one with costs that are simply too low to ignore.</p><p class="paragraph" style="text-align:left;">Chinese Labs are becoming incredibly sophisticated, but I don’t think anyone is particularly surprised at this,<b> considering one-third of “US researchers” are actually Chinese.</b></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.nytimes.com/2025/11/19/technology/ai-research-chinese-talent.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-s-first-frontier-model" target="_blank" rel="noopener noreferrer nofollow">The New York Times already highlighted that talent gap months ago</a>, and I believe they have older articles about it.</p><p class="paragraph" style="text-align:left;">US policy should be pressuring Frontier AI Labs to compete at the commodity token market. To me, this strikes as obvious, but for whatever reason, only Google seemed to understand that, only to then release Gemini 3.5 Flash at three times the original price of the third version.</p><p class="paragraph" style="text-align:left;"><i>Result? </i>Nobody uses that model, and the question of which model offers the best intelligence per cost is analogous to which Chinese model offers the best intelligence per cost.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Sadly, US Labs can’t compete on price with Chinese labs. And no, it’s not because Chinese model architectures are more frugal or that Chinese researchers are marked by the mandate of heaven. Sparse attention is something any researcher sees as a great compromise.</p><p class="paragraph" style="text-align:left;">No,<b> it’s much simpler: capital costs.</b></p><p class="paragraph" style="text-align:left;">When Anthropic serves you a token, they aren’t only paying for the operational cost of serving it; t<b>hey are also paying the cost required to build the data center in the first place</b>, and that’s the big one, actually.</p><p class="paragraph" style="text-align:left;">I really give US Labs a hard time, but I acknowledge it’s partially not their fault that, to compete, they have to pay $50-$100 billion per gigawatt. <b>It’s not normal for Meta to spend more on AI this year than the German state spends on defense.</b></p><p class="paragraph" style="text-align:left;">This is the product of a highly commoditized software running on top of non-commoditized, and thus very pricey, hardware. Let’s stop pretending this isn’t happening, and let’s start enacting policies that decrease those prices somehow.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=6018fbb8-9a29-4e2b-9e2e-698661256b00&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Why Anthropic’s Mess Doesn’t Matter &amp; The Era of Sophistication</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7b180518-2ff8-4186-9a9b-e0691be184b8/image.png" length="290404" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication</guid>
  <pubDate>Sun, 14 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-14T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Why Anthropic’s Mess Doesn’t Matter & The Era of Sophistication</h2><p class="paragraph" style="text-align:left;">As you surely know, currently, any non-American like me, including immigrants living in the US, which represents a huge portion of Anthropic’s own employees, <b>can’t use Fable/Mythos-class models by order of the United States Government (USG)</b>.</p><p class="paragraph" style="text-align:left;">This unprecedented export control, the first on an actual AI model, has already changed AI, even if they aren’t aware of it yet, forever.</p><p class="paragraph" style="text-align:left;">The consequences for Anthropic, OpenAI, and other suspects of receiving similar treatment are terrible.</p><p class="paragraph" style="text-align:left;">But not because they’ve reduced their addressable market a lot if this doesn’t get overturned, which is pretty bad already, <b>but because I believe it will start the era of sophistication</b>, something Labs really did not expect they would have to deal with.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">The Ephimerous Fable</h2><p class="paragraph" style="text-align:left;">As I wrote about in <a class="link" href="https://thewhitebox.beehiiv.com/p/sabotages-diffusion-and-a-2-trillion-company?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">my last issue of this newsletter</a>, Anthropic has finally revealed its new generation model to the world.</p><p class="paragraph" style="text-align:left;">But the release was a spectacular failure.</p><h3 class="heading" style="text-align:left;">From sabotage to export control</h3><p class="paragraph" style="text-align:left;">Simply put, Fable was the first model you and I had no right to use freely. Depending on the use case, if Anthropic didn’t like what it saw,<b> you would be instantly downgraded to a worse model.</b></p><p class="paragraph" style="text-align:left;">This included topics such as biology, cybersecurity, or general science, <b>areas that Anthropic considers are too dangerous to expose freely because those same topics can lead to very harmful outcomes</b>, like cyberweapons and all that fearmongering crap you can still learn to do without Fable.</p><p class="paragraph" style="text-align:left;">But I digress. For now, all we got was a model incapable of answering questions about mitochondria <a class="link" href="https://x.com/DeryaTR_/status/2064404588476948774?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">or </a><a class="link" href="https://x.com/DeryaTR_/status/2064404588476948774?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">unwilling to provide information on cancer research</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/175d0fa7-a438-4363-a6fc-5d3538721404/image.png?t=1781426872"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://x.com/DeryaTR_/status/2064404588476948774?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;">Then, Dario Amodei, the CEO, <a class="link" href="https://darioamodei.com/post/policy-on-the-ai-exponential?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">published a blog</a> talking about the huge impending risks that were coming with AI, stating, and I quote:</p><p class="paragraph" style="text-align:left;"><i>“Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety.”</i></p><p class="paragraph" style="text-align:left;">And, well, to say that he got what he asked for was an understatement. A few hours later, <b>Fable was banned</b>.</p><p class="paragraph" style="text-align:left;"><i>But why?</i></p><h3 class="heading" style="text-align:left;">The jailbreak that ended it all</h3><p class="paragraph" style="text-align:left;">First and foremost, let me be very clear: <b>I have zero doubt this will be overturned soon</b>; the fact that researchers at Anthropic, like Andrej Karpathy, can’t use the model they created is bonkers, <b>especially given that way more than half of Anthropic’s researchers aren’t American-born or naturalized citizens.</b></p><p class="paragraph" style="text-align:left;">But, anyway, the sequence, as currently reported, went as follows:</p><p class="paragraph" style="text-align:left;">Anthropic’s version is that the government contacted it at 5:21 PM ET on June 12 with an export-control directive barring access to Fable 5 and Mythos 5 by any foreign national, inside or outside the United States, <b>including Anthropic’s own foreign-national employees</b>.</p><p class="paragraph" style="text-align:left;">The government’s stated rationale was national security.</p><p class="paragraph" style="text-align:left;">Anthropic says the letter did not provide specific details, but its understanding was that the government believed there was a way to bypass, or “jailbreak,” Fable 5’s safeguards to use it to identify cybersecurity vulnerabilities.</p><p class="paragraph" style="text-align:left;">From all places, the leak came from Amazon, Anthropic’s largest investor. Amazon researchers reportedly produced a security report <b>claiming they could prompt Anthropic’s Fable 5, or the Mythos capabilities beneath it</b>, to provide cyber-relevant information that was supposed to be restricted.</p><p class="paragraph" style="text-align:left;">Amazon CEO Andy Jassy then raised those concerns with senior Trump administration officials, <a class="link" href="https://www.reuters.com/business/retail-consumer/amazon-voiced-concerns-about-anthropic-ai-models-before-us-governments-crackdown-2026-06-13/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">according to Reuters</a>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/DavidSacks/status/2065853007619588171?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">David Sacks’ framing</a>, an insider to this Administration, is that the export control was issued “reluctantly” after Dario Amodei refused to “fix the jail break or de-deploy the model.”</p><p class="paragraph" style="text-align:left;">Sacks said the administration’s hope is that Anthropic fixes the safety issue, that export controls are lifted, and that Fable returns to general release.</p><p class="paragraph" style="text-align:left;">In other words, the White House line is: Amazon and others surfaced a serious safety issue, Anthropic would not pause or fix it fast enough, so Commerce used export controls as the emergency lever.</p><p class="paragraph" style="text-align:left;">Nobody actually knows who’s telling the truth here. I genuinely don’t care, actually. But boy, <b>was Anthropic begging for this to happen</b>. Your entire marketing proposition can’t be <b><i>“let’s block every powerful model”</i></b><i>… </i>but rant when it’s yours that&#39;s being blocked.</p><p class="paragraph" style="text-align:left;">And to be very clear on my positioning, this is all incredibly stupid and simply a product of Anthropic crying wolf for far too long.</p><p class="paragraph" style="text-align:left;"><b>I don’t believe these models should be export-controlled</b>, and I don’t believe Anthropic’s leadership has the mandate of heaven to decide who gets to use the technology and when.</p><p class="paragraph" style="text-align:left;">I maintain that the safest way to create AI is to do it in the open. Besides, if you think you can block AI progress on the frontier by de-deploying one model, you’re out of your mind because the genie is not only out of the bottle, it has mapped the exits, copied the keys, and taught a thousand others how to escape.</p><p class="paragraph" style="text-align:left;">You can’t regulate yourself out of the “problem” of AI progress. Too late.</p><p class="paragraph" style="text-align:left;">Instead, we must see this for what it is: a company that spent 4 years announcing the end of times because they were creating the anti-christ, <b>and did so well that the USG now believes it’s, in fact, true </b>(when it’s not)<b>. </b>They came looking for regulatory capture to build a regulatory moat around them, and what they got was being regulated themselves.</p><p class="paragraph" style="text-align:left;">May I insist: <b>I do not think Fable should be blocked </b>(nor that it should exist at all; instead, deploy Mythos, the non-guardrailed model), <b>but if you ask to be regulated, you will get regulated.</b></p><p class="paragraph" style="text-align:left;">And if to top it off, your CEO and the Administration clearly hate each other, and well, that is not going to help. Personally, and I’m starting to say this more publicly,<b> this company needs a new CEO once it goes public.</b></p><p class="paragraph" style="text-align:left;">I may not like him, but Dario has done a fricking great job as a CEO in Anthropic’s private era, creating the most valuable private company in history, <b>but a public company CEO is a totally different beast</b>.</p><p class="paragraph" style="text-align:left;">It’s less about vision and ideals, and more about relationships; you have to be extremely likable, adept at uncomfortable conversations, and basically every single trait this guy clearly doesn’t have.</p><p class="paragraph" style="text-align:left;">This is particularly relevant when you consider the fact that, in my view, <b>one of the core customers of the frontier models will be the government</b>. If I had to bet, it would be the largest customer by far, and it wouldn’t be particularly close.</p><p class="paragraph" style="text-align:left;">The reason is that <b>most consumer and enterprise workloads will run on ‘commodity tokens’,</b> tokens from models that maximize “intelligence per dollar”, <b>because most tasks do not require the giga-brain Fable to work well.</b></p><p class="paragraph" style="text-align:left;">Factoring in the reality that the AI business is currently painfully unprofitable and will remain so for years, <b>will make frontier tokens very expensive</b>, probably out of reach for most people (if even available at all, given recent events).</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But let me give you a better explanation, backed by facts and research, of why this doesn’t matter at all.</p><h2 class="heading" style="text-align:left;">The Genie is on a Spaceship</h2><p class="paragraph" style="text-align:left;">Here’s a bold objective: I’m going to convince you that<b> this export control, even if it persisted, doesn’t matter in the slightest</b>, and I, who literally can’t use the model as I write these words, don’t care at all.</p><p class="paragraph" style="text-align:left;">The entire point of this export control, politics aside, is the assumption that if we shut the model down globally, nobody will be able to replicate or access its capabilities.</p><p class="paragraph" style="text-align:left;">But that, my dear reader, is categorically false. And the first proof requires me to acknowledge that I was wrong about some points I made recently.</p><h3 class="heading" style="text-align:left;">China and distillation</h3><p class="paragraph" style="text-align:left;">For quite some time, <b>you’ve read from me several times that China manages to stay close to the frontier by performing distillation</b>: they generate synthetic data from frontier models and train their models on it to improve and get closer to the frontier.</p><p class="paragraph" style="text-align:left;">And while distillation is, in fact, being used (by everyone, not just Chinese Labs, to the point it’s one of the topics Anthropic downgrades you to worse models),<b> I’m beginning to seriously doubt that this is the only explanation for China being so close in performance</b>. As Gen-Z would say, it’s “cope.”</p><p class="paragraph" style="text-align:left;">If we look at the latest stream of Chinese models, <b>Kimi K2.7 Code</b>, <b>GLM-5.2</b>, <b>Minimax M3</b>, as of course <b>DeepSeek V4 Pro</b>, not one of these has a performance level that can only be explained solely by distillation.</p><p class="paragraph" style="text-align:left;">These models are incredibly good and different. They don’t feel like copies of the American frontier; they have their own “personalities” and taste. I don’t really have the right words to explain this, but they just feel “different,”<b> which suggests they can’t possibly be just smaller models trained to imitate US ones</b>.</p><p class="paragraph" style="text-align:left;">I was guilty of making that simplification, but as the saying goes, <i>“To err is human; to correct oneself is wise.”</i></p><p class="paragraph" style="text-align:left;">In many cases, <b>US Labs are the ones leveraging Chinese models</b>. One great example is Cursor, which trained its Composer models by fine-tuning Moonshot’s Kimi K2.5/6 models.</p><p class="paragraph" style="text-align:left;">And although this is not hard proof, <a class="link" href="https://x.com/atomic_chat_hq/status/2065581878279549090?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">but circumstantial evidence</a>, you can already see examples where Chinese models perform better than US ones.<i> And at 3-10 times lower cost!</i></p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/prz_chojecki/status/2065741640635990128?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">Another example is ErdosBench</a>, a new mathematical benchmark to test models&#39; ability to handle complex, rare problems that require deep mathematical reasoning, and Kimi K2.7 Code ranks above all models except Fable.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0afd77a1-85e2-4f23-9854-467e26b7f595/image.png?t=1781458445"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://x.com/prz_chojecki/status/2065741640635990128?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;">The point here is that if you think this will slow down China,<b> I believe that would be completely wrong</b>. Especially when, at the same time,<i> the US blocks the world from using Fable while letting China buy highly competitive NVIDIA/AMD GPUs!</i></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And few people are as guilty as I am for once making that simplification. Moreover, <b>China doesn’t have a secret sauce either</b>, so the fact that their models are smaller, use sparse attention mechanisms, and have tinier training budgets comes at a cost: <b>worse overall performance</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But again, no, there’s no secret sauce. Chinese Labs are running a very similar recipe to US Labs: use Transformers, and progressively train larger models and run them for longer on the task, t<b>he quintessential architecture combined with the quintessential scaling laws.</b></p><p class="paragraph" style="text-align:left;">The recipe is mostly identical, and there’s no magical breakthrough.</p><p class="paragraph" style="text-align:left;">If anything, they could have a data difference, <b>but US Labs access to data is orders of magnitude higher </b>(because data is expensive to gather and these Labs have way more of it).</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://thewhitebox.beehiiv.com/p/clarifying-myths-in-the-us-vs-china-war?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication" target="_blank" rel="noopener noreferrer nofollow">I already explained this</a>, <b>but China’s real advantage is cost per watt</b>; they can deploy computing way cheaper than the US.</p><p class="paragraph" style="text-align:left;">As the data from Morgan Stanley below show, the cost per megawatt of Chinese chips is overwhelmingly lower (blue) than that of American chips (green), even those tailor-made for the Chinese market, like the H20.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d919f963-493a-4075-a44d-a5817e9e6e28/image.png?t=1781433455"/><div class="image__source"><span class="image__source_text"><p>Source: Morgan Stanley</p></span></div></div><p class="paragraph" style="text-align:left;">I will say I do think MS massively overshoots cost/token for American chips here; there’s no way the cost per million tokens is $10. SemiAnalysis uses a more realistic calculation below, <b>resulting in a number below $1.</b></p><p class="paragraph" style="text-align:left;">However, this is a very idealistic scenario, too, <b>because they use non-agentic sequences</b> (very short sequences) <b>and multi-token prediction </b>(every model prediction produces not one but three or four tokens).</p><p class="paragraph" style="text-align:left;">Sequences today are much longer, increasing cache size and thereby raising the average hardware intensity per workload (i.e., the average workload needs more GPUs), which dramatically worsens cost/token.</p><p class="paragraph" style="text-align:left;">On the other hand, MTP reduces cost/token by 3 or 4, as you get three or four times more tokens for roughly the same effort. However, MTP hurts performance, <b>so I’m not sure how representative it is of a real workload</b> (definitely used, though).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ae1d17ea-aece-49ad-9c43-42cc3bcd79ac/image.png?t=1781433653"/></div><p class="paragraph" style="text-align:left;">Thus, the real cost/token is probably somewhere in between. I’m going way off topic here, so to summarize, <b>the US has a capacity advantage</b> (it can deploy more), <b>but it’s not only not becoming cheaper but also more expensive to deploy each new GW</b>.</p><p class="paragraph" style="text-align:left;"><i>How sustainable is the buildout if marginal investment costs are rising instead of falling?</i></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But besides the increasing evidence coming from China, a new model<b>, or dare I say “model panels,”</b> dropped yesterday to prove that we have basically got the concept of “frontier” completely wrong.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-anthropic-s-mess-doesn-t-matter-the-era-of-sophistication">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=4b316d36-8b5c-4c9e-ae8e-aecdd9472135&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Sabotages, Diffusion, and a $2 Trillion Company</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fb40729c-e4bf-4c1f-ac29-a4dc3ab17980/image.png" length="252660" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/sabotages-diffusion-and-a-2-trillion-company</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/sabotages-diffusion-and-a-2-trillion-company</guid>
  <pubDate>Fri, 12 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-12T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back!</b> This week, we have lots to talk about. From AI satellites and an AI company sabotaging customers to what I believe is the future of edge AI: diffusion models.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>FRONTIER</b></span><br>Probably the Biggest Self-Own in AI Ever</h2><p class="paragraph" style="text-align:left;"><b>Anthropic</b> launched <b>Claude Fable 5</b> and <b>Claude Mythos 5</b> on June 9, introducing a new Mythos-class model family for advanced coding, knowledge work, vision, scientific research, and long-context tasks.</p><p class="paragraph" style="text-align:left;">The innovative aspect is that Fable 5 includes additional safeguards relative to Mythos in areas such as cybersecurity, biology, chemistry, and model distillation, due to risks perceived by Anthropic.</p><p class="paragraph" style="text-align:left;">These safeguards are applied using classifiers that detect certain requests in those categories. When detected, the response will instead be handled by Claude Opus 4.8, and users will be informed when this occurs.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Only the “selected few” with access to Mythos, unlike the peasants you and I are, will be able to use frontier-level capabilities to make biology or cybersecurity work.</p><p class="paragraph" style="text-align:left;">Much worse, for the particular topic of Frontier LLM development, Anthropic has done something unprecedented: they purposely downgrade, or dumbify the model so that it provides worse answers… <b>without telling you</b>.</p><p class="paragraph" style="text-align:left;">Have you ever met a company willing to sabotage your responses, without telling you, just because they disliked what you asked? </p><p class="paragraph" style="text-align:left;">This announcement caused massive controversy, and today they actually dialed it back: they won’t be dumbifying Fable anymore and will instead explicitly downgrade you to Opus 4.8, like in the other topics… or so they say (I don’t trust their word at this point).</p><p class="paragraph" style="text-align:left;">Both Fable 5 and Mythos 5 are priced at <b>$10 per million input tokens</b> and <b>$50 per million output tokens</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I don’t want to say I told… but I did. If you’re a regular of this newsletter, y<b>ou know I’ve not been shy about calling Anthropic out for perhaps years</b>.</p><p class="paragraph" style="text-align:left;">I never trusted them, and the world finally understands why.</p><p class="paragraph" style="text-align:left;">It all started once they began with their fearmongering about AI destroying humans, equating them to nuclear bombs, and all that condescending “I’m creating a God, only we should do it” stuff without giving a shred of proof.</p><p class="paragraph" style="text-align:left;">I will say it stuck with most people for a while. But I didn’t buy it because I understand the technology and its limitations, and I knew from the start it was all a marketing-slash-regulatory-capture strategy.</p><p class="paragraph" style="text-align:left;">To me, the most unforgivable thing is not that they decide I’m not worthy of using the frontier models, but the fact that a company trained on all of our data, while training models based on architectures developed by Google, not by them, <b>only to pull up the ladder behind them once they were ahead</b>, is just upsetting.</p><p class="paragraph" style="text-align:left;">Just picture a world where these guys get to decide who can use, and when, the most powerful technology humans have probably ever built. I don’t want to live in that world in the same way I wouldn’t want to live in Turkmenistan’s dictatorship.</p><p class="paragraph" style="text-align:left;">But I’m a positive guy, so I have positive takeaways:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Anthropic has damaged its reputation so badly that it will probably <b>have to pivot away from all this fearmongering crap for good</b>.</p></li><li><p class="paragraph" style="text-align:left;">This is going to invigorate/alert the US ecosystem that we need to improve open models, and fast. Companies like NVIDIA, Arcee, AI2, or even Google always understood the importance of open-source, and this should be a wake-up call for all of them</p></li><li><p class="paragraph" style="text-align:left;">The US Government comes out as a victor in this because its skepticism about Anthropic was clearly warranted, <a class="link" href="https://x.com/yanda/status/2064712629495705655?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">and many people are echoing this</a>: “Damn, the DoD was right.” <b>Now it’s time to pressure US Labs, especially OpenAI/Anthropic, to publish open research again</b>. Otherwise, you’ll still eventually lose to China because Anthorpic/OpenAI alone will not win against the entire Chinese ecosystem; thinking that’s possible is pure madness.</p></li><li><p class="paragraph" style="text-align:left;">I’m not a bigot, so if you ask me, <b>I’ll still recommend their models to clients because they are great (if not the best)</b>, no matter how much I dislike their leadership, because boy, are they capable of creating good models.</p></li></ol><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>SPACE</b></span><br>SpaceX Presents the AI1 Satellite</h2><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fb40729c-e4bf-4c1f-ac29-a4dc3ab17980/image.png?t=1781171737"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/SawyerMerritt/status/2064108916611420273?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">SpaceX, the new 2-trillion company (see below) has presented</a> what would be its first AI satellite running computing from space. It’s a massive satellite, more than 70 yards long, with computing power of up to 150kW at peak and capable of sustaining 120kW on average.</p><p class="paragraph" style="text-align:left;">For reference, that is roughly the electrical draw of a dense Blackwell Ultra-class AI rack on Earth.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">SpaceX is quite literally the definition of a moonshot. I haven’t bought into the IPO, but if they get this right, SpaceX is the ultimate AI play: they design the chips, they will manufacture them too, they have the rockets to push them into space, and they train the AIs that go into them while also being capable of supplying compute to other players like Anthropic, Cursor (which they acquired), or Google.</p><p class="paragraph" style="text-align:left;">It’s a great company, but in my view outrageously overvalued (more on this below).</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>COMPUTE</b></span><br>Google Enters Deal with SpaceX</h2><p class="paragraph" style="text-align:left;">Continuing with the star of the week, <a class="link" href="https://www.reuters.com/business/media-telecom/spacex-signs-cloud-deal-with-google-2026-06-05/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">as published by Reuters</a>, SpaceX has signed a <b>multi-year cloud services agreement with Google</b>, under which Google will pay SpaceX <b>$920 million per month</b> from <b>October 2026 through June 2029</b> for access to AI computing capacity.</p><p class="paragraph" style="text-align:left;">The deal gives Google access to about <b>110,000 Nvidia GPUs</b>, plus CPUs, memory, and related infrastructure.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">With this deal, SpaceX’s RPO compute backlog (the “agreed” revenue commitments it has signed with customers) amounts to $70 billion over roughly the next three years between the Anthropic and Google deals.</p><p class="paragraph" style="text-align:left;">For reference, SpaceX’s core business until recently was Starlink (Internet satellites), which has an ARR of $13.6 billion based on Q1 revenues. <b>This means the AI business, in just two deals, is already twice as large.</b></p><p class="paragraph" style="text-align:left;">The caveat, naturally, is margins.</p><p class="paragraph" style="text-align:left;">While the Starlink business has a 38% operating margin (meaning it still makes money after subtracting production and operating costs), <b>the AI segment</b>, despite its amazing revenue growth, <b>is enormously unprofitable</b> once we account for the huge depreciation costs associated with AI hardware and research & development costs (researcher salaries and training costs).</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>Recursive shows us the way to RSI</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b0bfc524-6e42-482a-b897-00f20c71062d/image.png?t=1781272340"/></div><p class="paragraph" style="text-align:left;"><b>Recursive AI</b>, a neo AI startup, <a class="link" href="https://www.recursive.com/articles/first-steps-toward-automated-ai-research?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">has released early results</a> from an automated AI research system that can propose ideas, implement them, run experiments, validate results, and use prior results to select subsequent experiments.</p><p class="paragraph" style="text-align:left;">Recursive says the system reached state-of-the-art results on three benchmarks: fixed-budget small language model training, small-model training speed, and GPU kernel optimization.</p><ul><li><p class="paragraph" style="text-align:left;">On <b>NanoChat Autoresearch</b>, it improved validation loss from <b>0.9372 BPB</b> to <b>0.9109 BPB</b>.</p></li><li><p class="paragraph" style="text-align:left;">On <b>NanoGPT Speedrun</b>, it reduced training time from <b>79.7 seconds</b> to <b>77.5 seconds</b> to reach the target validation loss.</p></li><li><p class="paragraph" style="text-align:left;">On <b>SOL-ExecBench</b>, it raised the mean score from <b>0.699</b> to <b>0.754</b> across 235 GPU kernel tasks.</p></li></ul><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">The system tested changes in model architecture, optimizer behavior, embeddings, attention precision, compiler settings, and fused GPU kernels. Recursive says it also screened results for reward hacks and variance before treating them as improvements. They are also open-sourcing artifacts from the runs so others can inspect and build on the outputs.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is the way. <b>Research that is created in the open to benefit all, not whatever Anthropic’s leadership with a God complex intends to do</b>.</p><p class="paragraph" style="text-align:left;">These RSI first steps are just that, minor wins that could one day lead us to a future where AIs indeed self-improve.</p><p class="paragraph" style="text-align:left;">The fact that AIs can modify themselves can be viewed with fear, sure, <b>but that’s precisely why this needs to be done in the open</b>, not behind walls of regulation and money, as some believe it must be.</p><p class="paragraph" style="text-align:left;">Money and power corrupt, so it’s vital that a technology that could one day become unstoppable be built in the open.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>STOCK MARKET</b></span><br>SpaceX is a two-trillion dollar company</h2><p class="paragraph" style="text-align:left;">SpaceX went public today, pricing its IPO at $135 per share and raising roughly $75 billion at an initial valuation of about $1.77 trillion.</p><p class="paragraph" style="text-align:left;">Shares were indicated to open higher, around $171, implying a valuation above $2 trillion. The stock is trading at $2.1 trillion, the same market value as TSMC, at the time of writing.</p><p class="paragraph" style="text-align:left;">For reference, <b>this makes SpaceX the seventh-largest company in the world by market cap</b>, just below TSMC and ahead of companies like Meta or Broadcom. Its value is greater than that of JPMorgan and Walmart combined.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Absolute madness. SpaceX trading higher than companies with:</p><ul><li><p class="paragraph" style="text-align:left;">38 times more revenues (Walmart)</p></li><li><p class="paragraph" style="text-align:left;">11 times more revenues (Meta)</p></li><li><p class="paragraph" style="text-align:left;">10 times more revenues (JP Morgan)</p></li><li><p class="paragraph" style="text-align:left;">7 times more revenues (TSMC)</p></li></ul><p class="paragraph" style="text-align:left;">It will be worth more than a potential corporation with almost 50 times the revenue (Walmart + JPMorgan).</p><p class="paragraph" style="text-align:left;">And look, I see the appeal in the company’s business. They have quite literally everything: a verticalized AI play, a monopoly on rockets and satellites… except the most important thing: <b>profits</b>.</p><p class="paragraph" style="text-align:left;"><b>The deals with Anthropic and Google really help</b> (the latter explained below) really help paint an improved picture, <b>adding $70+ billion to future revenues, but will do so across three years from now</b>, and will depend, especially in the case of Anthorpic, of cash flows that Anthropic’s business will not be able to guarantee (meaning, that money will come from investors or debt).</p><p class="paragraph" style="text-align:left;">I’ve maintained for a while now that this year’s IPOs, especially the upcoming Anthropic and OpenAI IPOs, will determine what we make of AI’s future.</p><p class="paragraph" style="text-align:left;">But seeing the—quite frankly delusional—fervency around SpaceX’s IPO makes me think that AI is going to be just fine because retail investors are willing to throw their money down the drain.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MONEY</b></span><br>OpenAI Considering Subscription Price Cuts</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.wsj.com/tech/ai/openai-considers-drastic-price-cuts-anticipating-war-for-users-with-anthropic-9b8c178e?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">According to the Wall Street Journal</a>, <b>OpenAI is considering cutting token prices to put pressure on Anthropic</b>, which would open a new chapter in the price wars that have been on a truce for more than a year now (token prices have barely moved in recent times).</p><p class="paragraph" style="text-align:left;">This is not something OpenAI can afford to do, but something to pressure Anthropic, which has a much smaller customer base, to reduce prices <b>at a time when OpenAI’s Codex is stealing a lot of users from Anthropic due to the higher rate limits</b> (because OpenAI has long had a much greater amount of compute than Anthropic, somethng the latter has sort of fixed in the last two months, but paying a high price for it).</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">While I understand OpenAI’s reasons for this,<i> these companies are really never going to make money, do they?</i></p><p class="paragraph" style="text-align:left;">The reason is quite simple: <b>it’s a commoditized technology</b> (otherwise, price wars wouldn’t be a thing) <b>running on non-commoditized hardware</b> (the hardware AI needs has some of the highest gross margins on the planet).</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Particularly daunting is the topic of subscriptions. SemiAnalysis ran a series of tests yesterday and found that subscriptions are, quite literally, cash-burning machines for these Labs, <b>with spend reaching $14k for a $200 subscription</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/be36f090-dec1-457c-a851-751847deb188/image.png?t=1781170077"/></div><p class="paragraph" style="text-align:left;">This shouldn’t be new to you, <b>considering I’ve long talked about the unprofitability of subscriptions</b>.</p><p class="paragraph" style="text-align:left;"><i>But why? </i>At the end of the day, the rationale is pretty straightforward:</p><p class="paragraph" style="text-align:left;"><b>AI is a technology with high marginal acquisition costs</b>, meaning each individual user can generate outsized costs due to the high hardware intensity of the average workload (<a class="link" href="https://thewhitebox.beehiiv.com/p/clarifying-myths-in-the-us-vs-china-war?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">as we discussed in our last issue</a>, the real problem is capital costs, not really inference costs).</p><p class="paragraph" style="text-align:left;">This means OpenAI or Anthropic can predict the profitability of a given subscription; you could make 95% gross margins, or -500%, depending on the user. <b>This breaks the historical “software contract,” in which, under CPU-based regimes, marginal costs were negligible.</b></p><p class="paragraph" style="text-align:left;">Pre-AI, for SAP, selling a $30/month seat would require, at most, $3-4 in production costs, give or take, <b>guaranteeing +85% gross margins no matter how “eager” that user was</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">As long as that remains true (hint: it will), <b>subscriptions don’t make any sense, and usage-based pricing is the only path to profitability</b>.</p><p class="paragraph" style="text-align:left;">This is why I enter every single board room I go into with this tattooed on my head: <b>assume subscriptions will disappear</b>; build an IT government model that assumes usage-based pricing on a cost-plus basis, where every token counts.</p><p class="paragraph" style="text-align:left;">OpenAI might make the subsidized era last a little longer, but it doesn’t change the future, which is a pay-as-you-go model.</p><p class="paragraph" style="text-align:left;">As for consumers, where usage-based models rarely work, it’s a non-issue because most consumer tasks (searching for stuff, asking questions, editing videos, etc.) <b>will mostly run</b><b> on edge hardware in one/two years’ time </b>(more on this later in this newsletter)<b>.</b></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>SOFTWARE</b></span><br>Wanna survive the Saaspocalypse? Do this</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://ramp.com/blog/introducing-ramp-applied-ai-solutions?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">Ramp has launched an </a><a class="link" href="https://ramp.com/blog/introducing-ramp-applied-ai-solutions?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">AI-applied solutions</a> service, in which Ramp engineers embed with enterprise finance teams to build customized AI agents and workflows on Ramp’s platform.</p><p class="paragraph" style="text-align:left;">Ramp says the offering is aimed at companies that are increasing AI spending but struggling to show measurable returns.</p><p class="paragraph" style="text-align:left;">It cites internal customer data showing that AI token spend across its 70,000+ customers has risen <b>13x since January 2025, </b><b>while only 21% report measurable results.</b></p><p class="paragraph" style="text-align:left;">Ramp says the service is <b>model-agnostic</b>, routing workflows to different AI models based on performance, cost, and task fit, rather than locking customers into a single provider.</p><p class="paragraph" style="text-align:left;">The company says deployments are designed to ship into production “in weeks,” with customers able to take ownership after handoff or keep Ramp involved as the system expands.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Even though I’m not a Ramp customer (never tried the tool),<b> this is the exact playbook a software company will have to follow to survive</b>.</p><p class="paragraph" style="text-align:left;">I’ve maintained for a long time two things about software’s future:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>It will be agent-first</b>, meaning it will be lightweight and<b> low margin</b> (and low prices) while being primarily used by agents, not humans</p></li><li><p class="paragraph" style="text-align:left;"><b>It will be customizable and adaptable</b>, feeling almost bespoke to the user</p></li></ol><p class="paragraph" style="text-align:left;">Ramp ticks both boxes with their decision. On the one hand, it’s clearly becoming an AI-native company, actively pushing research and fine-tuned models.</p><p class="paragraph" style="text-align:left;">But instead of just waiting for customers to build their own bespoke solutions to replace the Ramp license, <b>they offer that customization on top of the Ramp platform, providing the “be spokedness” customers want while still creating client lock-in on their product.</b></p><p class="paragraph" style="text-align:left;">If you’re an investor looking to see what software companies will exist in five years,<b> this is the type of sign you should be looking for</b>: companies like Palantir or Ramp that, instead of offering a generalist software, proactively move into the customer and become their bespoke solution.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Google Releases DiffusionGemma</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3321f33e-4940-47ee-9dd8-9c39a44a62db/Screen_Recording_2026-06-11_at_13.09.14.gif?t=1781176228"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">Google has released a new Gemma 4 variant</a> that can run on consumer hardware, delivering strong performance for its size.</p><p class="paragraph" style="text-align:left;">But the fascinating thing is the architecture itself: it’s a diffusion model, hence the name <b>DiffusionGemma</b>.</p><p class="paragraph" style="text-align:left;">In other words, unlike standard Large Language Models (LLMs), which generate one token (e.g., a word) every round, generating an ‘autoregressive’ sequence of tokens, <b>Diffusion models take a more immediate approach.</b></p><p class="paragraph" style="text-align:left;">They depart from a noisy overall picture, and iteratively “denoise” the slate to “unearth” the response.</p><p class="paragraph" style="text-align:left;">I’ve always thought of this intuitively as sculpting: you take a huge block of marble and “unearth” the “hidden” statue by chiseling away the excess. As the great artist Michelangelo once said:</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/fb372d13-22ae-4be7-b106-8eb5aa7259bf/image.png?t=1781176570"/></div><p class="paragraph" style="text-align:left;">This means diffusion models progressively transform what’s essentially noise into an actual result by performing several ‘denoising’ updates.</p><p class="paragraph" style="text-align:left;">However, <b>diffusion models introduce an unequivocal trade-off</b>: they are much faster because they generate results globally rather than sequentially, but in most cases this implies a decrease in performance.</p><p class="paragraph" style="text-align:left;">Google itself mentions this: <i>“For applications demanding maximum quality, we recommend deploying standard Gemma4”.</i></p><p class="paragraph" style="text-align:left;">But DiffusionGemma and other diffusion models do introduce one vital aspect that makes me particularly bullish about them: <b>they massively reduce the memory bottleneck</b>.</p><p class="paragraph" style="text-align:left;">Too long to explain here (in case you want the longer explanation, <a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/why-the-worlds-ai-will-run-on-diffusion-models-df2a18c581d8?sk=7be84c060313d017d2c3ad786555e0b5&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">read here</a>), but I’ll try my best to simplify. Computers work by moving data into a processor, which performs a series of operations, and most of the results and the input data are sent back to memory, creating a back and forth between the memory and the compute.</p><p class="paragraph" style="text-align:left;">This means that both are needed, and the slowest of the two is the bottleneck. <b>In AI, especially inference, memory is the bottleneck</b>, which is why AI is famously defined as “memory-bound.” In practice, this means that, on average, processors are somewhat idle, or “waiting” for data to arrive.</p><p class="paragraph" style="text-align:left;">For companies that make money by producing tokens, <b>that idle time translates literally into lost revenue.</b></p><p class="paragraph" style="text-align:left;">The reason inference is so bound by memory is easy to see. Think about how ChatGPT works, generating one word at a time. In practice, this means you push the model into the processor, the processor decides the next word, and the process repeats.</p><p class="paragraph" style="text-align:left;">The issue is that models are so large they break down before being pushed, so to make way for the next part of the model, we need to extract the previous part, increasing the amount of data moved in and out.</p><p class="paragraph" style="text-align:left;">Naturally, this means that if data movement is the bottleneck, <b>the way to squeeze the most performance is to make the most of every byte of data we push into the chip</b>.</p><p class="paragraph" style="text-align:left;">That is, if a GPU can allegedly perform 100 operations per byte of data the processor sees, the goal is to ensure we stay as close to that value as possible. This way, we guarantee that processors are running at a good pace.</p><p class="paragraph" style="text-align:left;"><i>And what does all this have to do with diffusion?</i> Simple: while we have to do this entire data dance to generate a single token in an autoregressive model like ChatGPT,<b> a diffusion model updates 256 tokens at each step. </b></p><p class="paragraph" style="text-align:left;">Careful, I’m not saying that every prediction pass churns 256 tokens, but a step in the denoising process. However, that results in extremely fast generation because once the denoising updates finish, you automatically get 256 tokens.</p><p class="paragraph" style="text-align:left;">For example, for two identical models, one autoregressive and one diffusion, if the denoising steps are 30 (as in the Sudoku example above), <b>while an autoregressive LLM generates 30 tokens, DiffusionGemma generates 256</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The reason I’m so bullish on diffusion models is that they will play a crucial role in edge hardware. Our smartphones and laptops don’t have nearly as much capacity for fast token generation as cloud servers do, and thus require algorithmic improvements to churn tokens faster.</p><p class="paragraph" style="text-align:left;">It’s not the intelligence level that makes edge models hard to adopt;<b> it’s how slow they are</b>. Diffusion models can massively increase token generation speed, making adoption much easier.</p><p class="paragraph" style="text-align:left;">As they reach “good enough” thresholds, <b>I believe they are going to be representative of a huge portion of the world’s generated tokens</b>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>EMERGENT CAPABILITIES</b></span><br>Using AI to See Who Stresses You the Most</h2><p class="paragraph" style="text-align:left;">One of the most powerful yet often-overlooked abilities of frontier AIs is <b>reverse engineering</b>: they can examine a product or service and determine how it works. For example, they can look at screenshots from an app and generate the source code without having seen it before.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/the2ndfloorguy/status/2064704204166635930?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">An X user</a> has connected their Whoop app, which controls a wristband that tracks heart rate and other metrics to assess stress or sleep quality, to Calendar, so it can track which meetings (and, thus, people) give them the most stress.</p><p class="paragraph" style="text-align:left;">To do this, they used the new Claude Fable to reverse-engineer the Whoop so that it could pull per-minute heart rate data and, in this way, associate what people give this man with more stress.</p><p class="paragraph" style="text-align:left;">What willpower and boredom do to someone, <i>right?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">More than the funny use of the technology,<b> my takeaway is how easy it’s going to be to build apps on the fly.</b></p><p class="paragraph" style="text-align:left;">Not necessarily things you can sell (with more apps available, getting attention will be harder than ever), but creating software as a way to solve personal problems, not necessarily as a product or service for others.</p><p class="paragraph" style="text-align:left;"><i>New project? </i>An app to track progress with the particularities of this one. <i>New son&#39;s hobby?</i> Build an app that helps your kid.</p><p class="paragraph" style="text-align:left;">The opportunities will be endless, and so will the cybersecurity risks, but we’ll leave that for another day.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ENTERPRISE</b></span><br>Harvey Shows The Way</h2><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/384ea4dd-bd52-462a-a6ca-f2afa7ac5cd8/image.png?t=1781173568"/></div><p class="paragraph" style="text-align:left;">Harvey, the AI legal startup, <a class="link" href="https://x.com/harvey/status/2064757424540749879?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">announced</a> a partnership with Trajectory Labs to post-train NVIDIA’s Nemotron 3 Ultra for legal-agent work, achieving frontier-level results in just 24 hours of fine-tuning on NVIDIA&#39;s open models using reinforcement learning (RL) and implying a 50x cost reduction.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">According to Harvey, the base Nemotron 3 Ultra model scored 0% on LAB’s all-pass metric before post-training. After less than 24 hours of post-training, the model reached 5.8% all-pass. That sounds mediocre<b>, but it’s almost as good as Claude Opus 4.6, a frontier-ish model.</b></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I’ve been saying for years that fine-tuning will play a vital role for enterprises, taking an open, non-frontier model and turning it into a frontier model (for that particular task), training it on task data.</p><p class="paragraph" style="text-align:left;"><b>I’ve long maintained that once enterprises start training models on company data, OpenAI and Anthropic will have a hard time selling into businesses</b>, because companies will have the option to use free or cheap models that can be run securely within their orgs and achieve frontier-level performance without the added premium costs.</p><p class="paragraph" style="text-align:left;">As I always say, enterprises don’t need generalist savant models; <b>they need models that do a given task well</b>, even if training them to do that well makes them worse in other areas. <i>Who cares?</i> Just pick another model for the other tasks.</p><p class="paragraph" style="text-align:left;">The frontier market will still exist, as some critical use cases require frontier-level performance. Think coding, drug discovery, science, cybersecurity (to some extent, not really true), or maths; in those areas, you want to maximize “intelligence.” But for everything else, what you want to maximize is “intelligence-per-dollar”.</p><p class="paragraph" style="text-align:left;">And in that regard, Anthropic/OpenAI models are nowhere to be seen:</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e726a744-237d-45e7-8d33-7e9950f9a2de/image.png?t=1781174041"/></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ROBOTICS</b></span><br>Humanoids take a human look</h2><p class="paragraph" style="text-align:left;">It was inevitable, but I can’t say I’m not spooked either way. Chinese company UbTech <a class="link" href="https://x.com/tphuang/status/2064140953971982406?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=sabotages-diffusion-and-a-2-trillion-company" target="_blank" rel="noopener noreferrer nofollow">has presented its humanoid companion</a> for pre-sale, getting 3,000 orders in a week. <b>The robots are human-like and are meant to serve as companions</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;"><i>How far has humanity fallen to requiring non-human humanoids to not feel lonely (or to get laid, which is even more concerning)?</i></p><p class="paragraph" style="text-align:left;">I can understand having a humanoid handling packages in a factory; I don’t think that’s anyone’s dream job. <b>But actively trying to assume the roles of other humans in basic human-to-human relationships is Kafkaesque</b>.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">The undeniable highlight of the week is the fact that a space company rebranded as an AI “datacenters in space” company, which is considerably unprofitable, <b>has managed to convince people that it’s valued at $2 trillion</b>, which makes one wonder, has investing changed forever?</p><p class="paragraph" style="text-align:left;"><i>Are we in 1997’s Alan Greenspan’s “irrational exuberance“ mode 30 years later?</i></p><p class="paragraph" style="text-align:left;">In the meantime, the release of DiffusionGemma by Google and Anthropic’s uncalled-for antics make it very clear that, now more than ever,<b> AI has to be open, and AI has to be small enough to run locally</b>. We can’t let this technology be controlled by misaligned, power-hoarding entities.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=884f5219-4e4d-4166-b4fc-b35ae616697d&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Clarifying Myths in the US vs China War</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cd41ad88-5201-47b2-971f-3807739f3bc1/image.png" length="97075" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/clarifying-myths-in-the-us-vs-china-war</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/clarifying-myths-in-the-us-vs-china-war</guid>
  <pubDate>Tue, 09 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-09T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Clarifying Myths in the US vs China War</h2><p class="paragraph" style="text-align:left;"><b>Welcome back!</b> This is an article I’ve wanted to write for a while now, exclusively focused on clearing up some of the most prevalent myths in this space, especially regarding China.</p><p class="paragraph" style="text-align:left;">With the world suddenly realizing that AI is expensive, the “battle” between the US and China for AI supremacy is hotter than ever. If the tide turns in favor of Chinese open models, the trillions of dollars at stake could be at clear risk.</p><p class="paragraph" style="text-align:left;">But as always, the picture is much more nuanced, and several of your beliefs right now that you have been told are simply blatantly false.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">Uncovering Several Myths (and Truths)</h2><p class="paragraph" style="text-align:left;">Ironically, <b>the first myth that needs to be debunked for good is the idea that China is “</b><b><i>catching u</i></b><b>p.”</b> The truth is that, despite all the hard evidence you may be seeing, it’s simply not true.</p><p class="paragraph" style="text-align:left;">But here’s the thing. As you’re about to witness, it might not matter at all.</p><h3 class="heading" style="text-align:left;">Don’t Trust the Benchmarks</h3><p class="paragraph" style="text-align:left;">The first myth is this graph below,<b> claiming that Chinese models </b>(they refer to open-weight models, models that are free to download, but most of the best ones are Chinese)<b> are only 4 months behind</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b4c43c64-7a08-4363-8ae3-57a01aa5081f/Open_models_lag_state-of-the-art_closed_models_by_4_months.jpg?t=1780565404"/></div><p class="paragraph" style="text-align:left;">The data seems hard to dispute; Kimi K2.6 seems to be at the level the US was at with GPT-5.3 Codex in December. However, <b>it’s false</b>.</p><p class="paragraph" style="text-align:left;">And the issue is precisely drawing that conclusion based on model benchmarks and not on the product.</p><p class="paragraph" style="text-align:left;"><b>And let me tell you that nobody more than I would want to see open-source catching up to proprietary solutions</b>. However, I’m not here to expose you to my desires, but to the truth.</p><p class="paragraph" style="text-align:left;">The reason is quite simple. In the absence of fundamental breakthroughs in algorithms that do not exist today, AI progress is driven by two scaling laws:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Larger training budgets</b>, which require bigger models</p></li><li><p class="paragraph" style="text-align:left;"><b>Larger thinking budgets</b>, which require larger reasoning sequences (i.e., thinking for longer improves performance)</p></li></ol><p class="paragraph" style="text-align:left;">And the US is enormously ahead in both.</p><p class="paragraph" style="text-align:left;">On the former, <b>US Labs have larger models</b> (in the order of ten times the size) and perhaps even <b>two orders of magnitude larger training budgets</b>.</p><p class="paragraph" style="text-align:left;">The largest known AI cluster in China is a <a class="link" href="https://www.scmp.com/tech/big-tech/article/3348502/shenzhen-activates-chinas-first-10000-card-ai-cluster-domestic-chips?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">10,000-Ascend 910C cluster in Shenzhen with up to 11,000 PetaFLOPs of compute</a>, or 11 ExaFLOPs. That number doesn’t say much by itself, but if we assume it is FP8, which most likely it is, <b>it’s a fourth of the compute of a single Google Ironwood 9,216-TPU pod</b> <a class="link" href="https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">at 42 ExaFLOPs</a>.</p><p class="paragraph" style="text-align:left;">And that isn’t even Google’s most powerful server, and they’ve scaled that number to a whopping 121 PetaFLOPs with the upcoming TPUv8 pods for training using FP4 precision.</p><p class="paragraph" style="text-align:left;">In apples to apples, that is 121 vs 22 ExaFLOPs, <b>6 times more, for approximately the same number of chips, 9,600.</b> That means a cluster with roughly the same number of chips has 11 times the compute potential.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/2478b224-b99f-4e13-b561-40156f823f76/image.png?t=1780908586"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/eighth-generation-tpu-agentic-era/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">I mean, a single NVIDIA Vera Rubin NVL72 server, recently deployed for the first time, has almost 60% higher compute capacity. One single server.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/99db6fd2-c61a-4049-ad46-1b3ae547e92e/image.png?t=1780566179"/></div><p class="paragraph" style="text-align:left;">In the meantime, two US models have already crossed the 10<sup>27</sup> FLOP training budget barrier, both Gemini 3.1 Pro and Claude Mythos, which, by the way, <a class="link" href="https://x.com/claudeai/status/2064394146916229443?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">just got released minutes ago</a>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">For reference,<b> that is about 100 times as much compute as was used to train GPT-4</b> (i.e., for that budget, you could have trained 100 GPT-4s).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/92755d9f-6a9f-4a4c-ba4e-24bf660a0df7/image.png?t=1780908722"/></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This is a massive disadvantage for China; it’s just is. On the second scaling law, <b>you need a lot of inference compute to be able to serve long sequences that increase performance</b>, compute that again China doesn’t have (relative to the US).</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">On the topic of training budgets, <i>how is it that China can “stay close” with budgets 100 times smaller?</i></p><p class="paragraph" style="text-align:left;">The answer is distillation, <b>which massively reduces your training requirements</b>. That is, by training your models on data from larger US models, you artificially close the gap with them very effectively and with much less compute, because you’re basically teaching your model to imitate the larger model.</p><p class="paragraph" style="text-align:left;">But distillation and its savings come at a cost that is not appreciable in the benchmarks: <b>generalization</b>.</p><p class="paragraph" style="text-align:left;">The fact that you trained a model on 10-100 times less compute means <b>your model is extremely unlikely to be as good as the former outside the prioritized domains</b>, something that benchmarks hide really well.</p><p class="paragraph" style="text-align:left;">You can purposefully train models to do well on benchmarks, but if you test those models outside that domain, you’re going to see the cracks pretty fast. But that doesn’t show in the benchmarks, <b>which means you appear to be better than you really are</b>.</p><p class="paragraph" style="text-align:left;">The other problem with benchmarks is that they hide reality from us; you aren’t seeing the actual product, just the models.</p><p class="paragraph" style="text-align:left;"><i>But what do I mean by that?</i></p><p class="paragraph" style="text-align:left;">When you’re testing for a benchmark, you’re doing so under perfect constraints. Probably batch 1 (meaning you’re serving the model in ideal conditions), with little regard to latency, and with huge compute budgets; you’re allocating maybe thousands of dollars to every task and tens of thousands to the overall benchmark in order to do the best you can.</p><p class="paragraph" style="text-align:left;"><b>But that’s not the reality your model will face in the real world</b>.</p><p class="paragraph" style="text-align:left;">The real world is not an ideal lab environment. In the real world, models need to fight for scarce compute, run in high batches (slower responses), and, crucially, <b>operate with limited thinking budgets so as not to bankrupt the serving company</b>.</p><p class="paragraph" style="text-align:left;">This means that the product, the actual deliverable that customers get, is NOT what the benchmark shows (and to be clear, this applies to US Labs too, just as much, it’s just that they have more compute).</p><p class="paragraph" style="text-align:left;">In real-time inference, Labs will quantize models, drop thinking budgets to meet demand, serve you distillations of the real thing, and others.</p><p class="paragraph" style="text-align:left;">All this shows that the real experience is not what benchmarks show. And as you can imply from all that I’m saying, for now, <b>this is a compute game all the way.</b></p><p class="paragraph" style="text-align:left;">Therefore, the point here is that it’s not surprising that Chinese Labs are becoming more frugal and thus innovating at the intelligence-per-dollar level; <b>they really have little option.</b></p><p class="paragraph" style="text-align:left;">But if they could compete at the Frontier with equal resources, they would behave exactly like US Labs are behaving (<a class="link" href="https://www.datacenterdynamics.com/en/news/alibaba-cloud-needs-10x-its-2022-compute-capacity-says-ceo-eddie-wu/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">Alibaba planning to 10x its compute</a>, <a class="link" href="https://economictimes.indiatimes.com/news/international/us/chinas-ai-reality-check-why-tech-leaders-say-beating-us-ai-giants-is-unlikely-anytime-soon/articleshow/126457670.cms?from=mdr&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">Zhipu co-founder acknowledging the gap could be widening</a>).</p><p class="paragraph" style="text-align:left;">But then, <i>how do we compare both?</i></p><p class="paragraph" style="text-align:left;">An actual fair comparison would be product benchmarks that measure model performance within their products and serve them to real users. Because when you do that, <b>you realize that the user experience is like night and day between American and Chinese products</b>.</p><p class="paragraph" style="text-align:left;">The only thing that could change this is a fundamental breakthrough that proves that models can be trained and served with less compute. The closest thing we have is Sapient Intelligence’s <b>HRM-Text</b> model.</p><p class="paragraph" style="text-align:left;">This is a billion-parameter model that I discussed<a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/training-ais-for-a-millionth-of-the-cost-93ca0357690a?sk=2dd1ba77293107dcd4ff12649f60a642&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow"> in detail here recently</a>. You can run it on your smartphone, and it proved to be quite decent and superior to GPT-3.5, despite requiring <b>~700x fewer training tokens</b> and an estimated <b>44,000x fewer training FLOPs</b> <b>(operations)</b>, as you can see in the graphs below.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/de1ebab1-50a2-447a-93f4-73c40bd56f01/image.png?t=1780567429"/></div><p class="paragraph" style="text-align:left;">Interestingly, <b>GPT-3.5 was the model ChatGPT used when it first launched in late 2022, and it was the state-of-the-art at the time</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Sadly, however, while promising, <b>this is not a particularly useful model</b>, and we need this to scale to more useful capability thresholds. If we saw a model that can run on an iPhone having the performance of the Chinese frontier today, well, that would be a complete revolution. However,<b> that’s simply not a reality today</b>.</p><p class="paragraph" style="text-align:left;"><i>But what about Chinese models served on US LLM providers? Are those a threat to US Labs?</i></p><p class="paragraph" style="text-align:left;">Well, that’s a completely different story.</p><h3 class="heading" style="text-align:left;">Usage and Prices</h3><p class="paragraph" style="text-align:left;">A few days ago, the CEO of <a class="link" href="https://www.lindy.ai/?pscd=try.lindy.ai&ps_partner_key=NTUzZmYxN2U5ODhm&ps_xid=8mgCvtyyy8F7G8&gsxid=8mgCvtyyy8F7G8&gspk=NTUzZmYxN2U5ODhm&gad_source=1&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">LindyAI</a>, a chatbot product that helps users increase productivity with AI models and has millions of users,<b> switched its underlying models from Anthropic’s Claude to DeepSeek.</b></p><p class="paragraph" style="text-align:left;">And not only are they saving millions of dollars, <b>but they claim an actual increase in performance.</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cd41ad88-5201-47b2-971f-3807739f3bc1/image.png?t=1780909737"/></div><p class="paragraph" style="text-align:left;">And <a class="link" href="https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">with companies like Uber saying “enough is enough” </a><a class="link" href="https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war" target="_blank" rel="noopener noreferrer nofollow">about uncontrolled AI spending</a>, many are waking up to something that was obvious in hindsight:<b> what really matters is intelligence per dollar</b>, and thinking that everyone would always prefer the frontier model was unequivocally shortsighted.</p><p class="paragraph" style="text-align:left;">Currently,<b> the prevailing sentiment is that Chinese models are cheaper and represent the Pareto frontier </b><b>in terms of intelligence per cost</b>, which suggests they are being widely adopted as we speak.</p><p class="paragraph" style="text-align:left;"><i>But are they? </i>Well, like the previous question, yes and no.</p><p class="paragraph" style="text-align:left;">The reality is that we already have plenty of evidence that open AI models are being adopted by US and European customers.</p><p class="paragraph" style="text-align:left;">Perhaps the best example of this is OpenRouter, a highly popular US LLM router that lets you access most models from a single place.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/219814c0-dbb2-410c-8ef1-a9da6fb60b7d/image.png?t=1780567694"/></div><p class="paragraph" style="text-align:left;">Another good example of increased open model usage is provided by Ramp, a company expense management company, showing that DeepSeek is becoming extremely popular amongst its enterprise clients.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a187d97a-3f88-4ce1-b6a7-caa423cce8eb/image.png?t=1780564267"/></div><p class="paragraph" style="text-align:left;">That looks like decent adoption, but the reality is that private tokens, those generated by Google/OpenAI/Anthropic/SpaceX, are several orders of magnitude larger. Nonetheless, Google’s latest figure is over 3.2 quadrillion tokens per month across its AI surfaces.</p><p class="paragraph" style="text-align:left;">That is <b>3,200,000,000,000,000+ tokens/month</b>.</p><p class="paragraph" style="text-align:left;">OpenRouter processes 34 trillion tokens per week, or 147T per month. That means Google alone processes 21 times as many tokens per month as OpenRouter, across all models it serves (including proprietary models), <b>signaling an enormous gap between private and open models.</b></p><p class="paragraph" style="text-align:left;">But Lindy’s transition from Anthropic to DeepSeek could be the start of the<b><i> “Huge Transition”,</i></b> where the decision-making algorithm for choosing a model is not <i>“let’s use OpenAI”</i> but <b><i>“let’s try the open models first, and if they don’t work, then we look at Anthropic”.</i></b></p><p class="paragraph" style="text-align:left;">All this could smell like a huge opportunity for China, considering the popular belief that Chinese models are unequivocally cheaper and offer better value for your buck. At this point, I believe this is widely accepted.</p><p class="paragraph" style="text-align:left;">However, <b>my issue with this entire conversation is that it is discussed from the wrong perspective.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h3 class="heading" style="text-align:left;">The enemy is inside</h3><p class="paragraph" style="text-align:left;">The issue is that the statement <i>“Chinese models offer 80% of performance per 20% the price”</i> is true but incomplete, because the last part is missing: <b><i>“but so are US open models”.</i></b></p><p class="paragraph" style="text-align:left;">Examples include the recently announced NVIDIA Nemotron 3 Ultra. Despite being half the size of Kimi K2.6 and three times smaller than DeepSeek v4 Pro (both Chinese models), i<b>t offers much better Pareto performance, is much faster, and is also cheaper.</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f4045f6d-9171-4b76-86b7-768f5f7c08f4/intempence_vs._Jurput_speco.jpg?t=1780564941"/></div><p class="paragraph" style="text-align:left;">And if we factor in closed models, like SpaceX’s Grok 4.3 model, it’s also clearly on the Pareto curve, offering a better price per token than Kimi K2.6.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c2820cd6-d52e-4c7c-9829-d972f899438a/image.png?t=1780910107"/></div><p class="paragraph" style="text-align:left;">Therefore, <b>there’s very little evidence that Chinese models are cheaper</b>. Conversely, I also want to use this moment to debunk a myth that China subsidizes inference costs. <b>It doesn’t</b>, and it’s easily provable in two ways:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">We have two public Chinese AI Labs, Minimax and Zhipu. Both have received CCP grants, but they are a small part of revenues (tiny in the case of the former, just 3%). Zhipu’s case is more favorable (grants account for 34% of revenue), but this is too little to make an argument that “the CCP is paying the inference bills”.</p></li></ol><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b7459fdb-0e98-4cc0-8efb-d16e0cb08892/image.png?t=1780568253"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><ol start="2"><li><p class="paragraph" style="text-align:left;">It takes a quick request to ChatGPT to compare Chinese endpoint prices with those of US LLM providers offering Chinese models to realize that the prices are basically identical, with only DeepSeek showing a higher likelihood of subsidizing.</p></li></ol><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/35873fe3-1784-46f9-8b13-d7bc201ed93b/image.png?t=1780568313"/></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Therefore, what investors in OpenAI and Anthropic should be scared of is not Chinese clouds offering competitive models, because inference has to be served from nearby due to latency constraints. <b>The threat is US inference providers</b>, companies like the Hyperscalers or neoclouds like Fireworks, <b>offering Chinese models on US soil while offering equally competitive pricing.</b></p><p class="paragraph" style="text-align:left;">So far, it doesn’t seem like China is ahead on anything. They are pretty good at training cheap models, but so is the US, and the benefits of these models are still being ripped by US companies.</p><p class="paragraph" style="text-align:left;">So, <i>where’s the issue then?</i></p><p class="paragraph" style="text-align:left;">Well, <b>there’s a place in this industry where China is ahead of the US, and it’s not even close.</b> An area where China is actually a threat.</p><p class="paragraph" style="text-align:left;">It has nothing to do with AI. It has nothing to do with talent. And it has nothing to do with a “secret sauce” nobody knows about.</p></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=clarifying-myths-in-the-us-vs-china-war">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=6a023fde-7d00-4887-a075-66fc8fb1db24&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>China and Uber, the US&#39;s Cookie Monsters</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/508a405e-7f75-4d58-9d3f-81c5a188bd1a/image.png" length="764905" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/china-and-uber-the-us-s-cookie-monsters</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/china-and-uber-the-us-s-cookie-monsters</guid>
  <pubDate>Wed, 03 Jun 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-06-03T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back!</b> This week, we have a long list of interesting news to talk about. From <b>China’</b>s latest SOTA model to the <b>US’</b>s next great small model, we also tackle <b>NVIDIA</b> and <b>Microsoft</b> super events, <b>Uber</b> clamping down on AI costs, <b>Bernie Sander</b>s’ latest AI rant, and many more.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Minimax M3, New Chinese SOTA?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/508a405e-7f75-4d58-9d3f-81c5a188bd1a/image.png?t=1780390132"/></div><p class="paragraph" style="text-align:left;">Minimax has released a new model that, on paper, is competitive with the best the US has to offer. Across several benchmarks, it holds its own against Opus 4.7 or GPT-5.5, alongside Opus 4.8, the bleeding edge of the industry.</p><p class="paragraph" style="text-align:left;">Despite being considerably smaller than its rivals (at least ten times smaller), it competes and offers a massive one-million-context window, <b>which suggests this model punches well above its weight class</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/this-new-chinese-ai-will-make-you-think-c670faadb1fa?sk=3e32e9606b876be0175ae400b11c54b8&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">I wrote about it in way more detail here</a>, but the crucial thing to highlight is that, much like all other Chinese Large Language Models (LLMs), <b>they perform some sort of compression over context.</b></p><p class="paragraph" style="text-align:left;">This stems from the fact that, while you might not remember what you had for dinner three days ago, you can recall decades-old childhood memories with ease, because the human brain is opinionated about what’s “worth remembering”.</p><p class="paragraph" style="text-align:left;">However, <b>this is extremely complicated to do with AIs</b>, to the point that we largely avoid doing so if we have the compute means to avoid it.</p><p class="paragraph" style="text-align:left;"><b>This results in models that do not compress context</b>; if you send them an entire 50-page report, they’ll store in context every single word in it, which means that an LLM’s context grows proportional to its length (and in a quadratic fashion, actually; tripling sequence length nine-folds the computational requirements).</p><p class="paragraph" style="text-align:left;">Chinese Labs, much more compute-constrained, <b>are forced to innovate in this regard</b>, and they do so by forcing this context compression.</p><p class="paragraph" style="text-align:left;"><i>But how?</i></p><p class="paragraph" style="text-align:left;">Say you’re reading a 100,000-word book and you want to predict what will happen in the next few hundred words. To do so, you’ll probably bear in mind what has recently occurred, while also taking into account a summarized view of what happened in earlier chapters. <b>You don’t remember every single thing that happened, only those things your brain considers important</b>.</p><p class="paragraph" style="text-align:left;">But if we give an AI a 100,000-word sequence to predict the 100,001st word, the model will store all previous 100,000 words. Not a single one is ignored. Even if a word is ‘uhm’ or ‘ehhh’, they are also attended to.</p><p class="paragraph" style="text-align:left;">This is as if, for you to decide what to eat today, you considered not only what you ate recently, but also what you ate a year ago from today, while also taking into consideration yesterday’s debate with your husband about which cushion color works best on your sofa.</p><p class="paragraph" style="text-align:left;">This seems dumb, but it’s exactly what is going on; <b>all options are considered.</b></p><p class="paragraph" style="text-align:left;">To handle this, <b>Minimax proposes a “hardware-aware” compression</b>. The context gets considerably compressed, in the same way other Chinese models like DeepSeek V4 do, <b>and it’s also designed to benefit GPU architectures the most</b>, basically by ensuring that, for every retrieved data bundle, as much of it as possible is ‘useful.’</p><p class="paragraph" style="text-align:left;">This is because, given the nature of DRAM, the memory used by GPUs, it’s faster to retrieve two contiguous data points than two separate ones: the first two can be retrieved directly, whereas the latter requires two retrievals (retrievals are fast, but they still add up to delay if done suboptimally).</p><p class="paragraph" style="text-align:left;">This is a marvel of engineering. <i>But is it really SOTA?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The jury is out, <b>but the answer is almost definitely no</b>. There’s no such thing as a free lunch, and the only reason Chinese Labs are running these compression mechanisms is that they&#39;re being forced to.</p><p class="paragraph" style="text-align:left;">US labs don’t run these mechanisms because they have the compute and the capital to avoid doing so. The rationale is simple: <b>if you force the model to compress, something that might be needed is forgotten, which obviously affects performance</b>.</p><p class="paragraph" style="text-align:left;">Chinese models are smaller, too, making it even harder to believe, because I’m afraid not, there’s no such thing as a Chinese secret sauce for now.</p><p class="paragraph" style="text-align:left;">But Chinese models are reaching a level of performance-per-dollar where they become ‘best value’ alternatives that could seriously undermine OpenAI&#39;s and Anthropic&#39;s ability to penetrate enterprise budgets.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>BIG TECH</b></span><br>NVIDIA’s Computex Event</h2><p class="paragraph" style="text-align:left;">A couple of days ago, <a class="link" href="https://www.youtube.com/watch?v=gxgi6D-Cf9I&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">Jensen Huang did a keynote at Computex</a>, Taiwan’s annual expo for the semiconductor industry, one of the most important events of the year. And Jensen had several things to announce.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Vera Rubin is ramping up to full production of “agentic AI factories,”</b> with Rubin-based products expected from partners in the second half of 2026.</p></li><li><p class="paragraph" style="text-align:left;">For AI infrastructure, <b>NVIDIA introduced DSX</b>, a software platform intended as a blueprint for building and operating AI factories, including simulation, power management, and coordination with energy providers.</p></li><li><p class="paragraph" style="text-align:left;">On PCs, NVIDIA announced <b>RTX Spark</b>, described as a new superchip for Windows PCs built around personal AI agents. NVIDIA also announced <b>DGX Station for Windows</b>, a deskside AI system aimed at enterprise users running very large models locally. <a class="link" href="https://youtu.be/OXneUrL3Fe0?si=TZIb-ZmuY_hyFyPm&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">As reviewed by Linus Tech Tips</a>, these laptops will be very powerful but insanely pricey (the 128 GB version goes for $8,000).</p></li></ol><p class="paragraph" style="text-align:left;">While the obvious highlight was Vera Rubin&#39;s servers being deployed finally, to me, this was not the key takeaway.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The key insight to highlight from Jensen’s keynote was a claim Jensen made: <b>he projected that CapEx per GW would grow from $50 billion today to around $100 billion soon</b>, clarifying that AI hardware is not only not becoming cheaper, but it’s also more expensive than ever.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">How companies expect to make money from this business, I truly don’t know. I always say that I don’t believe Anthropic or OpenAI can be profitable as long as NVIDIA, Hynix, and the semiconductor players are commanding huge gross margins.</p><p class="paragraph" style="text-align:left;">While AI software is commoditized, or at least there are several players competing, which forces prices to fall, <b>these same companies are paying larger premiums than ever to access the hardware intended to train and run these models</b>.</p><p class="paragraph" style="text-align:left;">Jensen would push back, saying that tokens/watts are falling, so every GW allows for much larger revenues, <b>but customers around the world are already struggling with token costs today</b>. So either our ability to generate tokens increases by several orders of magnitude per dollar, or this business won’t make sense for the foreseeable future.</p><p class="paragraph" style="text-align:left;">As Google shows below in the market section, <b>AI will require substantial liquidity to survive over the next few years</b>, including the inevitable participation of all of us, willing or not.</p><p class="paragraph" style="text-align:left;">Because if you think you have a choice, you don’t, because OpenAI and Anthropic are getting into your favorite index funds, and fast (Nasdaq has reduced the required time as a public company to get inserted from one year to 15 days ahead of the SpaceX IPO).</p><p class="paragraph" style="text-align:left;">And even if you don’t own index funds, they&#39;ll still get into your 401(k)s through index funds. Directly or indirectly, you’re going to be an owner of these companies whether you like it or not.</p><p class="paragraph" style="text-align:left;">And listen, “they”, and I mean the powers that be, be that the US, Germany, or Spain, don’t have a choice;<b> they need to make all of us part of their big bets to sustain them, in the same way they are going to push crypto stablecoins down our throats eventually, too</b>, in order to distribute this unpayable debt that the largest economies in the world have amassed.</p><p class="paragraph" style="text-align:left;">And even if you somehow avoid all of that, <b>you’re still going to pay with inflation</b>, because the combination of trillions of AI private credit and government debt has to be paid, and the only way is to deflate the value of our currencies.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>BIG TECH</b></span><br>Microsoft Also Had a Big Event, ‘Build’</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b2c02648-1f17-48a5-aa7b-bf1addde1dbd/image.png?t=1780487080"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://arena.ai/leaderboard/image-edit?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://news.microsoft.com/build-2026-live-blog/microsoft-build-2026-live/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">Yesterday, Microsoft had its Build 2026 event</a>, and I must say I am happy with what I saw. </p><p class="paragraph" style="text-align:left;">Microsoft announced <b>MAI-Thinking-1</b>, its first in-house reasoning model, described as a 35-billion-parameter model for multi-step reasoning, long-context work, and code generation. It also introduced MAI-Image-2.5, MAI-Transcribe-1.5, MAI-Voice-2, and MAI-Code-1-Flash for GitHub Copilot and VS Code, <b>up to 7 highly competitive models from scratch</b>.</p><p class="paragraph" style="text-align:left;">This shows that, finally,<b> the acqui-hire of Inflection two years ago is starting to bear fruit</b>. For instance, their image model now ranks second only to GPT-image 2 in a popular image-editing benchmark (thumbnail).</p><p class="paragraph" style="text-align:left;">Agents were also a highlight. <b>Microsoft introduced Microsoft Scout</b>, an always-on personal work agent built on OpenClaw and Work IQ, designed to operate across tools such as Teams, Outlook, OneDrive, and SharePoint. <b>You know my skepticism with current agents</b>, so we’ll see how that goes.</p><p class="paragraph" style="text-align:left;">Microsoft also previewed <b>Project Solara</b>, described as a chip-to-cloud platform for an “agent-first” computing model, and announced <b>Surface RTX Spark Dev Box</b>, a local AI development machine expected later in 2026 in the US, pending authorization.</p><p class="paragraph" style="text-align:left;">And on quantum computing, Microsoft introduced <b>Majorana 2</b>, saying its qubits are 1,000 times more reliable than the previous generation and that the company is targeting a commercially relevant quantum computer by 2029.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;"><a class="link" href="https://microsoft.ai/wp-content/uploads/2026/06/main_20260602_2.pdf?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">The paper they released on MA1 was super impressive</a>; a gold mine of research nuggets that tells me Microsoft AI is finally on the right track in terms of their AI efforts. The model was clearly not designed to score highly; it’s a Sonnet 4.6-level model with “just” one trillion parameters, but the process they used to train it signals sophistication and shows us a company that is finally serious about its internal AI efforts.</p><p class="paragraph" style="text-align:left;">And best of all,<b> committed to open-source too</b>.</p><p class="paragraph" style="text-align:left;">To highlight a particular value, <b>I was shocked by their super-low MFU of 20% despite being a training run</b>; this means GPUs were running at 20% of peak theoretical compute—it does not mean they used only 20% of the 8,000-strong cluster.</p><p class="paragraph" style="text-align:left;">This peak is unattainable, but such a low score, considering that<a class="link" href="https://arxiv.org/pdf/2407.21783?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow"> the Llama 3 team in 2023 had around 40%</a>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Unsurprisingly, the stock still fell after the event, quite a bit actually, which is what it has gotten all of us used to lately. However, in my opinion, that’s just short-termism by investors; if anything, <b>Microsoft improved in my eyes</b>, even if the outcomes of these efforts will take some time to show up in the P&L.</p><p class="paragraph" style="text-align:left;">But before we move on, I really have to mention two key metrics: one that Mustafa Suleyman mentioned, another that he disclosed without intending to (at least, on paper).</p><p class="paragraph" style="text-align:left;">For their MAI-Thinking 1 model, they claim that their <b>Maia inference system</b> (their new inference chip) <b>offers 1.4x better performance/dollar relative to NVIDIA’s GB200 server</b>.</p><p class="paragraph" style="text-align:left;">That is quite the claim, and I’m sure Jensen is going to call them out for it. But if true, NVIDIA investors should be worried unless Vera Rubin blows everyone and everybody out of the water. We’ll see.</p><p class="paragraph" style="text-align:left;">The other one merits a round of applause for me, because I quite literally nailed Mythos’ training budget, <b>at around 2×10</b><sup><b>27</b></sup><b> FLOPs</b>, <a class="link" href="https://medium.com/@ignacio.de.gregorio.noblejas/training-ais-for-a-millionth-of-the-cost-93ca0357690a?sk=2dd1ba77293107dcd4ff12649f60a642&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">as I predicted in this Medium article a week ago</a>, thanks to the fact Microsoft, well, told us:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/42083e64-7d63-4c88-ab17-c0b95827cd34/image.png?t=1780479501"/><div class="image__source"><span class="image__source_text"><p>Source: Microsoft</p></span></div></div><p class="paragraph" style="text-align:left;">But unless you’re an AI geek like me, that number probably doesn’t say much beyond the fact that it represents a number with 27 zeros, which suggests that it’s a lot.</p><p class="paragraph" style="text-align:left;"><i>But how much?</i></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This means that <b>you could train 100 GPT-4s</b> with the amount of compute that was used to train Mythos.</p><p class="paragraph" style="text-align:left;">That is the scale of AI today, and as we’re about to see with the Vera Rubin piece below, it’s only going to get crazier, which raises the question: <i><b>Is training on a gazillion tokens really the only way we have to improve AIs?</b></i></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>SPENDING</b></span><br>Uber Closes the Faucet</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">Uber has limited </a><a class="link" href="https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">its per-user AI spending to $1,500</a> after blowing past its annual budget in less than four months. It’s not surprising, given that both its CTO and COO voiced concerns about the ludicrous spending one can rack up with AI.</p><p class="paragraph" style="text-align:left;">In particular, <b>the former even added that they were struggling to see returns on such spending</b>, making this budget constraint a matter of time and leading to today’s decision.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">People may panic about this, but this is just normal. Expected. I got criticized a lot when I said ‘tokenmaxxing’ was a stupid strategy, as if generating more tokens would magically transform businesses.</p><p class="paragraph" style="text-align:left;">Yet the only meaningful transformation so far is your OpEx. Robert Solow once said, “PC are everywhere, but in the statistics.”</p><p class="paragraph" style="text-align:left;">I now propose the AI version:</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>LAW</b></span><br>Senator Sanders Wants 50% of AI for Everyone</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://x.com/BernieSanders/status/2061631422188626083?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">As published by Bernie Sanders on X</a>, the senator said he will introduce a bill <b>to give the public a 50% ownership stake in the largest AI companies in America</b>. He said the goal is to ensure that wealth created by AI is used broadly and to give the public the power to block company decisions that could harm Americans.</p><p class="paragraph" style="text-align:left;">Sanders’ Senate site identifies the proposal as the <b>American AI Sovereign Wealth Fund Act</b>. It would create a sovereign wealth fund through a <b>one-time 50% tax paid in stock, not profits, </b>from major AI companies such as <b>OpenAI, Anthropic, and xAI (SpaceX)</b>.</p><p class="paragraph" style="text-align:left;">Under the proposal, the federal government would receive voting shares and equal board representation at covered companies. Sanders argues the fund could grow with the value of AI firms and eventually support direct public benefits or public programs.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I’m all for finding ways to redistribute the benefits of AI across society. I also entertain the idea that these companies stole our data to train their models and that they are, in a way, in debt to society.</p><p class="paragraph" style="text-align:left;"><b>But fighting theft with theft, which is what this is, is not the solution</b>.</p><p class="paragraph" style="text-align:left;">Even if I’m not necessarily super pro-taxes (I pay an eye-watering amount of taxes in Spain relative to what I earn, basically crossing the confiscation barrier at this point),<b> I understand their value and believe they make much more sense in this case:</b> tax positive outcomes; don’t intervene in the search for justice.</p><p class="paragraph" style="text-align:left;">And to be clear, <b>I do think AI incumbents have to be very careful about inequality of outcomes</b>, or they are going to suffer massive social rejection. Maybe UBI (Universal Basic Income) could be another option, too.</p><p class="paragraph" style="text-align:left;">Something will have to be done eventually if AI is so transformational. But that something should not be theft.</p><p class="paragraph" style="text-align:left;"><i>Would I support a purely economic intervention?</i> That i<i>s, having the US taxpayers buy a 50% stake by paying $500 billion to OpenAI or Anthropic?</i></p><p class="paragraph" style="text-align:left;">Hell no.<b> That is a bailout these companies have not earned the right to.</b></p><p class="paragraph" style="text-align:left;">I don’t think taxpayers would agree to invest in companies that are so massively in debt, either. And I don’t think they should.</p><p class="paragraph" style="text-align:left;">But once these companies have positive cash flows, maybe we could discuss how to ensure societal benefits. But the solution can’t be theft, which is precisely what Bernie wants.</p><p class="paragraph" style="text-align:left;"><b>As a European, what Bernie says sounds super reminiscent of what European politicians say. </b>And that’s not a good thing. Far from it.</p><p class="paragraph" style="text-align:left;">Please don’t let US lawmakers make the same mistakes that have condemned Europe to irrelevance and ostracism. <b>Don’t let politicians ruin the US as they did with Germany’s industrial base</b>, or with Europe’s energy security, just because an RBMK reactor with a positive void coefficient (which doesn’t exist anymore) and three irresponsible Russian engineers led to the Chernobyl disaster 40 years ago. </p><p class="paragraph" style="text-align:left;">European politicians understood regulation and interventionism as the way to progress and wealth redistribution, <b>and all they have achieved is a dying continent at the mercy of the US and China</b>. I fear the US could fall into the same mistake.</p><p class="paragraph" style="text-align:left;">I travel a lot to the US, <b>and I see the same ideas that destroyed Europe’s future being thrown around too lightly</b>. And yes, there’s a very real issue with inequality in the US, with the top 1% owns 37% of income and the bottom 50% owns only 2.5% of wealth.</p><p class="paragraph" style="text-align:left;">That is not sustainable. But theft is never the answer.</p><p class="paragraph" style="text-align:left;">And the funniest thing of them all: <i>has AI actually proven to be that incredibly unequalizing force?</i> Something that, in the words of Senator Sanders, “<i>could become smarter than us and function independently of our control” and thus warrant intervention?</i></p><p class="paragraph" style="text-align:left;">As I’ve reiterated countless times, <b>this is just doomer porn</b>. But just like I don’t think the US Government should steal 50% of a company from its owners, <i>wouldn’t it be funny if it happened to the same people who pushed the doomer narrative in the first place?</i></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>STOCK MARKET</b></span><br>Anthropic, Ready to IPO?</h2><p class="paragraph" style="text-align:left;">Quite possibly the news of the week, <a class="link" href="https://www.anthropic.com/news/confidential-draft-s1-sec?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">Anthropic has filed its confidential S-1 </a>for review to reserve its right to go public.</p><p class="paragraph" style="text-align:left;">The company said it has confidentially submitted a draft Form S-1 registration statement to the <b>US Securities and Exchange Commission</b> for a proposed IPO of its common stock.</p><p class="paragraph" style="text-align:left;">The filing gives Anthropic the option to go public after the SEC completes its review, <b>but the company said the offering remains subject to market conditions and other factors</b>.</p><p class="paragraph" style="text-align:left;">Anthropic did not disclose how many shares it may offer or the expected price range. The announcement was made under <b>Rule 135 of the Securities Act</b>, meaning it is not an offer to sell securities or a solicitation to buy them.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Well, as a matter of fact,<b> the company didn’t disclose a single thing</b>. Extremely secretive despite having an alleged booming business “that will be profitable this quarter”, <i>right?</i></p><p class="paragraph" style="text-align:left;">This industry is just so full of shit it’s actually amusing to me.</p><p class="paragraph" style="text-align:left;">For what it’s worth, it’s a great company, nobody doubts that. But just like with SpaceX, the problem will be its valuation (and surely the same will apply to OpenAI).</p><p class="paragraph" style="text-align:left;">For context, <b>they have just closed a funding round at $965 billion</b>, which means it has to IPO way higher than that, easily above $1.5 trillion, which would make it as valuable as Meta, and maybe even closer to $2 trillion. Make that make sense.</p><p class="paragraph" style="text-align:left;">If that’s the case, that’s hilariously overpriced, <b>but I do think some investors will be blind enough to purchase it at that price</b>. We’ll see.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>STOCKS</b></span><br>Google’s Historical $85 Billion Equity Issuance</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.reuters.com/legal/transactional/alphabet-raise-80-billion-equity-capital-ai-spending-2026-06-01/?utm_source=chatgpt.com" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://www.reuters.com/legal/transactional/alphabet-raise-80-billion-equity-capital-ai-spending-2026-06-01/?utm_source=chatgpt.com" target="_blank" rel="noopener noreferrer nofollow">reported by Reuters</a>, Alphabet, Google&#39;s parent, plans to raise <b>$85 billion</b> in equity to fund the expansion of its AI infrastructure.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Alphabet says the money will support AI-related computing capacity, data centers, and custom chip development. The move follows a sharp increase in planned capital spending, with Alphabet’s 2026 capex now guided at <b>$180 billion to $190 billion</b>, up $5 billion from the previous estimate. Alphabet shares fell after the announcement, and the gap with NVIDIA, the most valuable company on the planet, is now almost a trillion, up from “only” $200 billion a few weeks ago.</p><p class="paragraph" style="text-align:left;">Interestingly, these $85 billion add to the $85 billion they’ve already borrowed over the last year.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">A few weeks ago, <b>I showed a graph showing that Hyperscaler FCF </b>(Free Cash Flow, the amount of cash they have available for discretionary spending) <b>was down sharply</b>, meaning these companies were literally running out of money.</p><p class="paragraph" style="text-align:left;"><b>Google’s $85 billion raise clearly indicates that they are “all in” on AI and will do </b><b>whatever it takes to win</b>, even at the expense of shareholders like me.</p><p class="paragraph" style="text-align:left;">Probably the highlight here is the participation of Berkshire Hathaway, which seems to have made up its mind on which AI horse they are betting on amongst the Hyperscalers (I wouldn’t count Apple as an AI bet just yet).</p><p class="paragraph" style="text-align:left;">This is like IPOing again, <b>and $85 billion less that could have gone to Anthropic and OpenAI</b>.</p><p class="paragraph" style="text-align:left;"><i>Do the public markets have the $500 billion we’ll need to provide to Google, SpaceX, Anthropic, and OpenAI?</i> I don’t think people realize how uncertain this question is.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>HARDWARE</b></span><br>Vera Rubin is Finally Here</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/cafa7695-b81e-4307-a35d-a7c2ec1b4174/image.png?t=1780389564"/></div><p class="paragraph" style="text-align:left;">As mentioned above, <b>Jensen Huang has confirmed that both Microsoft and Dell/CoreWeave have deployed their first Vera Rubin NVL72 servers</b>, the next-generation AI chips.</p><p class="paragraph" style="text-align:left;"><b>And the differences in performance</b>, especially relative to memory and memory bandwidth (the key bottlenecks in inference, which is the primary driver of compute) to previous generations, <b>are simply astonishing</b>.</p><p class="paragraph" style="text-align:left;">For example, in Rubin, <b>each GPU inside a server can communicate more data to one another than a Hopper GPU could within its own package</b>. This significantly increases the amount of data GPUs can share, thereby considerably elevating performance.</p><p class="paragraph" style="text-align:left;">Think of it this way. In inference, you are bottlenecked by how much data can be moved. This means that your GPUs are “waiting” idle for data to arrive—they aren’t really idle, but running at very low capacity relative to what they could be doing. Therefore, by increasing the amount of data that can be shared, you’re significantly alleviating the bottleneck,<b> leading to more tokens per second and per watt</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The one thing I need to point out is that some people are saying that in Vera Rubin, the speed between GPUs is faster than inside the package of an H100. <b>But speed here is the wrong word; it’s bandwidth.</b></p><p class="paragraph" style="text-align:left;">GPU-to-HBM speed, the time it takes a GPU to read data from its HBM chips, is on the order of hundreds of nanoseconds, easily an order of magnitude faster than the time it takes to reach other GPUs.</p><p class="paragraph" style="text-align:left;">People confuse bandwidth (how much data can be sent per second) with speed (how fast it moves) all the time. Don’t be that person.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>GOOGLE</b></span><br>Gemma 4 12B is Here</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/da086df6-c12d-4140-9613-75003dbdc21c/image.png?t=1780510064"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-and-uber-the-us-s-cookie-monsters" target="_blank" rel="noopener noreferrer nofollow">Google has introduced Gemma 4 12B</a>, a mid-sized open model designed to run multimodal AI locally on laptops. The model handles text, vision, and native audio inputs, and Google says it can run on consumer hardware with <b>16GB of VRAM or unified memory</b>.</p><p class="paragraph" style="text-align:left;">Google describes Gemma 4 12B as a bridge between its smaller, edge-focused E4B model and its larger 26B Mixture-of-Experts model, with benchmark performance “nearing” that of the 26B model <b>while using less than half the memory footprint</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Besides looking like a great model for its size, <b>a key technical change is its encoder-free multimodal architecture</b>: instead of separate encoders for images and audio, visual and audio inputs are integrated more directly into the language model backbone. For audio, Google says it removed the audio encoder and projects raw audio into the same space as text tokens.</p><p class="paragraph" style="text-align:left;">This is genuinely interesting and something that I believe will become much more common with smaller models.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>OpenAI Launches ‘Sites’</h2><p class="paragraph" style="text-align:left;"><b>OpenAI “Sites”</b> is part of the new <b>Codex</b> release, announced today. OpenAI describes Sites as a preview feature that l<b>ets Codex create and share interactive, hosted websites and apps.</b></p><p class="paragraph" style="text-align:left;">Sites can turn ideas, analysis, and plans into dashboards, planners, review workspaces, project boards, galleries, and lightweight tools. The sites can be shared with anyone in a user’s workspace through a URL.</p><p class="paragraph" style="text-align:left;">OpenAI said Sites is rolling out in preview for <b>Business and Enterprise teams</b> via the Codex app, and enterprise admins can enable it in admin settings. The company also said it is working with early partners, including <b>Wix, Base44, Replit, Lovable, Figma, Webflow, and Emergent, as it builds a Sites partner ecosystem.</b></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">AI is innovating once again with yet another AI site generator. I swear this entire industry is three products being reinvented again and again.</p><p class="paragraph" style="text-align:left;">What I will say is that the sites look incredibly clean, something OpenAI’s models have been really bad at historically. The reason seems to be its partner ecosystem, which might have helped improve such skills.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Rants aside, <b>OpenAI seems to be looking to turn Codex into an enterprise platform</b>, not just a tool to code, but a tool to build anything you want. For now, <b>that probably pulls it closer to a slop machine than something actually useful</b>; AIs still make a lot of mistakes.</p><p class="paragraph" style="text-align:left;">And to be fair, this reminds me all too much of Confluence, the Atlassian tool that lets teams have sites, projects have repos, and all that stuff that nobody can justify its value, but everyone pretends it does.</p><p class="paragraph" style="text-align:left;">Not trying to be overly cynical here, <b>but this feels extremely tailored to the status quo in corporate America rather than something that helps it progress</b>. Take this with a pinch of salt because I’m a professional corporate hater, though.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">A very eventful week this was. But here are the four most important takeaways in my view:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>China keeps innovating</b>. There’s no way around it. Whether it’s by folding logical circuits or by creating SOTA-ish models with a fraction of the US’s resources, China will compete no matter what.</p></li><li><p class="paragraph" style="text-align:left;"><b>AI is getting more expensive than ever</b>. From Mythos huge training run to Jensen increasing the cost per GW, AI, the technology that needs to be adopted by all of us to make this giant bet work, is ironically becoming more expensive, not cheaper; I’m sure this is a great sign! If not, ask Uber</p></li><li><p class="paragraph" style="text-align:left;"><b>Great models are getting smaller</b>. The silver lining is that while the frontier becomes more expensive, we are getting incredibly good at making great small models. <i>Wouldn’t it be ironic that, after all this money, most of the benefits of this technology were built around free models? </i>Well, that could actually happen.</p></li><li><p class="paragraph" style="text-align:left;"><b><i>Where is all this IPO money going from?</i></b> 2026 is going to test our ability as investors to provide liquidity to this industry. But here’s the thing: <b>don’t expect me to participate; not at those prices.</b></p></li></ol><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Until the next one!</b></span></p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=bef77665-6aa2-453d-a5e6-046ccda6a6bd&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>Why History Won&#39;t be Kind to AI</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/533a819c-ff8d-4a46-8f7e-d73a6bf7c22d/image.png" length="136228" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/why-history-won-t-be-kind-to-ai</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/why-history-won-t-be-kind-to-ai</guid>
  <pubDate>Sun, 31 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-31T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Why History Won’t be Kind to AI</h2><p class="paragraph" style="text-align:left;">I’ve long talked about AI’s limitations on a technical basis; <b>it’s not the magical technology some make it to be</b>. Huge potential, still mostly unmet.</p><p class="paragraph" style="text-align:left;">But it’s hard to deny it’s a transformational technology. The question is more like: <i>when? When will it change the world?</i></p><p class="paragraph" style="text-align:left;">We’ll go back in history to answer that question. And you might not like the answer I have for you. And at the end, I’ll give you my two cents on what I think will happen.</p><p class="paragraph" style="text-align:left;">You might not like that either.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">A PC, a container, and a steam engine walk into a bar</h2><p class="paragraph" style="text-align:left;">We, humans, love to ignore history. But as Mark Twain (allegedly) once said: <i>“History doesn’t repeat, but it rhymes.” </i>So,<i> what does history tell us about AI’s chances to change the world?</i></p><p class="paragraph" style="text-align:left;">And the short answer is that it will, <b>but it will take time, maybe too much time</b>.</p><h3 class="heading" style="text-align:left;">The Solow Paradox and Shadow AI</h3><p class="paragraph" style="text-align:left;">Nobel Laureate Robe<i>r</i>t Solow had a great quote back in 1987: <i>“You can see the computer age everywhere but in the statistics.”</i></p><p class="paragraph" style="text-align:left;">At that time, it had been almost two decades since a group of engineers at Intel created the microprocessor in 1971, a development that eventually led to the personal computer.</p><p class="paragraph" style="text-align:left;">But in 1990, three years after saying those famous words, <b>just 20 million personal computers were sold</b>. But PCs are hardly the only example; <b>electricity is another great one</b>. While the first commercial power station was built in 1882, it was not until 1920, four decades later, <a class="link" href="https://www.nber.org/papers/w11093?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">that electricity surpassed steam as the dominant form of horsepower in the US economy</a>.</p><p class="paragraph" style="text-align:left;">We can go even further back to another great example: <b>steam engines</b>. In their case, <b>the diffusion took even longer</b>. James Watt patented his steam engine in 1769.</p><p class="paragraph" style="text-align:left;">By 1830, its penetration into the economy, measured by productivity, was still trivial, <a class="link" href="https://warwick.ac.uk/fac/soc/economics/research/workingpapers/1989-1994/twerp339.pdf?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">and we didn’t see its real impact until the third quarter of the nineteenth century</a>, almost a century after Watt’s patent.</p><p class="paragraph" style="text-align:left;">Circling back to PCs, <b>it wasn’t until the mid 1990s that the PC revolution really took off</b>, and it took an outrageous 3 decades to reach 50% PC adoption.</p><p class="paragraph" style="text-align:left;">It was particularly slow in its contribution to the most important metric of them all: productivity. For the first two decades since its conception, the Internet revolution had very little impact on macro numbers, if any.</p><p class="paragraph" style="text-align:left;">And it might be the case that it happens to AI, too.</p><h3 class="heading" style="text-align:left;">AI’s downsizing problem</h3><p class="paragraph" style="text-align:left;"><i>What if a technology is just too good at its job? </i>Several months ago (maybe more than a year, actually), <b>I argued that AI could actually decrease the size of the economy</b>. </p><p class="paragraph" style="text-align:left;">The rationale was based on two ideas:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">I believed that, over time, <b>AI would</b><b> be extremely deflationary and shrink the output of those industries it affected</b>.</p></li></ol><p class="paragraph" style="text-align:left;">With enough sophistication (not the case today), AI lowers barriers to competition in its areas of impact, commoditizing those industries and leading to price wars in which prices fall faster than demand grows.</p><p class="paragraph" style="text-align:left;">In other words, while price reductions do increase demand (Jevons’ Paradox), because GDP (Gross Domestic Product) is measured as price times quantity, <b>if prices fall faster than quantities rise, the overall market shrinks.</b></p><p class="paragraph" style="text-align:left;">Agriculture is always a great example. Despite producing more food than at any time in history, agriculture’s impact on major economies has nosedived (<a class="link" href="https://ourworldindata.org/grapher/agriculture-share-gdp?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">except </a><a class="link" href="https://ourworldindata.org/grapher/agriculture-share-gdp?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">in some African countries</a>, where agriculture accounts for a minimal share of national GDP, with the US’s share falling below 1%).</p><p class="paragraph" style="text-align:left;">Nonetheless, while in 1840 agriculture accounted for more than 60% of total jobs in the US, most developed economies have since transitioned to services as the main driver of employment.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/533a819c-ff8d-4a46-8f7e-d73a6bf7c22d/image.png?t=1780217771"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://ourworldindata.org/structural-transformation-and-deindustrialization-evidence-from-todays-rich-countries?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;">The answer as to why is quite simple: huge increases in productivity (like the one we see in Sweden below) decreased prices severely, way more than demand rose, <b>overall decreasing agricultural output.</b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f9621acd-ff49-4a24-a419-49492fc4c8b6/image.png?t=1780217874"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://ourworldindata.org/grapher/labor-productivity-agriculture-sweden?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><ol start="2"><li><p class="paragraph" style="text-align:left;"><b>Baumol’s cost disease</b></p></li></ol><p class="paragraph" style="text-align:left;">Adding to the fact that many transformational technologies can shrink output, we have to add Baumol’s cost disease to the picture.</p><p class="paragraph" style="text-align:left;">Baumol’s cost disease is the idea that some sectors become more expensive not because they are getting worse, but because they cannot raise productivity as fast as the rest of the economy.</p><p class="paragraph" style="text-align:left;">Productive sectors have higher wages, attracting more talent. To retain talent, less productive sectors raise wages, <b>creating a downward spiral of declining productivity</b>.</p><p class="paragraph" style="text-align:left;">A beautiful example is a string quartet. </p><p class="paragraph" style="text-align:left;">In the 18th century, you needed four musicians and a certain amount of time to play a Mozart quartet. Today, <b>you still need four musicians and roughly the same time</b>. Productivity has barely improved.</p><p class="paragraph" style="text-align:left;">But those musicians live in an economy where other sectors, like manufacturing or software, have become far more productive, so wages across the economy rise. <b>To keep musicians from leaving for better-paid jobs elsewhere, orchestras must also raise wages</b>, even though each performance is not much more “productive” than before.</p><p class="paragraph" style="text-align:left;">This has a big effect on the macro picture, <b>as a nation’s GDP becomes dominated by essential, low-productive sectors</b>. Healthcare, education, and industries where productivity doesn’t change as fast as others but are essential, concentrate the majority of the output because quantity is not negotiable (they are essential), and prices continue to rise over time, driven by non-productive wage hikes.</p><p class="paragraph" style="text-align:left;"><i>The lesson?</i> Fascinatingly, if AI’s impact is not pervasive across all industries, it may be extremely hard to see, which poses a very real danger to the industry: <b>AI’s invisible output.</b></p><h3 class="heading" style="text-align:left;">AI’s invisible output</h3><p class="paragraph" style="text-align:left;">One thing most tech bros in San Francisco miserably fail to understand is that <b>being useful does not automatically mean being more productive</b>. At least not in the way we measure productivity.</p><p class="paragraph" style="text-align:left;">Everyone and their grandma can see AI is useful, but justifying the alleged massive rises in productivity is proving way harder than we thought.</p><p class="paragraph" style="text-align:left;">If we are spending several basis points of global GDP on AI, we should expect in return a technology that massively transforms productivity, <b>meaning the world starts to do a lot more with a lot less</b>.</p><p class="paragraph" style="text-align:left;"><i>But are we?</i> We’ll talk about costs later, but the other big issue we face is what some are calling ‘invisible output’. In the sea of tokens being generated worldwide,<b> many productivity gains “benefit no one</b>.” Let me explain.</p><p class="paragraph" style="text-align:left;">What two years ago would have required consulting my lawyer for any trivial issue can now, at least for some of them, be handled by AI.</p><p class="paragraph" style="text-align:left;">As a self-employed person with two companies in Spain, I receive a lot of attention from ‘Hacienda’ (our IRS), receiving a decent amount of letters. These are famously cryptic, but I don’t need my lawyers anymore for such things; I just ask ChatGPT.</p><p class="paragraph" style="text-align:left;"><b>This means that my lawyer is not getting paid less</b>. Of course, one could take the other side of the argument and say, <i>“Well, but your lawyer can now attend to more people.” </i></p><p class="paragraph" style="text-align:left;">Maybe, but that’s the thing; there’s a certain threshold of capabilities AIs can take away from what before would have been transactional interactions feeding into our GDP output metrics.</p><p class="paragraph" style="text-align:left;">Don’t get me wrong, <b>you would be unprecedentally foolish to think you don’t need a human lawyer anymore</b>; it’s just that, in some way, AIs have taken the role of <i>“consigliere for all menial stuff”</i> that increases users’ productivity without appearing elsewhere.</p><p class="paragraph" style="text-align:left;">Or, if somewhere, it’s on the AI Lab’s top line, if not for the fact that they have massively commoditized their businesses or, in the cases of ChatGPT or Gemini Search, <b>offer many of such tokens for free</b>.</p><p class="paragraph" style="text-align:left;">But even if that ChatGPT supporting my business does appear in OpenAI’s top line (I am a paying subscriber), it’s on an entirely different sector and mixed with a lot of other stuff, making it almost impossible to measure accurately.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">It must be mentioned that one can also be critical of the measurements themselves. For example, <b>the way we measure the productivity of the public sector is, by definition, flawed.</b></p><p class="paragraph" style="text-align:left;">Since you don’t pay for many of those services, the only way we can measure them is by considering production costs and using volume as the output.</p><p class="paragraph" style="text-align:left;">A public school in Spain is “free”, so to measure its productivity, we just look at costs (wages and others) and measure output by the “volume of educational services delivered”, using metrics such as student-hours and even academic attainment. But as there’s no price, how productive they really are, how much value it’s delivering, well, it’s hard.</p><p class="paragraph" style="text-align:left;">Nevertheless, the fact that the number of new graduates is growing doesn’t say anything about the quality of education. In fact, in Spain, <b>we call this “titulitis”,</b> a phenomenon in which we have a huge number of “highly educated” unemployed workers while construction workers, carpenters, and plumbers have all the work they want, and more.</p><p class="paragraph" style="text-align:left;"><i>Is Spain’s public university sector really that productive if we’re sending most of these kids directly into unemployment?</i></p><p class="paragraph" style="text-align:left;">To be clear, I’m not saying public education is not a great success of society (I myself benefited from it during my undergraduate years, and I’m eternally thankful for it); <b>I’m just questioning the quality of the measures of productivity</b>.</p><p class="paragraph" style="text-align:left;">But leaving this quite hard-to-deal-with problem aside for a moment, let’s go back to history to answer: <i>why did it take so long for other technological disruptions to transform the economy?</i></p><h3 class="heading" style="text-align:left;">Understanding the WHY</h3><p class="paragraph" style="text-align:left;">In most cases, it was a mixture of three things: <b>a solution looking for a problem to solve</b>, <b>unsophisticated approaches</b>, and, of course, <b>costs</b>.</p><p class="paragraph" style="text-align:left;">On the latter, <b>steam engines and electricity were simply very early to the party</b>. In the steam engine’s case, its very slow diffusion was held back by fuel inefficiency and very slow price declines.</p><p class="paragraph" style="text-align:left;">This is the easiest example answer; <b>it just wasn’t worth it until it was</b>, and the technology had to develop and become cheaper to be fairly adopted.</p><p class="paragraph" style="text-align:left;"><i>Sounds familiar?</i></p><p class="paragraph" style="text-align:left;">Electricity’s case was a little bit more nuanced, and quite frankly, much more interesting.</p><p class="paragraph" style="text-align:left;"><b>Although the “War of the Currents” between Edison and Tesla didn’t help</b>, as Edison created a lot of fear around Tesla’s alternating current despite being much safer and, in hindsight, the only viable solution for long-distance transmission (we need very high voltages to transmit enough power without current, and thereby losses, being small enough, and also we only knew how to step down voltage using AC transformers),<b> in this case the technology was quite ready.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">In particular, not until we reframed our factories to electricity did electricity become adoptable. Surprisingly, <b>electricity was only widely adopted in the 1920s, almost three decades after the first transformer was built.</b></p><p class="paragraph" style="text-align:left;"><i>What changed? </i>Unit drives.</p><p class="paragraph" style="text-align:left;">Not until factories were retrofitted to handle several individual unit drives (motors with their own, individual control units), forcing a complete redefining of the factory’s layout, <a class="link" href="https://warwick.ac.uk/fac/soc/economics/research/workingpapers/1989-1994/twerp339.pdf?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai" target="_blank" rel="noopener noreferrer nofollow">did electricity finally make sense</a>.</p><p class="paragraph" style="text-align:left;">Put another way, part of the delay in exploiting the potential industrial productivity gains offered by electricity was simply due to the durability of old manufacturing plants that embodied technology adapted to the regime of mechanical power derived from water and steam.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><i>Sounds familiar?</i></p><p class="paragraph" style="text-align:left;">With PCs, it was actually mostly about being a solution looking for a problem to solve. <b>A hilarious example of this was Apple’s marketing at the time</b>, where they couldn’t even describe the use case and “begged” users to tell them what it was.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a8f8094f-c8e7-489c-a016-89c6ea16a17b/image.png?t=1780221363"/></div><p class="paragraph" style="text-align:left;"><i>Sounds familiar?</i></p><p class="paragraph" style="text-align:left;">It had some price dynamics, too: not until Intel was forced to drop prices due to competition and computers became affordable (especially with MOS Technologies selling its 6502 for $25, $150 in today’s dollars), was Steve Wozniak “allowed” to tinker with the technology, <b>leading to the first Apple computer prototype.</b></p><p class="paragraph" style="text-align:left;">Now that we know what determines diffusion, clear use cases, system readiness, and cost, it’s time to see how AI is faring in each one. <b>And, ladies and gentlemen, it’s not good</b>.</p><p class="paragraph" style="text-align:left;">Behind the paywall, we analyze AI’s situationship across all levers of diffusion, while also using history again, a last-century, <b>lesser-known revolution that offers incredible insight into how most AI investments are going to turn out</b>.</p><p class="paragraph" style="text-align:left;">Because AI will succeed. Most investors, however…</p></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=why-history-won-t-be-kind-to-ai">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=6a3c2418-c585-435e-a68f-59681ef5d831&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>China Sends a Message, Opus 4.8, &amp; More</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8298bdda-23c4-43b8-91fd-ed2a94b6eb42/image.png" length="910693" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/china-sends-a-message-opus-4-8-more</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/china-sends-a-message-opus-4-8-more</guid>
  <pubDate>Thu, 28 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-28T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back!</b> Today, we tackle a broad list of topics.</p><p class="paragraph" style="text-align:left;">From your weekly ratio of <b>AI solving maths problems</b>, Huawei’s new <b>chip breakthrough</b>, <b>Uber</b> sounding the token alarm, a trillion-dollar company going up by <b>19% on a single day</b>, and <b>Opus 4.8</b>, among other interesting news.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>Google’s AlphaProof Nexus Solves 9 Erdős Problems</h2><p class="paragraph" style="text-align:left;">As published on <b>arXiv</b>, <a class="link" href="https://arxiv.org/pdf/2605.22763v1?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">a Google DeepMind-led paper introduces AlphaProof Nexus</a>, a framework that uses large language models and the Lean proof assistant to search for formally verified mathematical proofs.</p><p class="paragraph" style="text-align:left;">The authors report that their strongest agent solved <b>9 of 353 open Erdős problems</b>, <b>including two questions that had been open for 56 years</b>, at an inference cost of <b>“a few hundred dollars” per problem</b>. It also proved <b>44 of 492 OEIS conjectures</b> after autoformalization and manual review.</p><p class="paragraph" style="text-align:left;">The idea is that AlphaProof Nexus takes a Lean theorem with missing proof steps, lets LLM-based prover subagents revise proof sketches, and checks progress through Lean. In layman’s terms, it lets an agent “guess and verify” different approaches to solving maths problems.</p><p class="paragraph" style="text-align:left;">To be clear, Google’s proof-solver, Alphaproof, has existed for a while now. The difference is that Nexus gives Alphaproof as a tool to another AI agent.</p><p class="paragraph" style="text-align:left;"><i>But why does this work so amazingly well?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The key is the use of the Lean proof assistant and compiler that can automatically verify correctness. This is a perfect counterbalance to an LLM’s biggest problem: hallucinations, as well as helping it improve under automatic verification.</p><p class="paragraph" style="text-align:left;"><b>Remember that in AI, we can only learn what can be measured</b>. AIs excel at those areas where verification is simple; areas where discerning a model’s response quality is simple or even automatic, as is the case in maths.</p><p class="paragraph" style="text-align:left;">Lean enables the AI to engage in a learning loop in which it can try new ways of solving problems and receive automatic feedback, <b>making learning a matter of computation</b>. </p><p class="paragraph" style="text-align:left;">As I always say, there’s a reason we have amazing coding agents while AIs are terrible at writing. It’s not magic; <b>it’s one task that is easily verifiable versus one that is very hard</b> (what is great writing? It’s highly subjective).</p><p class="paragraph" style="text-align:left;">But to me, <b>the highlight here is the costs</b>. They took only a couple of hundred dollars per solved problem. That is great news because it means frontier prices on hard problems are falling.</p><p class="paragraph" style="text-align:left;">On the flip side, AI getting better at verifiable domains is not surprising; it’s maths, a matter of optimizing against a known and measurable objective. There’s zero reason to believe an AI can’t optimize against such problems.</p><p class="paragraph" style="text-align:left;"><i>But finding a way to train AIs on non-verifiable domains?</i> Well, that’s still a mystery to this day.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>HARDWARE</b></span><br>Huawei’s New Chip Breakthrough</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8298bdda-23c4-43b8-91fd-ed2a94b6eb42/image.png?t=1779894630"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://f004.backblazeb2.com/file/chinaxiv/english_pdfs/chinaxiv-202605.00224.pdf?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">Huawei has published a paper</a> that is making quite the noise. The claim is that it has found a new way to scale chips despite US export controls: <b>LogicFolding</b>.</p><p class="paragraph" style="text-align:left;">The problem China faces is simple. Advanced chips have historically improved by shrinking transistors, which lets companies pack more compute into the same chip area. But shrinking transistors below the 7–5 nanometer range requires extremely advanced EUV lithography tools, mostly made by ASML, <b>which China cannot access due to US-imposed export controls</b>.</p><p class="paragraph" style="text-align:left;">Therefore, Huawei’s answer is not to shrink the transistor, <b>but to change the chip’s geometry</b>. Instead of spreading circuits only across a flat surface, LogicFolding stacks compute logic vertically. In theory, this increases density, shortens some wires, reduces power lost to wiring delays, <b>and allows the chip to do more work without needing a more advanced manufacturing node</b>.</p><p class="paragraph" style="text-align:left;">In other words, <b>Huawei says it achieved a major increase in density without moving to a smaller transistor node</b>, something this time the US can’t prevent with US controls. This is not the same as catching up to TSMC or NVIDIA, nor does it make export controls irrelevant, <b>but it surely hurts in Washington</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The bottom line is that <b>China now has a way to competitively scale compute </b><b>on-chip without ASML&#39;s EUV tools</b>, at least for smartphones.</p><p class="paragraph" style="text-align:left;">Huawei’s reported density numbers for its 2026 smartphone CPU (graph above) appear competitive with, or even better than, Apple’s 2024 A18 Pro chip (iPhone 16 and MacBook Neo) despite Apple using a much more advanced TSMC node (3 nanometers).</p><p class="paragraph" style="text-align:left;"><b>Huawei is targeting AI processors by 2030-2031 with density comparable to future 1.4A-class chips</b>. We can&#39;t jump to conclusions so soon, but if true, China’s 2030 chips could be competitive with US 2028–2030 chips on a chip-by-chip basis, massively closing the gap.</p><p class="paragraph" style="text-align:left;">As we have discussed multiple times, China is already competitive at the server/system level today, so if that event materializes, <b>it could be hard to see who&#39;s ahead</b>.</p><p class="paragraph" style="text-align:left;">All of this, again, without access to EUV litho tools (yet).</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>COSTS</b></span><br>Uber Sounds the Token Alarm</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.youtube.com/watch?v=y_mQ6xLcKyc&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">In a recent podcast</a>, Uber’s COO said the <b>ride-hailing company isn’t seeing a clear increase in productivity from using AI coding services</b> despite their use by its engineering teams. That has prompted executives to discuss how to get a handle on token consumption costs. </p><p class="paragraph" style="text-align:left;"><i>“If you‘re not actually able to draw a direct line to how much useful features and functionality you’re shipping to your users, [the costs become] harder to justify,”</i> he said.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Next up, rain is wet. Jokes aside, <i>who could’ve known that ‘tokenmaxxing’, generating as many tokens as you can as a sign of productivity, was not a well-thought-out strategy?</i></p><p class="paragraph" style="text-align:left;">Needless to say, it was the CTO of that same company who first raised the alarm that their AI spending had skyrocketed far beyond what they had anticipated, <a class="link" href="https://www.theinformation.com/newsletters/applied-ai/uber-cto-shows-claude-code-can-blow-ai-budgets?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">burning through an entire yearly budget by March</a>.</p><p class="paragraph" style="text-align:left;">Now, the world is realizing what readers of this newsletter have known for months: <b>AI is much more expensive than we thought</b>, and not only that, but it will get worse as companies stop subsidizing tokens and start charging the real value.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And to be clear, this is hardly a call to stop using AI.</p><p class="paragraph" style="text-align:left;">Instead, <b>it’s about being smart about your AI use</b>. Start measuring return on investments, like in every single technology you’ve ever used. “Selling intelligence” makes it look like you have to ‘tokenmaxx’.</p><p class="paragraph" style="text-align:left;">Well, no, because if you do that, <b>you’re going to start ‘bankruptmaxxing’.</b></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MEMORY</b></span><br>Micron Goes Up 20% on a Single Day</h2><p class="paragraph" style="text-align:left;">As I write these words, the three DRAM memory companies, Samsung, SK Hynix, and Micron, <b>have crossed the psychological threshold of one trillion dollars</b>, a remarkable growth rate. For instance, Hynix has grown 11x in a year.</p><p class="paragraph" style="text-align:left;">But Micron set all the records three days ago when it went up by 20%, or $200 billion, in a single day after UBS upgraded its price target to $1.6k, roughly double what it was at the time, which would mean more than double where it was. And the stock flew.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">We’ve talked a lot about the importance of memory in the AI trade. They are key to both lines of progress: <b>making models bigger requires more memory capacity</b>, and <b>making sequences longer requires more memory capacity and bandwidth</b>.</p><p class="paragraph" style="text-align:left;">But I can’t help but feel uneasy that a stock already worth $800 billion at the time can go up 20% on a single analyst quote. <i>Are we approaching the peak?</i></p><p class="paragraph" style="text-align:left;">Honestly, <i>who knows?</i> I will only say this: I have more than 25% of my liquid net worth on Samsung and Hynix, so we&#39;d better not be at the peak.</p><p class="paragraph" style="text-align:left;">Interestingly, <b>both Samsung and Hynix trade at lower price-to-earnings multiples than the average S&P 500 stock</b> (Micron has a much higher PE), and their forward PEs (multiples over projected future earnings) are still insultingly low because these companies are going to make so much money next year.</p><p class="paragraph" style="text-align:left;"><i>How much, you say?</i> Just look at Morgan Stanley’s very high-level estimate of the Bill of Materials for the upcoming VR200 chip (Vera Rubin from NVIDIA). Memory has grown by 6x from one generation to the next.</p><p class="paragraph" style="text-align:left;"><b>And that memory does not include HBM</b>, which is included inside the GPU line item. At least for the next two years, these three companies are going to print money.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3e15833e-ca7f-4c3c-b2e6-bb5ce0beae31/image.png?t=1779869322"/><div class="image__source"><span class="image__source_text"><p>Source: Morgan Stanley</p></span></div></div><p class="paragraph" style="text-align:left;">The question, as always with these stocks, is whether we should be valuing them based on PE or PB. Historically, due to their cyclical nature, <b>semiconductors have been evaluated by multiple-to-book value</b> (assets - liabilities), because you couldn’t discount future cash flows because they were so uncertain.</p><p class="paragraph" style="text-align:left;">At PB, their multiples sit around 10 on average. But here’s the thing: even if you insist on valuing them by book, you’re still going to have to rerate them because their books are growing incredibly fast due to the amount of cash they are receiving.</p><p class="paragraph" style="text-align:left;">Importantly, new investments are not being financed, <b>so the asset base is growing rapidly while liabilities grow very little, and thus the book value will continue to grow</b>.</p><p class="paragraph" style="text-align:left;">However, <b>many people now actually believe memory is a secular market,</b> <b>and these companies will finally have predictable revenues </b>(unlikely considering it’s still hardware), but here’s my two cents: <b>it will remain cyclical but with much higher frequency</b>, meaning it might remain cyclical, but new cycles will come fast.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>DEBT</b></span><br>SoftBank is Playing a Dangerous Game</h2><p class="paragraph" style="text-align:left;">I’ve talked about AI leaning more and more into debt. But the following news just sets a new degenerate standard. <a class="link" href="https://www.bloomberg.com/news/features/2026-05-19/softbank-founder-son-s-devotion-to-openai-s-altman-spooks-some-insiders?embedded-checkout=true&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">New Bloomberg reporting</a> explains that SoftBank has borrowed against its own OpenAI shares… to buy more of them.</p><p class="paragraph" style="text-align:left;">It’s like the start of a joke, except that it’s not one.</p><p class="paragraph" style="text-align:left;"><b>SoftBank has committed roughly $60 billion to OpenAI</b>, and internal advisors who questioned the size of the bet say founder Masayoshi Son shut them down. Former SoftBank insider Habib Imam described the position as &quot;a bet on a worldview about AGI&quot; and added, &quot;you can&#39;t hedge a worldview.&quot; </p><p class="paragraph" style="text-align:left;">To fund the commitment, SoftBank sold its remaining Nvidia stake, took out a $40 billion bridge facility, and layered on a margin loan, all at around 8% interest. SoftBank&#39;s last bet of this scale was WeWork, which imploded in 2019.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">These bets make sense (I wouldn’t be making those bets either way, though) as long as the price of the underlying shares keeps going up. Because if it starts going down, lenders are going to come crushing down on you.</p><p class="paragraph" style="text-align:left;"><b>This is just one additional reason to view the IPOs of Anthropic and OpenAI as the ‘it moment’ for the industry</b>. If they are a resounding success, we’re into something. If they show the slightest flakiness, oh boy.</p><p class="paragraph" style="text-align:left;">And to be clear, the issue is not with the Hyperscalers; they could cut back on investments somewhat if those IPOs fail, but they will survive, and investors know it.</p><p class="paragraph" style="text-align:left;">The biggest problem is the debt side of things; what happens to all the AI companies that are largely dependent on external financing to survive, <i>the CoreWeaves of the world?</i></p><p class="paragraph" style="text-align:left;">Because guess what, they are great businesses, but no matter how great they are, <b>they are still massively cash-flow negative</b>. My enduring feeling is that there’s really too much money at stake here to let all this come down; people just don’t want to hold cash and will go to unprecedented lengths to justify these investments.</p><p class="paragraph" style="text-align:left;">I could be wrong, though.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>LAW</b></span><br>Law AIs Aren’t Really That Good Yet</h2><p class="paragraph" style="text-align:left;">Harvey has released some early results of its <a class="link" href="https://www.harvey.ai/blog/legal-agent-benchmark-initial-results?utm_source=chatgpt.com" target="_blank" rel="noopener noreferrer nofollow">Legal Agent Benchmark (LAB)</a>, which show that frontier AI agents still complete fewer than <b>10%</b> of complex legal tasks end-to-end under LAB’s strict “all-pass” grading standard.</p><p class="paragraph" style="text-align:left;">Interestingly, there’s a clear winner in this category: <b>Claude Opus 4.7 led the tested models at 7.1%</b>, followed by Sonnet 4.6 at 5.4%, Opus 4.6 at 4.2%, <b>GPT-5.5 at 2.1%</b>, <b>and Gemini 3.5 Flash at 0.8%,</b> with the latter two showing a really poor performance.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">It’s important that this task and domain-specific benchmarks start to emerge to show how hard reality can hit. <b>These models can do an okay-ish job, but fail miserably if you don’t hold their hand, at least today</b>.</p><p class="paragraph" style="text-align:left;">For now, they remain an interesting coworking tool that, used well, can really push what one can do. The other side of the coin is costs. <b>Opus 4.7</b>, the top scorer, <b>cost about $50.90 per task and took roughly 22 minutes per run</b>, while faster or cheaper models scored lower.</p><p class="paragraph" style="text-align:left;">This doesn’t sound cheap at all, and continued use of these models can be a nightmarish expense.</p><p class="paragraph" style="text-align:left;">And talking about Opus…</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MODELS</b></span><br>Anthropic’s Opus 4.8 Out. Mythos Soon?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8580a016-b9ce-4213-8ec3-57ad2443a1f6/image.png?t=1779995544"/></div><p class="paragraph" style="text-align:left;">Minutes ago, as I was writing these words, <a class="link" href="https://www.anthropic.com/news/claude-opus-4-8?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">Anthropic launched Claude Opus 4.8</a>, an upgraded version of its flagship AI model, while also preparing to release a more advanced model, the long-awaited <b>Claude Mythos</b>.</p><p class="paragraph" style="text-align:left;">Anthropic says Opus 4.8 improves on earlier Opus models in areas such as coding, reasoning, financial analysis, computer use, and browser-agent tasks. <b>The company highlights “honesty” as a key change</b>, saying the model is more likely to flag uncertainty and avoid unsupported claims.</p><p class="paragraph" style="text-align:left;">Unsurprisingly, <b>Anthropic is offering Opus 4.8</b><b> at the same price as its predecessor</b>, despite performance improvements. The Verge also reports that Anthropic is adding controls for how much “effort” Claude applies to tasks, affecting token use and cost.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Minutes later, they announced a monstrous $60 billion Series H round, <a class="link" href="https://x.com/AnthropicAI/status/2060061347522433422?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=china-sends-a-message-opus-4-8-more" target="_blank" rel="noopener noreferrer nofollow">valuing the company at a whopping $963 billion post-money</a>. As I explained in a recent post, the discussions were about a new $30 billion round,<b> but it has turned out to be twice that size.</b></p><p class="paragraph" style="text-align:left;">As for Opus 4.8, it seems like it&#39;s fully state-of-the-art, but come on, at this point we&#39;re both better than this and know that benchmarks rarely tell the same story. For instance, Opus 4.7 looks better on benchmarks than GPT-5.5, <b>and that couldn’t be further from the truth in practice</b>.</p><p class="paragraph" style="text-align:left;">The next days will tell us how much better Opus 4.8 really is. But one thing’s for sure: <b>it’s not going to be cheaper</b>.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">AI is truly hitting a hard wall when it comes to costs, which have risen to a point that people simply can’t ignore them. I believe Uber is the first of many companies to realize that either you become sophisticated and frugal in your use of AI, or your AI bills will become a problem.</p><p class="paragraph" style="text-align:left;">And the fact that Anthropic has released Opus 4.8 without a price reduction, despite the heat they are getting, tells me they just can’t.</p><p class="paragraph" style="text-align:left;">Yes, <b>they now command a $47 billion annual run rate</b>, but if people truly start believing AI is expensive, these companies’ revenue growth frenzy may come to a halt. And right quick.</p><p class="paragraph" style="text-align:left;">Remember, as I always say, t<b>he marginal cost of software is no longer zero once you include AI</b>, which means every user counts.</p><p class="paragraph" style="text-align:left;">Beyond the tech, we continue to see basic market degeneracy to the point where people are borrowing against their own shares to buy more. <b>This is the type of behavior that, if things go south, we will all regret.</b></p><p class="paragraph" style="text-align:left;">The silver lining is that AI continues to make progress in maths, but we can’t pretend to be eternally excited when all good news comes from verifiable domains like coding or maths.</p><p class="paragraph" style="text-align:left;"><b>The world is much more than coding and maths; we need results elsewhere</b>, too, like in law. And the truth is that, once you evaluate AI in those domains, the news aren’t that exciting anymore.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=6781de54-cd08-48cb-b7d0-73f544038b9d&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>The New Power Frontier: The 800 VDC Revolution</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dac5b0d6-d616-4f4a-b22d-cd387c20f755/image.png" length="2827262" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/the-new-power-frontier-the-800-vdc-revolution</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/the-new-power-frontier-the-800-vdc-revolution</guid>
  <pubDate>Mon, 25 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-25T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>The Next Power Frontier</h2><p class="paragraph" style="text-align:left;">Welcome back! Today is one of those newsletters in which I’ve learned just as much as what you’re going to learn yourself. And it also has a huge upside because it’s an article that will set a precedent for you about data centers for the foreseeable future.</p><p class="paragraph" style="text-align:left;">Because to allow progress to continue, data centers will undergo a profound repurposing, especially regarding power equipment, the great forgotten bottleneck.</p><p class="paragraph" style="text-align:left;"><i>The deadline?</i> <b>As soon as next year.</b></p><p class="paragraph" style="text-align:left;">This implies a huge amount of player and opportunity reshuffling, which means the lamest and most predictable part of the entire AI stack has just become exciting.</p><p class="paragraph" style="text-align:left;">I’m going to tackle everything to describe where data centers are headed, the challenges that lie ahead, and the financial opportunities that emerge in the process, <b>including the hottest new market emerging from AI</b> and the company that, as of today, represents my next investment; a company that is equally risky as it has huge upside potential, probably the most out of all the ones I’ve talked about.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">Bring me the heat</h2><p class="paragraph" style="text-align:left;">As you surely know, we use accelerators (e.g., GPUs) to train and serve AI models. But to get those GPUs to do their job, we need a lot more stuff around them: memory, storage, CPUs, interconnect gear, and, of course, all the power-related equipment: VRMs, PSUs, busbars, sidecars, breakers, switchgear, converters, rectifiers, transformers, and many more.</p><p class="paragraph" style="text-align:left;">Building a data center is not like building a warehouse; everything is connected. But until now, we were in easy mode, <b>as AI has followed the traditional approach to building data centers</b>.</p><p class="paragraph" style="text-align:left;">But for the beasts that are arriving as soon as next year, <b>current data centers are not enough, not even close</b>; we need something else.</p><p class="paragraph" style="text-align:left;"><i>But why? </i>Well, absurd amounts of power in very small places.</p><h3 class="heading" style="text-align:left;">Power is rising fast</h3><p class="paragraph" style="text-align:left;">Just five years ago, the state of the art in accelerators was the NVIDIA A100 (the chip that, by the way, gave us the first ChatGPT).</p><p class="paragraph" style="text-align:left;">These chips were located in servers in groups of 8, <b>and each server required 6,500 watts at max power</b>, less than what my burger joint’s fryer needs to operate.</p><p class="paragraph" style="text-align:left;">This meant that you could house these chips in traditional data centers, which could allow between 5-10kW per rack (a rack is the chassis where the AI server, or other non-AI servers, are housed.</p><p class="paragraph" style="text-align:left;">A data center offering 15-20 kW per rack was considered pretty dense at the time, before the AI revolution. However, <b>things have changed</b>.</p><p class="paragraph" style="text-align:left;">Today, an NVIDIA GB300 NVL72, the most powerful AI server the world has to offer, <b>can “offer” peaks of 150 kW</b>, more than a tenth of a megawatt.</p><p class="paragraph" style="text-align:left;">This has already implied having to build data centers from scratch just to be able to house servers of this size, <b>and led to truly enormous data centers that can draw hundreds of megawatts and, in some cases, gigawatts</b>, in order to house just a few thousand of these servers.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But wait, it gets worse.</p><p class="paragraph" style="text-align:left;">With the 2027 server lineup, AI servers like NVIDIA’s Rubin Ultra or AMD’s MI550 series, we cross 500 kW per rack, reaching 600 kW per single AI server, <b>with a clear line of sight to</b><b> 1 MW servers before the end of the decade</b> (Interestingly, China already crossed that line with the Huawei CloudMatrix, a 500kW beast. However, that is a corridor-scale server, not a single rack).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/118f76f4-2376-4b26-b194-d1e47b9b4f41/image.png?t=1779695429"/><div class="image__source"><span class="image__source_text"><p>Be aware that the x-axis is compressed. Source: Citrini</p></span></div></div><p class="paragraph" style="text-align:left;">But before I explain to you why this completely changes the picture of how you approach building a data center, we must first answer: <i>why, in the name of Jesus, do</i><i> we need that much power?</i></p><h3 class="heading" style="text-align:left;">The New Scary Accelerator</h3><p class="paragraph" style="text-align:left;">In AI hardware land, you have two ways of “progressing”:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Packing more compute into each accelerator, which has been the primary driver of progress for decades</p></li><li><p class="paragraph" style="text-align:left;">Packing more memory to offer higher read/write speeds</p></li></ol><p class="paragraph" style="text-align:left;">Not trying to linger too much on this topic, <b>you have to think about accelerators as two workers in a factory</b>. The memory worker feeds work to the compute worker, which performs the job and returns the result. It’s a back-and-forth between both.</p><p class="paragraph" style="text-align:left;">Usually, you optimize one or the other, but in AI, you need both because, no matter how fancy Anthropic or OpenAI sound, we only really know how to progress by packing more compute and data. That’s the only play in the book.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>First scaling law</b>: Models are getting bigger, increasing both memory requirements and compute requirements (every single prediction needs more “work”)</p></li><li><p class="paragraph" style="text-align:left;"><b>Second scaling law</b>: We also need longer sequences (e.g., models need to remember more in context and execute more tools, which is crucial for agents). This increases memory requirements in particular.</p></li></ol><p class="paragraph" style="text-align:left;"><b>This means that each accelerator has grown significantly in size along both axes, compute and memory</b>. For instance, if we compare the current NVIDIA chip in production, the B200, to the 2027-slated Rubin Ultra chip, it’s literally half the size (approximately 2.5-ish more compute and 4-ish the memory required).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/dc7d8fad-0015-4433-8691-226b4381d861/image.png?t=1779696057"/></div><p class="paragraph" style="text-align:left;">But what’s more striking is that the power required has more than doubled between these two chips (orange and dark blue below):</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6918e17a-49ff-4764-8bd0-8c7bd164001f/image.png?t=1779696109"/></div><p class="paragraph" style="text-align:left;">2,500W doesn’t sound like a lot considering your hairdryer requires that or more. The problem is that size is still literally the surface area of a little more than a credit card, generating incredible amounts of power… and heat, in a very concentrated place.</p><p class="paragraph" style="text-align:left;">And when you factor in that transistors operate at less than 1 volt (~0.7 V), using Watt&#39;s law, <b>we have 3,571 Amps of current flowing through this credit card-sized chip</b>. </p><p class="paragraph" style="text-align:left;">Not remotely comparable, but for reference, you only need 0.1 A, or 35,000 times less, and a stroke of bad luck, to kill a human, and your (~2,300 W) hairdryer requires around 10 Amperes to work in Europe (~20 in the US), only 3,500 times less than the numbers we’re throwing around here.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And once we start talking about logic-to-logic hybrid bonding, <a class="link" href="https://www.huaweicentral.com/huawei-logicfolding-architecture-everything-you-need-to-know/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-new-power-frontier-the-800-vdc-revolution" target="_blank" rel="noopener noreferrer nofollow">as Huawei intends to do</a>, we can make double use of our GPUs as ovens, too. But I digress.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But having servers that require such large amounts of power and, indirectly, current means one thing: <b>losses</b>.</p><h3 class="heading" style="text-align:left;">The traditional way</h3><p class="paragraph" style="text-align:left;">The power a data center needs has to come from a source. That source can be on-site, as most Hyperscalers would love to do, in order to avoid the grid altogether and avoid dealing with angry neighbors with higher electricity bills.</p><p class="paragraph" style="text-align:left;">In reality, however, <b>most of the data centers remain connected to the grid</b>, and many new projects will still require grid connection:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1cc62721-c946-41b3-be58-b67d7e4c1108/image.png?t=1779697112"/><div class="image__source"><span class="image__source_text"><p>Source: Sightline</p></span></div></div><p class="paragraph" style="text-align:left;">This means data centers receive from an external source a huge amount of power at very high voltage, using alternating current, or AC.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><i>But why AC? What is AC?</i></p><p class="paragraph" style="text-align:left;">Most power plants are about heating something that vaporizes water. This water gets funneled through a turbine, which starts to move. At the end of the turbine, we have a rotor with magnets that, by default, create a magnetic field. Around those magnets, we have static wire coils.</p><p class="paragraph" style="text-align:left;">As the turbine rotates, so do the magnets, <b>creating a changing magnetic field around the coils</b>. Using Faraday’s Law, the changing magnetic field induces a voltage in the wires, creating current, which is then sent to the world.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3d9bc430-6528-499d-9fb5-deafb33156fe/image.png?t=1779698046"/></div><p class="paragraph" style="text-align:left;">Interestingly, because each coil sometimes sees the magnet&#39;s north pole and sometimes its south, the current “changes direction” continuously, creating the alternating flow, giving us ‘alternating current’.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">As mentioned, to minimize transmission losses, electricity is transmitted at very high voltages, <b>so that the current is as small as possible</b>. We’re talking about tens of thousands of volts, thousands of times more than what GPUs need.</p><p class="paragraph" style="text-align:left;">Thus, <b>we need to step down those voltages</b>. For that, <b>we use a transformer</b>, a piece of equipment that hasn’t changed much in more than 100 years (until now, as we’ll see later).</p><p class="paragraph" style="text-align:left;">This machine puts two coiled wires side by side, the second shorter than the first, and we apply the same intuition as in the generator; as the current changes direction, the change in current direction changes the magnetic field, which induces a current in the other wire (the ‘iron core’ helps keep the magnetic field constrained rather than spreading out).</p><p class="paragraph" style="text-align:left;">In a step-down transformer, the second coil is shorter, so the induced voltage is smaller.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/ded874fa-976a-47ba-9016-9a841f7e4108/image.png?t=1779699086"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://www.facebook.com/mdrashidk2/posts/understanding-the-internal-diagram-of-a-transformera-transformer-is-an-essential/24040104622347384/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-new-power-frontier-the-800-vdc-revolution" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">With this transformer (many of them in a sequence, actually), we gradually step down the voltage to around <b>54 V</b> (you usually have other transformers in between), <b>which is the legacy voltage we work with in normal data centers</b>. Overall, the process looks a lot like this:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/3cc309cb-0fa6-4b8b-b508-ca705cc39101/image.png?t=1779700134"/></div><p class="paragraph" style="text-align:left;">We don’t have to go deep with the intermediate steps like UPS/Battery and switchgear (the former is to ensure chips keep receiving electrons if the grid fails, the latter is about protection and isolation).</p><p class="paragraph" style="text-align:left;">The important part here is steps 5 and 6.</p><ol start="1"><li><p class="paragraph" style="text-align:left;">We first use a rectifier to transform AC to DC (chips need Direct Current)</p></li><li><p class="paragraph" style="text-align:left;">We also do a step down from 54 to 12 V and then to 0.7 V, both internally in the rack using VRMs (Voltage Regulator Modules) in power shelves</p></li></ol><p class="paragraph" style="text-align:left;">Currently, <b>that power equipment is inside the rack NVIDIA ships to customers</b>, meaning the rack receives AC at a certain voltage, which is rectified (converted to DC) and stepped down once or even twice.</p><p class="paragraph" style="text-align:left;">This power equipment that does this requires space in the rack, which means less space for compute.</p><p class="paragraph" style="text-align:left;">For example, the Blackwell GB300 NVL72, NVIDIA’s most powerful server, and the one being deployed as we speak by your favorite data center constructor, <a class="link" href="https://developer.nvidia.com/blog/nvidia-800-v-hvdc-architecture-will-power-the-next-generation-of-ai-factories/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-new-power-frontier-the-800-vdc-revolution" target="_blank" rel="noopener noreferrer nofollow">requires 8 power shelves to handle the steps outlined above</a>.</p><p class="paragraph" style="text-align:left;"><i>The issue?</i> <b>We’re basically reaching the limit</b>. If we wanted to do the same with the 2027 Kyber Rack (the Rubin VR200 NVL576, also known as ‘Rubin Ultra’), you would need the entire server for power, literally,<b> leaving absolutely zero space for compute</b>.</p><p class="paragraph" style="text-align:left;">We need an alternative. And we need it fast.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-new-power-frontier-the-800-vdc-revolution">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=the-new-power-frontier-the-800-vdc-revolution">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=273c44e5-3376-4b57-ab74-f22271e05d1e&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>A Historical Discovery &amp; a Bunch of Lies</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f38c56f9-8e65-4789-8486-a5b66c32f991/image.png" length="197492" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/a-historical-discovery-a-bunch-of-lies</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/a-historical-discovery-a-bunch-of-lies</guid>
  <pubDate>Fri, 22 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-22T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><b>Welcome back!</b> This week, we have the first truly AI-only discovery in history, <b>and it’s one of the greatest mysteries in maths</b>.</p><p class="paragraph" style="text-align:left;">We also discuss a new model trained at a millionth of the cost of frontier models, <b>which is still very competitive with models tens of times larger</b>, as well as a rundown of market news, including a review of <b>Anthropic</b>&#39;s claim that it’s about to become operationally profitable.</p><p class="paragraph" style="text-align:left;">We’ll end with <b>CEO digital twins</b> and Google’s I/O event.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Enjoy!</b></span></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>FRONTIER RESEARCH</b></span><br>A First in AI Discovery</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">As published by OpenAI</a>, the company says one of its<b> internal reasoning models</b> has disproved a longstanding conjecture tied to the “planar unit distance problem,” a famous question in discrete geometry first posed by Paul Erdős in 1946.</p><p class="paragraph" style="text-align:left;">The problem asks how many pairs of points in a plane can be exactly one unit apart. For decades, mathematicians believed the best constructions were essentially based on square-grid arrangements. <b>OpenAI says its model has found a new infinite family of constructions that outperform those assumptions</b>, yielding a polynomial improvement over the previously accepted approach.</p><p class="paragraph" style="text-align:left;">According to OpenAI, the proof was reviewed by external mathematicians. Princeton mathematician Noga Alon described the problem as one of Erdős’s favorite open questions, while researchers including Arul Shankar and Jacob Tsimerman said the construction was unexpectedly sophisticated and difficult to derive manually.</p><p class="paragraph" style="text-align:left;">Timothy Gowers, Fields Medal recipient, <a class="link" href="https://x.com/OpenAI/status/2057176201782075690?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">called it a huge milestone in AI’s push toward mathematics</a>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is actually a first in the industry, and I’m not exaggerating.</p><p class="paragraph" style="text-align:left;">For the longest time (very likely until this event), <b>AI has shown zero capability to improve the frontier of human knowledge autonomously</b>. Previous Erdős solutions in which AI participated included the active involvement of renowned mathematicians. Instead, here the AI was just given the problem and the goal, as you can see below:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f4b73e85-e171-47a9-908f-94f90a4d9476/image.png?t=1779443137"/><div class="image__source"><span class="image__source_text"><p><a class="link" href="https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-proof.pdf?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">Source</a></p></span></div></div><p class="paragraph" style="text-align:left;">Interestingly, <b>OpenAI claims this model is not a specialized model for maths</b> but a general-purpose model that happened to be incredibly good at maths, a claim very similar to the one made by Anthropic regarding Mythos&#39;s cybersecurity capabilities.</p><p class="paragraph" style="text-align:left;">The only caveat to all of this is that, while OpenAI published the reasoning (125 pages long), it was a redacted version, summarized from the original. Why OpenAI chooses not to disclose the actual reasoning trace is unknown, <b>but it certainly reduces trust in its claims</b>.</p><p class="paragraph" style="text-align:left;">That said, it’s important that we reflect on why this happened in the first place. As the researchers explain, the key here is exploration. Intelligence is commonly described as a mixture of intuition and search: our minds suggest possible directions, and we explore them to find the best solution.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Interestingly, <a class="link" href="https://arxiv.org/abs/2605.20579?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">three hours later, the proof had already been improved by a human</a>. From what I’ve been told, this is common with proofs, but I guess we can say a little bit of our faith in humanity’s utility has been restored.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>A model trained for a millionth of the cost?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://sapient.inc/introducing-hrm-text/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">Sapient Intelligence has introduced HRM-Text</a>, a fully open 1.15-billion-parameter reasoning language model designed to achieve strong benchmark results with far less training data and compute than conventional models. </p><p class="paragraph" style="text-align:left;">Sapient says HRM-Text was trained on about <b>40 billion tokens</b> of structured data, compared with the <b>4–36 trillion tokens</b> used in many large open pretraining runs, thousands of times fewer tokens.</p><p class="paragraph" style="text-align:left;">The company claims the model can be pretrained in roughly one day for about <b>$1,000</b>, with an int4 size of <b>0.6 GB, meaning it can easily run even on your smartphone</b>.</p><p class="paragraph" style="text-align:left;">The reported benchmark results include <b>56.2% on MATH</b>, <b>82.2% on DROP</b>, <b>81.9% on ARC-Challenge,</b> and <b>60.7% on MMLU</b>, benchmarks that are in no way state-of-the-art these days, but were three years ago.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Sapient says these results were independently verified in April 2026 and reflect only the base model, without post-training, fine-tuning, or reinforcement learning.</p><p class="paragraph" style="text-align:left;"><i>But how on Earth did they pull this off?</i></p><p class="paragraph" style="text-align:left;">For starters, <b>average data quality in the pretraining set is much higher</b>. That is, instead of just feeding the entire world’s data, they select high-quality data and train just on that.</p><p class="paragraph" style="text-align:left;">Beyond that, the model is a <b>hierarchical latent recurrent architecture</b>. <i>But what is that?</i></p><p class="paragraph" style="text-align:left;">Large Language Models like ChatGPT have a very simple functioning mechanism. They receive a set of words, and they produce the next one. This one gets appended, and the process repeats for the next.</p><p class="paragraph" style="text-align:left;">Here, when measuring how an AI deploys compute, <b>the amount of internal computation is secondary to the amount of tokens it generates</b>. In other words, LLMs mostly leverage more compute by generating more tokens.</p><p class="paragraph" style="text-align:left;">The HRM models work slightly differently because they include a recursive process inside before answering. In other words, just like you may think for longer internally before you answer a question, <b>HRMs can do the same (kind of; it’s a loose analogy)</b>. The point here is that just like humans and reiteratively reflect on their own thoughts, HRMs can loop over theirs.</p><p class="paragraph" style="text-align:left;">This ‘recurrence’ is what allows them to generate more compute despite being way smaller. But we both agree this is something not worth using today, relative to the accessible models.</p><p class="paragraph" style="text-align:left;">So, <i>what should we take from this?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The key to all of this is the message it sends. In the current paradigm, you first build knowledge, and then you build reasoning (i.e., “intelligence”). That is, first you teach the model everything about the world, which gives it intuition, and then you train it to use that knowledge to reason.</p><p class="paragraph" style="text-align:left;">The problem here is two-fold:</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>It’s incredibly expensive</b>, requiring you to train extremely large models on extremely large datasets to build sufficient knowledge.</p></li><li><p class="paragraph" style="text-align:left;"><b>It incentivizes memorization</b>. For many responses, we don’t really know whether the model has memorized the response or is actually reasoning it.</p></li></ol><p class="paragraph" style="text-align:left;">But the fact that this model worked suggests an alternative option. Despite being extremely small relative to its peers, <b>it matches their “intelligence” capabilities while lagging in the knowledge benchmarks </b>(because it wasn’t trained on as much data).</p><p class="paragraph" style="text-align:left;">The authors theorize that <b>we might be able to decouple knowledge from reasoning</b>, meaning we could train incredibly powerful, small reasoners that rely on recursion to add extra compute.</p><p class="paragraph" style="text-align:left;">They envision a future, first envisioned by Anthropic’s new star hire, Andrej Karpathy, <a class="link" href="https://thewhitebox.beehiiv.com/p/anti-hype-guide-part-2-karpathy?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">and defined as the reasoning core</a>, in which models are small, don’t know too much about anything, but are great reasoners.</p><p class="paragraph" style="text-align:left;"><i>And if they need knowledge?</i> Well, they just look it up with a search API.</p><p class="paragraph" style="text-align:left;">And while we don’t know for sure if this is a viable route, we do know that sample efficiency, getting models to be smart without requiring “infinite” data, <b>has a lot of room for improvement</b>.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ECONOMICS</b></span><br>Microsoft Starts Canceling Claude Code Licenses</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.theverge.com/tech/930447/microsoft-claude-code-discontinued-notepad?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">As published by The Verge</a>, Microsoft is reportedly canceling most internal <b>Claude Code</b> licenses and moving employees to <b>GitHub Copilot CLI</b> by <b>June 30, 2026</b>.</p><p class="paragraph" style="text-align:left;"><a class="link" href="https://the-decoder.com/microsoft-pulls-claude-code-licenses-and-pushes-developers-back-toward-its-own-ai-tool/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">A cost motive seems central to what is going on</a>; the June 30 deadline coincides with the end of Microsoft’s fiscal year, and multiple reports say cost control was likely a major factor behind the move. <b>Claude Code’s token-based usage costs reportedly rose as adoption expanded internally</b>, making the tool expensive at enterprise scale.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Well, well, well, <i>who would’ve thought?</i> Well, I expect you did, <b>because I’ve been telling you this technology is very expensive for a while now.</b></p><p class="paragraph" style="text-align:left;">Given the amazing growth we’ve seen so far this year from Anthropic and OpenAI, enthusiasts in the space have just assumed AI has fixed its problems and that the moon is next.</p><p class="paragraph" style="text-align:left;">Well, no.</p><p class="paragraph" style="text-align:left;">For instance, in the event I had two days ago with executives in Spain, which included top leadership from Microsoft, BBVA, Telefonica, and others,<b> many </b><b>were completely unaware of the real costs of AI</b> and are only now starting to realize that this technology can get expensive.</p><p class="paragraph" style="text-align:left;">The “illusion” is caused by subscriptions, which hide AI’s real cost. Once these subscriptions end (soon), and companies have to pay for AI’s real cost, oh boy.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>HARDWARE</b></span><br>Anthropic’s $45B deal with SpaceX</h2><p class="paragraph" style="text-align:left;">One of the best things about AI companies IPOing this year is that they are finally going to stop gaslighting us with the state of their financials.</p><p class="paragraph" style="text-align:left;">As SpaceX has filed its S-1, <b>we now better understand the real costs of AI, no matter what Anthropic tells you</b> (a money-loser for now; we’ll talk about this below).</p><p class="paragraph" style="text-align:left;">As explained by Bloomberg based on this filing, <a class="link" href="https://www.bloomberg.com/news/articles/2026-05-20/anthropic-to-pay-spacex-nearly-45-billion-for-computing-deal?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc3OTM0MDE0MywiZXhwIjoxNzc5OTQ0OTQzLCJhcnRpY2xlSWQiOiJURkNUNldLSUpISUkwMCIsImJjb25uZWN0SWQiOiJBOEExRDhFQTI5OTc0OTRGQTQ1QUE2REJBMjAwNTM3MSJ9.7GmTLgNTHuRQhQxgg38WqTvHCpZXe6DAGkd4qH3ckIA&utm_source=tldrai&leadSource=uverify%20wall" target="_blank" rel="noopener noreferrer nofollow">Anthropic has agreed to pay $1.25 billion a month</a>, or $15 billion a year, for the next three years, with May and June being paid with a heavy discount as the hardware ramps up (we’ll see why this is particularly important later).</p><p class="paragraph" style="text-align:left;">This is a huge bump to SpaceX’s revenues in its effort to justify an incredible $2 trillion market cap for a company with fewer revenues than Anthropic.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;"><i>But how much compute is SpaceX giving Anthropic for such a huge amount of money?</i> We don’t know, but we can work something out. </p><p class="paragraph" style="text-align:left;">Assuming the deal was signed at, say, <b>$5/hour per GPU</b>, a slightly discounted price from the standard price for top-end GPU prices on spot markets, that’s $43,800/year per GPU. That gives roughly 342,466k GPUs.</p><p class="paragraph" style="text-align:left;">The Colossus datacenters are a mixture between H100s, H200s, and B200s, but let’s simplify by assuming they are all B200s, 72 per server, so we have 4,756 B200 server-equivalents.</p><p class="paragraph" style="text-align:left;">Each GB200 NVL72 server draws 120 kW, giving around 570 MW, or more than half a GW of compute, approximately half of what both data centers will eventually dispose of, assuming the rest is given to the Cursor team, <b>which has been basically unofficially acquired by SpaceX</b>.</p><p class="paragraph" style="text-align:left;">This event basically marks the undeniable pivot of SpaceX-xAI toward becoming more of a neocloud than a model-layer competitor. Unless Grok 5 turns out to be other-worldly, <b>it seems we’ve lost the first competitor in the frontier AI race</b>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>MARKETS</b></span><br>Anthropic, First Profitable AI Lab?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f38c56f9-8e65-4789-8486-a5b66c32f991/image.png?t=1779443917"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-propel-anthropic-into-its-first-profitable-quarter-7edbf2f4?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-propel-anthropic-into-its-first-profitable-quarter-7edbf2f4?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">reported by the Wall Street Journal</a>, Anthropic expects to report its first quarterly <b>operating profit</b> in Q2 2026, driven by projected revenue of <b>$10.9 billion</b>, up from <b>$4.8 billion</b> in Q1. The company projects <b>$559 million in operating profit</b> for the June quarter.</p><p class="paragraph" style="text-align:left;">The growth is tied to rising enterprise demand for Claude and AI coding tools. <b>The figures are still projections, not final results</b>, and the WSJ notes that comparisons with rivals are complicated by differing accounting practices. Anthropic’s heavy compute spending also means profitability may not continue across the full year.</p><p class="paragraph" style="text-align:left;"><i>But how should we take these numbers?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">It’s estimated that there are about 50 quintillion tons of salt on Earth. <b>That’s the size of the pinch of salt I want you to take from this</b>.</p><p class="paragraph" style="text-align:left;">This is just absurd number-cooking to portray the company in the best light possible ahead of the IPO. Even the WSJ had to add, <i>‘It is unclear what accounting methods Anthropic has used to book revenue and costs’.</i></p><p class="paragraph" style="text-align:left;">Of course, it’s unclear, buddy, it’s totally made up. For starters, it’s based on a projection of Q2 revenues. Second, there’s no way that includes all training costs; they surely just included the training costs for the models they managed to monetize (they very likely excluded Mythos training costs, for instance).</p><p class="paragraph" style="text-align:left;">The revenue growth is truly incredible, yes, but there are many things we have to be cautious about here. First, as legendary investor Charlie Munger once said, <i><b>“EBITDA earnings are bullshit earnings</b></i>.”</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">And the earnings Anthropic is presenting here aren’t even GAAP-compliant numbers; <b>they are </b>✨<b>adjusted</b>✨<b>, whatever the hell that means, because they don’t elaborate.</b></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This is literally saying that if we take this, this, and this, and discard this, this, and that, ✨we are profitable✨. Well, <b>that’s not how profitability works</b>.</p><p class="paragraph" style="text-align:left;">And let me be clear; they aren’t saying they are profitable but “operationally profitable”, meaning the core business makes money, but that’s precisely the point I’m trying to make; <b>there’s no such thing as “operationally profitable” in a capital-intensive business.</b></p><p class="paragraph" style="text-align:left;">Needless to say, SpaceX is hemorrhaging money, according to the S-1 filing, and OpenAI is allegedly bleeding money on another level entirely, <b>which tells me some companies are being honest and others aren’t</b>.</p><p class="paragraph" style="text-align:left;">And before you accuse me of being biased, let me clarify one thing: there are businesses where showing adjusted earnings can make sense. Sometimes, your business has a one-off, huge cost that should not recur, which skews your GAAP-compliant numbers.</p><p class="paragraph" style="text-align:left;">In that case, as it shouldn’t repeat, <b>it can make sense to clarify this</b>, in the same way an unexpected healthcare payment in April that turned the entire month into a loss doesn’t mean you’re irresponsible with your own spending.</p><p class="paragraph" style="text-align:left;"><b>But ‘adjusted earnings’ are not applicable to a capex-intensive business</b>; the huge capex reinvesting process never ends; the fact that your hardware reinvestments turn all this to red must be mentioned if it’s the norm; <b>it’s not a one-off thing; you’re spending more than you earn on a regular basis</b>.</p><p class="paragraph" style="text-align:left;">Instead, AI companies with huge investments should be measured by cash flow. I don’t care if you’re profitable if you aren’t capable of offsetting your own investments.</p><p class="paragraph" style="text-align:left;">And let’s not forget that Anthropic&#39;s numbers excluded stock-based compensation, which is historical, given that a large portion of their employees’ salaries is made up of equity.</p><p class="paragraph" style="text-align:left;">The main point I’m trying to make, and this applies not only to Anthropic but to every single AI Lab, is that we should always judge them by cash flow, how much money is coming in and out of them, not adjusted numbers or operational profits. Capex-intensive companies can’t be judged by how profitable your operations are <b>if you need to invest two times that to run the business in the first place</b>.</p><p class="paragraph" style="text-align:left;">Anthropic won’t actually be profitable in the years ahead; let’s stop pretending they will.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>GEOPOLITICS</b></span><br>Europe’s nightmare in a photo</h2><p class="paragraph" style="text-align:left;">What appears massive and exciting is actually terribly revealing. <a class="link" href="https://www.bloomberg.com/news/articles/2026-05-20/french-companies-bid-for-10-billion-europe-ai-gigafactory-site?taid=6a0da8212bdfad00011cfb24&utm_campaign=trueanthem&utm_content=business&utm_medium=social&utm_source=twitter&embedded-checkout=true" target="_blank" rel="noopener noreferrer nofollow">As </a><a class="link" href="https://www.bloomberg.com/news/articles/2026-05-20/french-companies-bid-for-10-billion-europe-ai-gigafactory-site?taid=6a0da8212bdfad00011cfb24&utm_campaign=trueanthem&utm_content=business&utm_medium=social&utm_source=twitter&embedded-checkout=true" target="_blank" rel="noopener noreferrer nofollow">reported by Bloomberg</a>, a French consortium is bidding to host one of the EU’s planned AI gigafactory sites, with a project valued at about <b>€10 billion </b>(<b>$10 billion-plus)</b>.</p><p class="paragraph" style="text-align:left;">The bid is led by <b>AION</b>, a group involving major French technology, telecom, finance, and energy players, including <b>Orange, Iliad/Scaleway, Capgemini, Artifact, Bull, Ardian, and EDF</b>, according to Reuters. The proposed site would be built in <b>France</b> and is intended to provide large-scale computing capacity for AI model training and deployment.</p><p class="paragraph" style="text-align:left;">The project is tied to the EU’s broader effort to build sovereign AI infrastructure and reduce reliance on US and Chinese cloud and AI providers.</p><p class="paragraph" style="text-align:left;">The European Commission has been backing AI “factories” through EuroHPC, with total public and member-state investment in supercomputing and AI factory infrastructure expected to reach <b>€10 billion</b> over 2021–2027.</p><p class="paragraph" style="text-align:left;">AION’s plan reportedly targets an initial phase of around <b>100 megawatts</b> of capacity, with a longer-term goal of reaching <b>one gigawatt </b>(no way they are getting that without spending at least four times this amount).</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Celebrated as a massive achievement, <b>this is 5% of Google’s 2026 AI capital expenditures</b>, or 1.42% of what the five US Hyperscalers, Oracle, Meta, Amazon, Microsoft, and Google, are going to spend in the US this year.</p><p class="paragraph" style="text-align:left;">For reference, <b>this represents around 200 MW of compute at most</b>, enough to build a good inference data center, but nowhere near the required scale for frontier training, which has just moved into gigawatt-scale levels.</p><p class="paragraph" style="text-align:left;">The only thing going for Europe here is that, despite higher electricity costs, they don’t matter much, as ~90% of the TCO (Total Cost of Ownership) is capital costs. On the flip side, <b>the EU doesn’t remotely have the level of liquidity the US has</b>.</p><p class="paragraph" style="text-align:left;">And before you accuse me of being a Europhobe (ironic considering I’m European), you don’t have to believe me, as an even worse vibe was given by Arthur Mensch, Mistral’s CEO, who said that Europe <a class="link" href="https://www.businessinsider.com/mistral-ceo-warns-europe-2-years-avoid-us-ai-dependence-2026-5?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow"><i>“Had two years before becoming a US vassal state.”</i></a></p><p class="paragraph" style="text-align:left;">He also shared their plan to have around 1 gigawatt of compute capacity available by 2029 (five times the capacity announced above). For reference,<b> that is already less than what OpenAI or Anthropic have individually today</b>, and don’t get me started with Google or Meta.</p><p class="paragraph" style="text-align:left;">For instance, Google is estimated to have 4 million H100 equivalents. At 1.2kW per GPU (accounting server overhead, not just GPU power alone), <b>that’s around 4.8 GW</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/a688735a-91f9-4bf4-b45a-6ac6ebfa6bb0/image.png?t=1779445673"/></div><p class="paragraph" style="text-align:left;">Unless we find a way to produce intelligence with much fewer electrons, which many companies in France and the UK think is possible, though, and Sapient might have proven above, too, <b>Europe has pretty much lost this race already</b>.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>EVENTS</b></span><br>Google I/O, kind of a bummer</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://blog.google/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">Google I/O 2026, their national product event</a>, focused almost entirely on AI, with the most critical announcements centered on Search, Gemini, and Android XR.</p><p class="paragraph" style="text-align:left;"><b>Google said Search is receiving its biggest upgrade in more than 25 years</b>, adding a new AI-powered search box and agent-style features that can act on user requests inside Search.</p><p class="paragraph" style="text-align:left;">The company also introduced the <b>Gemini 3.5</b> model family, making <b>Gemini 3.5 Flash</b> the default model in the Gemini app and Search. Google described the update as part of a broader shift toward more “agentic” AI, where Gemini can perform longer, multi-step tasks.</p><ul><li><p class="paragraph" style="text-align:left;">Google also unveiled <b>Gemini Omni</b>, a new multimodal model line designed to work across text, images, audio, and video, and expanded AI tools for developers through Google AI Studio and its Antigravity platform. <a class="link" href="https://x.com/chrisfirst/status/2057612088651252062?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">This model does have some seriously impressive demos</a>.</p></li><li><p class="paragraph" style="text-align:left;">Another major announcement was <b>Gemini Spark</b>, an always-on AI agent designed to work across Google Workspace and third-party apps.</p></li><li><p class="paragraph" style="text-align:left;">Google also highlighted new AI shopping tools, including <b>Universal Cart</b>, which is meant to let users complete purchases across multiple merchants.</p></li><li><p class="paragraph" style="text-align:left;">On the hardware front, Google returned to smart glasses and extended reality. The company showed Android XR glasses, including models in partnership with Xreal, Warby Parker, and Gentle Monster, with launches expected later in 2026.</p></li></ul><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Overall, the event had three takeaways:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Boy, does Google suck at product</p></li><li><p class="paragraph" style="text-align:left;">AI is becoming more expensive</p></li><li><p class="paragraph" style="text-align:left;">Google Search is now AI Search</p></li></ol><p class="paragraph" style="text-align:left;">Few people in this world are more bullish on Google than me,<b> but we have to address the fact that Google’s go-to-market strategy sucks</b>.</p><p class="paragraph" style="text-align:left;">They are the most innovative company on the planet, second to none, a literal gold mine, but then the execution is ‘meh’. Google Cloud’s dashboard is famously complicated, and the new suite of AI products is chaotic, overly overlapping, and hard to follow.</p><p class="paragraph" style="text-align:left;">I’ve been saying this for years: <b>Google’s biggest threat is itself</b>. They have unlimited data, more computing power than anybody else, and are cash-rich. They should have won the race by now.</p><p class="paragraph" style="text-align:left;">And yet, here we are, trailing OpenAI and Anthropic.</p><p class="paragraph" style="text-align:left;">Moreover,<b> I will say it’s becoming clear that Google DeepMind is not “LLM-pilled”</b>; the omni model seems to indicate that, to them, multimodality is the way to AI, and that LLMs by themselves won’t get there. We’ll see how that turns out.</p><p class="paragraph" style="text-align:left;">As mentioned above, <b>Google has also made a decisive transition in its search business toward AI</b>, which has huge implications for several businesses around the world that made money from traffic. This is long overdue, but traffic seems to be collapsing as people simply trust whatever the AI overview says, <a class="link" href="https://x.com/aboutberlin/status/2057423342496293243?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">as shown below by a popular Berlin guide website for tourists</a>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/c0ad1308-e548-4632-962f-d6150e2abf07/image.png?t=1779446938"/></div><p class="paragraph" style="text-align:left;">But the biggest takeaway for me was the confirmation that even Google, the most compute-rich company, is being forced to raise prices.</p><p class="paragraph" style="text-align:left;">The new flash model, Gemini 3.5 Flash, is around 3 times more expensive than the previous version per token <b>and more expensive overall once you factor in longer sequence generation</b>.</p><p class="paragraph" style="text-align:left;">For all the bright minds Google has, nobody seems to understand that what people want is good intelligence at a lower cost, not an Einstein-level bankrupting force. Because guess what, most enterprise processes don’t need “Einstein intelligence” to work well.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>STUPIDITY</b></span><br>What in the…</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.wsj.com/tech/ai/execs-are-deploying-digital-twins-to-do-their-work-9547b375?mod=djemfoe&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-historical-discovery-a-bunch-of-lies" target="_blank" rel="noopener noreferrer nofollow">As published by The Wall Street Journal</a>, executives are increasingly using AI “digital twins” trained on their own speech, writing, and work history to handle parts of their workload. The systems can answer questions, prepare messages, deliver presentations, or appear in public-facing formats while imitating the executive’s style.</p><p class="paragraph" style="text-align:left;">One example cited is <b>Reid Hoffman</b>, LinkedIn co-founder and Greylock partner, whose “Reid AI” was trained on 22 years of his content. According to the report, <b>the AI twin has made more than 75 public appearances in multiple languages since 2024</b>, and Hoffman says it can cut his workload by up to 50% when used.</p><p class="paragraph" style="text-align:left;">Other executives are using similar tools inside companies to scale leadership communication, answer employee questions, and support performance or management tasks. The article frames the trend as part of the broader rise of workplace AI agents, in which software not only assists with single tasks but also serves as a proxy for specific people or roles.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I don’t know, man. <b>I don’t think this is what the world needs</b>. At this point, I’m generating extreme fatigue with respect to anything remotely sounding like AI-generated.</p><p class="paragraph" style="text-align:left;">If you’re going to send a fricking replica of you to talk to me, I’m going to hate you with a passion. It’s distateful.</p><p class="paragraph" style="text-align:left;"><i>Can we stop pretending that AI can substitute humans in human interactions?</i> It’s just incredibly naive to think this is how the world will look in the future. <i>And do we want that?</i></p><p class="paragraph" style="text-align:left;">AIs must stay on tool-land; <b>they must be used to improve our lives, not act on our name</b>. I can deal with an AI customer support chatbot if it solves my problem; that’s still a purely transactional, I-need-help type of interaction I don’t need a human for. But that’s just about the only example where I could accept talking to an AI instead of a human.</p><p class="paragraph" style="text-align:left;">And to be clear, I talk to AIs all the time whenever I discuss my own work, reflect on a new study, etcetera.</p><p class="paragraph" style="text-align:left;">But these are transactional conversations where I’m clearly using the AI as a tool. CEO digital twins may make sense in certain situations, but if they become the norm, <i>do I even need the CEO then?</i></p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">We ended our newsletter in the worst possible way, with stupid claims and ideas.</p><p class="paragraph" style="text-align:left;">But let’s try to stay positive, because OpenAI’s math discovery is truly a first in the industry and one that must be celebrated. One of my 2026 predictions was precisely discovery, and discovery we got.</p><p class="paragraph" style="text-align:left;">However, after proofreading the entire piece, I’ve felt uneasy; something doesn’t feel right. It feels like we’re running out of time.</p><p class="paragraph" style="text-align:left;">The many delusional claims about Anthropic’s finances, a few back-of-the-envelope math coming from a company desperate to improve its finances ahead of an IPO, <b>and yet even analysts I deeply respect for their critical thinking are celebrating as if Anthropic was actually profitable</b>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">To be clear, and I’ve shared this at my last event, <b>as long as debt markets are willing to finance Hyperscaler corporate debt at the same risk level</b> as if it were the fricking United States Government paying, they have no reason to stop spending.</p><p class="paragraph" style="text-align:left;">But I wouldn’t be surprised if the net return was negative.</p><p class="paragraph" style="text-align:left;">And it’s sad because we all know what AI is capable of, we all know it’s an amazing technology with a great future. But the incredible level of investment we&#39;re making without waiting for the technology to prove its value in real life could threaten the entire industry.</p><p class="paragraph" style="text-align:left;">The absence of critical thinking makes me believe many people are simply incapable at this point of thinking in first principles, <b>and that’s a terrible, terrible sign</b>.</p><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=3ca1334e-1a17-45f1-a8fd-e239d4087833&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>My Honest &amp; Sober Opinion on AI</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/084a5765-893e-40e5-bc8e-f5a902d9f3ea/ChatGPT_Image_May_17__2026__01_24_12_PM.png" length="1315587" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/my-honest-sober-opinion-on-ai</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/my-honest-sober-opinion-on-ai</guid>
  <pubDate>Sun, 17 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-17T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>My Honest & Sober Opinion on AI</h2><p class="paragraph" style="text-align:left;">Next week, I’m giving a keynote to executives from top companies here in Spain, and I’ve been tasked with giving a state of the union on AI, from markets to product.</p><p class="paragraph" style="text-align:left;">It’s designed to be a review that keeps you instantly up to date on the industry and all the tricks it tries to play on you. It’s a sober, <b>no-bullshit</b> analysis that will help you make better decisions along the way.</p><p class="paragraph" style="text-align:left;">If you ever wanted to ask me, <i>“What’s your honest opinion on the state of AI?”</i> <b>this is what I would answer.</b></p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">The Macro Picture in 2 Minutes</h2><p class="paragraph" style="text-align:left;">First, we are going to recap the state of the industry from a macro perspective:<b> how big it is and where money is coming from</b> through a series of graphs you can go over in 2 minutes.</p><p class="paragraph" style="text-align:left;">From the looks of it, <b>the level of investment we’re seeing in AI is unprecedented</b>, especially in terms of the steepness (how fast we are deploying so much capital).</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6d10c5f1-1506-47eb-87a3-c366358e7ec9/image.png?t=1779029115"/><div class="image__source"><span class="image__source_text"><p>Source: Dell’Oro, IEA</p></span></div></div><p class="paragraph" style="text-align:left;">However, once we factor in global GDP, we realize that the situation is not that impressive or unique in history:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/384b645f-9068-4f5b-8988-d1cd89bd6cf0/image.png?t=1779029340"/><div class="image__source"><span class="image__source_text"><p>Source: Dell’oro, IEA</p></span></div></div><p class="paragraph" style="text-align:left;">Irrespective of that, <b>the stock market has reacted in good measure</b>, rallying to new all-time highs at the time of writing, and catapulting the ‘winners’, companies exposed to AI in a positive way, <b>and “sepulting” the losers</b>, those that are negatively affected by AI.</p><p class="paragraph" style="text-align:left;">Today, the stock market is a market of AI winners and losers.</p><p class="paragraph" style="text-align:left;">Which is to say, in the eyes of the market, <b>your correlation</b>, positive or negative, relative to AI, <b>determines your fate</b>. Perhaps no better example of this is the following, where I’ve plotted the performance of the CPU and Memory chips (a weighted basket of the most popular stocks in each) relative to a SaaS index.</p><p class="paragraph" style="text-align:left;">And while the latter is down 15% year-to-date, memory stocks are up an average of 200% (again, weighted), meaning they’ve tripled in value, and CPU stocks have, on a weighted average basis, doubled.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/d4e7f001-0404-4bd4-a881-f92aa41680fd/image.png?t=1779029531"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">This has, of course, concentrated indices (and earnings) in many countries around AI (we have some extreme cases like Korea), but the US is an example of this too:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6b4d5aa5-8502-40d3-8e0c-73fd77da569b/image.png?t=1779029743"/></div><p class="paragraph" style="text-align:left;"><b>This has stoked fears of an imminent bubble explosion</b> (it’s a bubble, <a class="link" href="https://thewhitebox.beehiiv.com/p/a-final-answer-is-ai-really-a-bubble?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai" target="_blank" rel="noopener noreferrer nofollow">but the implications are harder to answer</a>), but if we compare it to other technology-driven bubbles, <b>it’s nowhere near the same levels of ‘bubbleness’</b>, at least relative to the “dot-com” bubble of the early 2000s:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/37661c90-007e-4c8a-b6c9-429644157d79/image.png?t=1779029665"/><div class="image__source"><span class="image__source_text"><p>Source: Pictet Asset Management</p></span></div></div><p class="paragraph" style="text-align:left;">Seeing all of this, the natural question is: <i>who’s paying?</i></p><p class="paragraph" style="text-align:left;">The default answer everyone resorts to is US and Chinese Hyperscalers, especially the former (Meta, Amazon, Microsoft, Google, and Oracle), with capital expenditures that are more than impressive and grow incredibly year over year.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b6bf9c8b-e628-4ad3-9cb3-fe0cbc35f10d/image.png?t=1779030484"/><div class="image__source"><span class="image__source_text"><p>Source: Morgan Stanley</p></span></div></div><p class="paragraph" style="text-align:left;">However, <b>these guys don’t have infinite money</b>, and it’s having an impact that&#39;s emptying reserves really fast.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/64db08a6-4b71-4750-a484-e41d0b71768d/image.png?t=1779030614"/><div class="image__source"><span class="image__source_text"><p>Source: Bloomberg</p></span></div></div><p class="paragraph" style="text-align:left;">Thus, if you dig deeper, you realize that the source is actually ‘sources’ across vastly different levels or associated risk; a mixture between Big Tech cash and debt, and private equity, credit, and VCs.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/1b14f27a-a25f-4fd1-893b-c9d5d2c5e300/image.png?t=1779030814"/></div><p class="paragraph" style="text-align:left;">The conversation most enthusiasts aren’t prepared for is that, above, <b>there’s almost a trillion dollars of money that has to come from high-risk yields</b> (~$800 billion) and dogshit securitizations, where we’re going to see some pretty large fuckups (~$150 billion).</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But some of you may be inclined to argue that revenues are finally piling in and financing much of this, <b>given</b><b> Anthropic’s impressive revenue growth</b>. But as you probably know, <b>they are trailing this massive investment by a considerable margin </b>(no more than $100 billion/year by year’s end). </p><p class="paragraph" style="text-align:left;">Yes, we’ve all heard the claims of massive AI revenues from the Hyperscalers, but every time I hear that argument, I just show them this:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/69c88dbe-801f-4a7d-a7d8-c63ee8d4fbee/image.png?t=1779030114"/><div class="image__source"><span class="image__source_text"><p>Source: The Wall Street Journal</p></span></div></div><p class="paragraph" style="text-align:left;">Guys, a lot of this money is “made up.” <i>But what is holding enterprises back from adopting this amazing technology yesterday?</i></p><p class="paragraph" style="text-align:left;">Oh boy, where do I start.</p><h2 class="heading" style="text-align:left;">State of the Technology</h2><p class="paragraph" style="text-align:left;">Most people just assume AI is ready. It’s not. <b>At least not for the most meaningful, economically valuable tasks</b>. </p><p class="paragraph" style="text-align:left;">And to know “where AI makes sense, and where it doesn’t”, as someone looking to adopt the technology, be that a CEO or a tech enthusiast, the same rules apply; the only thing that changes is scale (for good and bad).</p><p class="paragraph" style="text-align:left;">There are three factors you should consider: the particularities of <b>inference</b>, <b>economic</b>, and <b>technological</b>.</p><h3 class="heading" style="text-align:left;">Why AI inference is a real mess</h3><p class="paragraph" style="text-align:left;">To truly understand the situation in the AI industry, we need to grasp AI inference and the issues it entails.</p><p class="paragraph" style="text-align:left;">AI inference, serving AI models to users, <b>is what pays the bills for the entire industry</b>. At this point, <b>AI is almost synonymous with Generative AI</b>, the field in which models generate responses back to you, with examples like ChatGPT or Claude, at least in terms of investment.</p><p class="paragraph" style="text-align:left;"><b>These responses are made up of ‘tokens’</b>; words in text models, image patches for image and video generators, and so on.</p><p class="paragraph" style="text-align:left;">Since everything is tokenized (both the input you provide and the model&#39;s response), <b>AI is charged by the token</b>, making your ‘AI costs’ a function of both processed and generated tokens.</p><p class="paragraph" style="text-align:left;">In other words, the entire business is summarized in one simple inference (pun intended): <b>tokens equal revenue; the more tokens I generate, the more revenue I make</b>. Thus, the goal is to process and generate as many tokens as possible and charge more than they cost me.</p><p class="paragraph" style="text-align:left;">Sounds simple. However, the problem here lies in “charge more” and generating tokens in a way that I can make a return.<b> Because how much I have to charge to actually make money, once we account for all costs, is way more than we can charge now</b>. But we’ll get to that later.</p><p class="paragraph" style="text-align:left;">For most consumers, this is all irrelevant because we pay for subscriptions, as AI Labs try to hide the complexity of it all.</p><p class="paragraph" style="text-align:left;">But at heart, <b>everything is unitary</b> (meaning the Labs incur a cost for each token they process and generate, and hope the subscription value stays above the total cost of those tokens).</p><p class="paragraph" style="text-align:left;">And sadly, subscriptions are not profitable, and AI Labs lose money in the vast majority of them. For example, Anthropic recently claimed the average developer costs up to <a class="link" href="https://www.businessinsider.com/anthropic-claude-code-token-estimates-2026-4?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai" target="_blank" rel="noopener noreferrer nofollow">$13 per day</a>, or ~$260 for a 20-day work month. <b>In reality, nothing about AI is profitable these days</b>.</p><p class="paragraph" style="text-align:left;">And before we tackle the actual implications, I think it’s important to give you an idea of why the economics of AI are so broken and why the price of the technology, compared to previous technological improvements, <b>is actually going up</b>.</p><p class="paragraph" style="text-align:left;"><i>Are the new models just too big and costly? </i>Well, yes and no. Models are getting larger and costlier, but improved hardware largely eliminates that impact and, in fact, $/token is falling with each new hardware generation.</p><p class="paragraph" style="text-align:left;">Which is to say, the problem is not the hardware (though that is half true, as we see below), <b>but mostly the cost of purchasing it</b>; it’s just that they are spending so much money to grow that they have little option; there’s really nothing they can do about it because the hill of costs they are trying to climb is too steep.</p><p class="paragraph" style="text-align:left;">Let me explain why the cost structure of AI providers is so broken. <b>The real problem lies in the unit economics of what sits behind the AI, the hardware</b>, especially in the two big hurdles below.</p><ol start="1"><li><p class="paragraph" style="text-align:left;"><b>Our hardware isn’t optimized for inference</b> (serving AI models to users), which is the main driver of most compute demand. This means<b> the cost of producing tokens is still too high relative to how much companies can charge for them</b>.</p></li><li><p class="paragraph" style="text-align:left;"><b>Capital costs:</b> As mentioned, the cost of getting set up destroys the entire business case from the start.</p></li></ol><p class="paragraph" style="text-align:left;">You will have heard that AI inference is super profitable, with some people even quoting gross margins of 90%. That means that, for every dollar of sales, the gross profit is 90 cents.</p><p class="paragraph" style="text-align:left;"><b>The problem is that this number conveniently ignores both R&D</b> (Research and Development) and also the fact that<b> they quote gross margins to ignore operating costs</b>, and most importantly, cash flows, which is also telling, because that’s where the real problems reside (<b>capital costs</b>). Once you account for those,<b> the situation is much, much worse </b>(huge losses, basically).</p><p class="paragraph" style="text-align:left;"><i>But why is hardware suboptimized for inference?</i></p><p class="paragraph" style="text-align:left;">The reason is the steepness of the curve that defines the inference trade-off:</p><ul><li><p class="paragraph" style="text-align:left;">If we maximize token throughput (generating tokens), we sacrifice user latency, which skyrockets, and users say goodbye to you.</p></li><li><p class="paragraph" style="text-align:left;">If we minimize latency (optimize for tokens/second, also known as interactivity) to improve the user experience, <b>you&#39;re running your hardware at a massive discount </b>relative to what you could be making with it.</p></li></ul><p class="paragraph" style="text-align:left;">This gives us the famous throughput vs. interactivity curves, like the one below, which basically shows that<b> the more interactivity you want</b> (higher tokens/second per user), <b>the fewer tokens you get per</b><b> GPU</b>. As LLMs are billed by the token,<b> that means less revenue</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7e869f3d-55fc-42d3-8ee6-5b42d92be5ab/1_M9DZektyzqJccoBHOh8k1A.png?t=1779009688"/></div><p class="paragraph" style="text-align:left;"><i>And can’t we just stay at the top of the curve?</i> Well, it’s not that simple. Because AI models are commoditized, customers churn if your responses are slow.</p><p class="paragraph" style="text-align:left;">Therefore, <b>most</b><b> inference providers operate at the lower end of the curve</b>, ensuring a good user experience even if that means less revenue.</p><p class="paragraph" style="text-align:left;">Of course, as I was just suggesting, one option is to make the workload more ‘GPU-friendly’. In inference, that means moving up the curve to the left, increasing the number of tokens each GPU produces, <b>and thereby securing higher revenue relative to the hardware you have</b>.</p><p class="paragraph" style="text-align:left;">But if you do that, customers churn, so you’re forced to live in ‘suboptimized land’.</p><p class="paragraph" style="text-align:left;">The problem with this decision is that <b>you’re still running a race against time</b>, because your chips depreciate really fast and your hardware utilization is worse; <b>you’re paying a premium for hardware you’re then running at a discount</b>, like buying a Ferrari to drive it only through downtown Tallahassee.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">So, even if you want to offer the best user experience possible, you have to account for the fact that <b>every such decision makes it harder for you to ever make money.</b></p><p class="paragraph" style="text-align:left;">At this point, they have two options: <b>premium inference</b> and <b>new hardware</b>.</p><p class="paragraph" style="text-align:left;"><b>The former offers ultra-fast tokens at a huge premium</b>. If you’re going to be forced to live in the lower ends of the curve, producing way fewer tokens than I would need to make money at standard prices,<b> I’m going to offer a premium service and charge many times the usual price</b>.</p><p class="paragraph" style="text-align:left;">Nonetheless, Cursor and Anthropic both charge six times the price for tokens that are 2.5 times faster than the standard ones. <i>And guess what?</i> <a class="link" href="https://newsletter.semianalysis.com/p/cerebras-faster-tokens-please?utm_source=post-email-title&publication_id=6349492&post_id=197494856&utm_campaign=email-post-title&isFreemail=false&r=42kvwg&triedRedirect=true&utm_medium=email" target="_blank" rel="noopener noreferrer nofollow">People are paying</a>.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;"><b>The other option is using hardware with a better design for such workloads</b>. Although the inference trade-off is unavoidable (it’s just maths), the thing is that the nature of our hardware, particularly GPUs, which are built for workloads very different from AI inference decoding, just makes it even worse.</p><p class="paragraph" style="text-align:left;">One example that is more designed to live in that part of the curve is Cerebras, <a class="link" href="https://thewhitebox.beehiiv.com/p/a-real-transformer-robot-a-historical-ipo-more?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai" target="_blank" rel="noopener noreferrer nofollow">the hottest recent IPO</a>, which offers such high memory bandwidth per chip (the main bottleneck in inference) that even for <b>extremely sparse workloads </b>(e.g., using an entire cluster to serve a single user, the type of stuff you have to pull off in inference), <b>the amount of compute I get out of the chip is enormous </b>(left side of the below graph).</p><p class="paragraph" style="text-align:left;">However, as you can see below,<b> the issue is that for denser workloads</b> (i.e., training or the inference prefill stage), <b>Cerebras loses its appeal almost immediately</b> (and Cerebras is highly custom hardware, so it&#39;s very expensive to build, making the use case particularly off-putting).</p><p class="paragraph" style="text-align:left;">Nonetheless, for workloads like training, <b>a single R200 chip</b> (NVIDIA Rubin Chip, retailing at ~$50k) <b>delivers better throughput than a $1 million Cerebras chip</b>, offering 20 times the performance at 1/20th the price.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/081fd49b-ab60-4268-9452-de7a760c910f/0_duMf0LAi6ZFUxHhc.jpg?t=1779009688"/></div><p class="paragraph" style="text-align:left;">This means that, in reality, <b>there’s really no perfect solution</b>, and the ideal hardware depends on the situation. So,<i> what do we do if there’s no one-size-fits-all solution?</i></p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">The fact that the picture changes so much across different hardware types makes it very clear that the future of AI inference lies in <b>disaggregation</b>, with solutions such as NVIDIA’s SuperPoD or <a class="link" href="https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai" target="_blank" rel="noopener noreferrer nofollow">Amazon’s deal with Cerebras</a> to run inference on a mix of Trainium and WSE chips.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">In both instances, <b>‘GPU-like’ accelerators are used for the denser parts of the workload</b>, and the sparse parts are streamed to hardware such as Cerebras’s WSEs or Groq’s LPUs to keep throughput high (and thus, revenues piling in) while still offering a great user experience to users (fast and cheap tokens).</p><p class="paragraph" style="text-align:left;">So, <i>what’s the takeaway here?</i> Well, to me, it’s that the picture is way more complicated than many investors and enthusiasts alike think.</p><p class="paragraph" style="text-align:left;">Which leads well into what all this means to your wallet.</p><h3 class="heading" style="text-align:left;">AI on the margins</h3><p class="paragraph" style="text-align:left;">If you’re a regular reader, you know I’ve been screaming off the rooftops for quite some time that <b>AI is way more expensive than we think</b> (and the section above proves why). In a way, <b>we are getting spoiled as AI Labs are in full ‘Silicon Valley’ mode</b>, where growth is all that matters, and profits can come later.</p><p class="paragraph" style="text-align:left;">They are getting all of us hooked to this technology, and one day they’ll start raising prices, and there will be nothing you can do about it (well, as we’ll see later, there’s one thing).</p><p class="paragraph" style="text-align:left;">In fact, <a class="link" href="https://medium.com/@theimperative/diary-of-a-scared-executive-b5e440995c3d?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai" target="_blank" rel="noopener noreferrer nofollow">that day has already come</a>, and prices are going up from Copilot to Claude Code. Frontier AIs are no longer guaranteed to be priced lower than before. If anything, <b>they are raising prices</b>.</p><p class="paragraph" style="text-align:left;">Besides pushing unitary economics higher (charging more per token), <b>they are also transitioning to usage-based pricing</b>, ensuring they earn a margin on every token they provide instead of charging a flat subscription price.</p><p class="paragraph" style="text-align:left;">This leads us to one of the most important graphs in all of AI: understanding how software’s cost curves change with this technology.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/4ad3f291-dbc5-4fe2-9c45-bd6e9a12414a/image.png?t=1779009341"/></div><p class="paragraph" style="text-align:left;">In traditional software, you were faced with an important capital investment every few years to buy equipment. After that purchase, even though user numbers increased, costs remained roughly stable and, importantly, predictable.</p><p class="paragraph" style="text-align:left;">Cloud simplified things even more. You still had an upfront cost (the cloud migration), but from then on, <b>you’re just dealing with stable costs all the way</b>. Don’t get me wrong, they might have a slight upward trend over time, but not too crazy.</p><p class="paragraph" style="text-align:left;">And crucially, <b>onboarding new users incurs negligible additional cost</b>. For instance, onboarding Jack from marketing to HubSpot represents an almost negligible increase in costs, just the price of the seat.</p><p class="paragraph" style="text-align:left;">As most software is purchased through SaaS deals, you can’t see what’s going on under the hood, but for HubSpot, <b>that new user represents negligible cost increases over the underlying hardware</b>, which is why traditional software has such high margins.</p><p class="paragraph" style="text-align:left;">For you, it’s still a good deal most of the time because while seat prices do go up almost religiously every year, <b>you can predict the behavior of your IT hardware spending</b>, even though you may not realize you’re paying a huge premium for that software.</p><p class="paragraph" style="text-align:left;">Well, <b>all this goes to shit with AI</b>.</p><p class="paragraph" style="text-align:left;">Because with this technology, <b>every new user counts</b>. And to prove this, I’m going to show you one of the wildest metrics in the history of this technology, and that’s saying something, because you’re not prepared to know how much some users are spending on AI.</p><p class="paragraph" style="text-align:left;">Behind the paywall, we move from the industry-level picture to the practical reality of deploying AI: the <b>cost traps</b> that appear once usage scales, the <b>technical limits</b> that still make automation harder than many assume, the <b>hidden fragilities</b> companies need to design around, the <b>operating principles</b> that separate serious AI adoption from superficial experimentation, and a set of <b>final recommendations</b> to adopt AI for anyone remotely serious about this technology.</p></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=my-honest-sober-opinion-on-ai">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=2e331d5e-a2ea-47fa-be17-ee98e13100f0&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>A Real Transformer Robot? A Historical IPO, &amp; More</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/eef88c87-32b8-4e4d-b4f9-29c9b8200c0b/image.png" length="267241" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/a-real-transformer-robot-a-historical-ipo-more</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/a-real-transformer-robot-a-historical-ipo-more</guid>
  <pubDate>Fri, 15 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-15T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/60b717c4-3783-475e-b9fe-eb8e3be93988/image.png?t=1759994834"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>TLDR;</h2><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Welcome back!</b></span> This week, we have <b>robots</b> from the future, an <b>historical IPO</b> (with detailed technical coverage at the end of the newsletter), a star Lab’s first model release, OpenAI’s personal finance tool, and many other news items you must know.</p><p class="paragraph" style="text-align:left;"><b>Enjoy!</b></p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a5bcefd-e9fc-4b09-bacf-d4b9eb8ea3c3/image.png?t=1759994596"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>A New Type of Generative Model?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/9a9e4625-edf0-4295-b9d1-73caf12dac08/image.png?t=1778850469"/></div><p class="paragraph" style="text-align:left;"><b>Thinking Machines</b>, one of the most highly valued AI Labs in the private markets and packed with star researchers and engineers from other top Labs, <a class="link" href="https://thinkingmachines.ai/blog/interaction-models/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">is finally showing its teeth</a>.</p><p class="paragraph" style="text-align:left;">And the results are very promising because they aren’t simply following the same playbook as other top Labs; <b>they are actually pushing novel research</b> (and quite publicly, too).</p><p class="paragraph" style="text-align:left;">Their first major model release is an <b>interaction model</b>, trained to offer a strong balance between intelligence and interactivity. The trade-off here is obvious because <b>interactivity goes against both scaling laws of intelligence</b>:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">The smarter you want a model to become, the larger it is (and thus slower to run)</p></li><li><p class="paragraph" style="text-align:left;">The more thinking budget you give it, the slower its response will be, too.</p></li></ol><p class="paragraph" style="text-align:left;">The solution is an architecture that may sound familiar if you’re a regular reader of this newsletter because it draws strong similarities to how robotics AI is being approached, but also includes unique features.</p><p class="paragraph" style="text-align:left;">It’s inspired by Daniel Kahneman’s System 1 and System 2 ways of thinking. Known as Thinking, Fast and Slow, <b>it argues that our brains think at two speeds, one fast one slow:</b></p><ul><li><p class="paragraph" style="text-align:left;">one is intuition-based, <b>fast</b>, unconscious, like flinching when someone is about to hit you,</p></li><li><p class="paragraph" style="text-align:left;">and the other is <b>slow</b>, conscious, and deliberate, like solving a maths test.</p></li></ul><p class="paragraph" style="text-align:left;">Following this principle, and as the thumbnail suggests, here we have an architecture with a fast model, known as the <b>interaction model</b>, which is the one in charge of interacting with the user, and the <b>background model</b>, which is in charge of the ‘slow thinking’ and also of making the necessary tool calls (like calling a search API to check something on the Internet).</p><p class="paragraph" style="text-align:left;">This way, the model can still handle long-range, reasoning-heavy tasks that require time while still appearing snappy and “real-time.” <b>The result is a system that breaks with the way AIs have historically interacted with us, on a turn-by-turn basis</b>.</p><p class="paragraph" style="text-align:left;">Instead, both you and the model interact continually across all time slots. Therefore, the AI can talk alongside you and across multiple modalities, such as audio, video, or text.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0f85d8a6-ea45-4fc7-bf15-78b04fb8e5cb/image.png?t=1778850743"/></div><p class="paragraph" style="text-align:left;">These allow this “model” to do stuff we hadn’t seen before, certainly not at these levels of interactivity:</p><ul><li><p class="paragraph" style="text-align:left;"><b>It can handle real-time multimodal interaction</b>: audio, video, text, screen input, and tool outputs in one continuous stream.</p></li><li><p class="paragraph" style="text-align:left;"><b>It can process interaction every 200 ms</b>, so it understands pauses, interruptions, overlapping speech, timing, hesitation, and turn-taking.</p></li><li><p class="paragraph" style="text-align:left;"><b>It can speak while listening</b>, allowing simultaneous conversation, live translation, corrections, and backchanneling.</p></li><li><p class="paragraph" style="text-align:left;">It can watch what you are doing live, spot bugs, follow a screen, count reps, comment on a video feed, or react to visual changes.</p></li><li><p class="paragraph" style="text-align:left;">It can coordinate with tools and background agents, meaning one part of the system can keep talking to you while another searches, reasons, browses, or executes longer tasks.</p></li><li><p class="paragraph" style="text-align:left;">It can generate or update interfaces while interacting, making it useful for live coding, dashboards, design work, research workflows, and interactive tutoring.</p></li></ul><p class="paragraph" style="text-align:left;">Furthermore, the model is very smart and competitive with alternatives from OpenAI, such as GPT -2, in real time, across several benchmarks.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">As we’ve learned from Cerebras&#39; IPO and from OpenAI and Anthropic’s super-successful fast modes, <b>there’s a lot of value in optimizing interactivity; people don’t like to wait</b>.</p><p class="paragraph" style="text-align:left;">It’s interesting that TML went this route to start their model releases, but it certainly looks good.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>CODING AGENTS</b></span><br>Are Task-specific harnesses the real deal?</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0f39f6fc-719f-41f5-9bc0-76935c69ed4a/image.png?t=1778840357"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://poetiq.ai/posts/recursive_self_improvement_coding/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Poetiq&#39;s Meta-System</a> autonomously developed a coding harness from scratch through recursive self-improvement, achieving new state-of-the-art performance across several benchmarks and frontier models.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Using only standard APIs, with no fine-tuning or access to model weights (meaning they cannot adapt the model to the problem), <b>the harness achieved state-of-the-art results on LiveCodeBench Pro</b>, a contamination-resistant C++ coding benchmark that evaluates pure programming ability without relying on public ground-truth code or tool use.</p><p class="paragraph" style="text-align:left;">Importantly, this harness is both model-agnostic, initially optimized for Gemini 3.1 Pro, raising its score from 78.6% to 90.9%, surpassing Google&#39;s own Gemini Deep Think (Google’s own harness and multi-agent setting), but they also applied the same harness, unchanged, to other models, and still had GPT 5.5 High improve from 89.6% to 93.9% <b>and boosted Kimi K2.6 by nearly 30 percentage points</b>.</p><p class="paragraph" style="text-align:left;">Importantly, it’s also task-specific, which is the key secret here; <b>the harness is built for the task in particular, which is what makes models see such huge bumps in performance</b>.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I do still believe most of the real value comes from <b>actually fine-tuning models on the task.</b></p><p class="paragraph" style="text-align:left;">But the idea of task-specific harnesses, which make models perform particularly well on a specific task without additional training, <b>is a neat middle-ground way to customize models you can’t control</b> (i.e., proprietary models) to your task.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>AI BEYOND LLMs</b></span><br>Time-series models scale</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.datadoghq.com/blog/ai/toto-2/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Datadog has released Toto 2.0</a>, a new family of open-weight time-series forecasting models designed to predict numerical signals over time.</p><p class="paragraph" style="text-align:left;">The models range from 4 million to 2.5 billion parameters and are available on GitHub and Hugging Face.</p><p class="paragraph" style="text-align:left;"><b>Datadog says the release is intended to test whether time-series foundation models improve as they scale</b> and reports that each larger Toto 2.0 model improves over the smaller one, with no sign of saturation at the largest 2.5-billion-parameter size.</p><p class="paragraph" style="text-align:left;"><i>But why do we need this?</i></p><p class="paragraph" style="text-align:left;">It’s easy to see the value in a model like ChatGPT, which, with scaling, gives it the ability to speak and reason better, <i>but a model that simply predicts lines on a graph?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is needed anywhere organizations depend on numerical signals that change over time and need to anticipate what comes next.</p><p class="paragraph" style="text-align:left;">In Datadog’s core market, that means forecasting cloud infrastructure metrics such as latency, error rates, traffic, CPU usage, memory consumption, database load, GPU utilization, and cost, so teams can detect anomalies earlier, prevent incidents, and plan capacity more efficiently.</p><p class="paragraph" style="text-align:left;">More broadly, the same type of model can be useful for retail demand planning, energy consumption forecasting, financial time series, logistics, weather-sensitive operations, and any environment where thousands or millions of signals are too many to model manually.</p><p class="paragraph" style="text-align:left;">So there’s definitely value, especially inside certain sectors, and seeing that the principles of scaling that have made LLMs so valuable apply to other areas too is great news.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7effc623-4559-41b9-978a-2f478190bdf8/image.png?t=1759994520"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>PUBLIC MARKETS</b></span><br>Cerebras’ Historic IPO</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/eef88c87-32b8-4e4d-b4f9-29c9b8200c0b/image.png?t=1778853881"/></div><p class="paragraph" style="text-align:left;">If markets weren’t crazy enough as they are, we’ve <b>the biggest IPOs of the year</b>, Cerebras Systems, to the mix, <a class="link" href="https://www.reuters.com/legal/transactional/cerebras-set-debut-stock-market-gripped-by-ai-mania-2026-05-14/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">which went public yesterday</a>.</p><p class="paragraph" style="text-align:left;">The AI chipmaker priced its IPO at $185 per share, raising $5.55 billion. Trading opened far above that level, <b>with shares starting at $350</b>, briefly climbing higher, and closing at $311.07, up 68% on the day.</p><p class="paragraph" style="text-align:left;">Today, it’s down <a class="link" href="https://stocktwits.com/symbol/CBRS?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow" style="text-decoration: none; font-style: normal;"><span style="color:#DC2626;">$CBRS ( ▼ 4.21% )</span></a> .</p><p class="paragraph" style="text-align:left;"><i>But should you invest?</i></p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">To me,<b> there’s a mixture between delusion and warranted optimism here</b>. It’s a great company, and the chip is an engineering marvel, <i>but</i> <i>what’s the right price?</i></p><p class="paragraph" style="text-align:left;">I was asked to do a simplified due diligence for an investment banking client last week. In the portfolio section at the end of this newsletter, <b>you’ll find what I told them</b>. I cover everything, from the technological perspective, to financials and supply chain, as well as how much dent they can make in the AI market.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>RESEARCH</b></span><br>Ineffable partners with NVIDIA</h2><p class="paragraph" style="text-align:left;">Ineffable, the AI Lab co-founded by David Silver, one of the founding fathers of Reinforcement Learning (RL) and the creator of AlphaGo, the first superhuman AI, <a class="link" href="https://www.ineffable.ai/blog/nvidia-ineffable-intelligence-team-up-to-build-the-future-of-reinforcement-learning-infrastructure?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">has partnered with NVIDIA</a> to develop—potentially—new hardware for a fundamentally different approach to AI training using Large Language Models (LLMs).</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">The reason I’m highlighting this is <b>to make you aware of this lab</b>, as it&#39;s fundamentally different from what we have right now. They don’t believe the current state of AI is in the right direction.</p><p class="paragraph" style="text-align:left;">As David Silver himself explained, <i>“Researchers have largely solved the easier problem of AI: how to build systems that know all the things humans already know,”</i> Silver said. <b><i>“But now we need to solve the harder problem of AI: how to build systems that discover new knowledge for themselves. That requires a very different approach — systems that learn from experience.”</i></b></p><p class="paragraph" style="text-align:left;">In plain English, <b>he basically thinks LLMs are fine but not the endgame.</b></p><p class="paragraph" style="text-align:left;"><i>But how are they going to do that?</i> We don’t actually know, but there are several hints along the way. The first is the paper he coauthored with Rich Sutton, the other Godfather of reinforcement learning, <a class="link" href="https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">the era of experience</a>.</p><p class="paragraph" style="text-align:left;">It’s an interesting read, but simplified, it’s the idea that while modern AIs learn from human data, <b>learning by imitating us and then adding a sprinkle of exploration and experience-gathering</b>.</p><p class="paragraph" style="text-align:left;">Instead, <b>they believe that truly powerful AIs turn the balance upside down</b>; it’s fine if AIs learn from human data, but the vast majority of learning should come from their own experiences.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">The secret might be feedback. In other words, for AIs to learn without being told whether they are making progress (the verifying principle; we can only learn what we can verify), <b>they will experience the consequences of their actions and use that as verification</b> (e.g., I jump from a height of 10 feet and break my leg; never jumping from that high again).</p><p class="paragraph" style="text-align:left;">The real world becomes the verifier, and that’s when you truly learn from experience.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">The idea makes sense. But just like Yann LeCun’s ideas with JEPAs, for which he’s started a new AI Lab to build world models that have nothing to do with LLMs, <b>I need to see the receipts. I like what I read, but I need proof</b>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>VENTURE CAPITAL</b></span><br>Another $30 billion, really?</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.ft.com/content/9deae3c6-716d-4f4d-8b09-434d8519f847?syn-25a6b1a6=1&utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Anthropic has agreed terms</a> for a $30bn fundraising that would value the artificial intelligence company at $900bn, according to the Financial Times.</p><p class="paragraph" style="text-align:left;"><b>The round is expected to close as soon as this month</b>, though terms could still change before completion. The deal would nearly triple Anthropic’s previous valuation and place the Claude maker above OpenAI’s reported $852bn valuation.</p><p class="paragraph" style="text-align:left;">The valuation increase reflects the company’s rapid revenue growth, <b>with annualized revenue projected at about $45bn</b>, up from $9bn last year. Big tech groups are not expected to take part in the round.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">I’m old enough to remember this company raised the same amount of money <a class="link" href="https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">3 months ago</a>, <b>which puts their annualized spending at a whopping $120 billion</b>.</p><p class="paragraph" style="text-align:left;">So even if people can linearly extrapolate and assume Anthropic will double revenues by the end of this year (that’s what the current growth suggests), <b>which they probably won’t </b>because, even if there was such demand, <b>they largely don’t have the compute to serve it, they are still spending way more.</b></p><p class="paragraph" style="text-align:left;">This industry is finally seeing revenues piling in, but the elephant in the room is that, <b>for every new dollar that comes in, more than one dollar has to come out</b>, so spending is accelerating even faster than revenues. Ironically, this might be an opportunity for Cerebras, as we’ll talk about later.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>PICTURE</b></span><br>Netflix’s INKubator</h2><p class="paragraph" style="text-align:left;"><a class="link" href="https://www.animationmagazine.net/2026/05/netflix-staffing-for-inkubator-ai-powered-experimental-animation-studio/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Netflix is staffing up a new internal animation studio called INKubator</a>, which is expected to focus on AI-assisted animated shorts and specials. The unit is hiring producers, software engineers, and CG artists to work on experimental “GenAI-native” production pipelines. </p><p class="paragraph" style="text-align:left;"><b>The studio has not been formally announced by Netflix</b>, but listings and LinkedIn profiles suggest it quietly began taking shape earlier this year. The initiative appears focused first on short-form animation, while some listings point to longer-term ambitions for feature-quality content.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">Seeing the progress image/video generation models are making, especially with OpenAI’s GPT-image-2, I’m sorry, but this is a no-brainer we will have to get used to.</p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">For instance, an X user tested the following: He uploaded an image they had created using AI that resembled a Monet painting and asked others to critique it. <b>The criticisms were varied and mostly negative</b>: “there’s no cohesion”, “it’s all borked nonsense”, “it’s a mess”, and a long tail of others.</p><p class="paragraph" style="text-align:left;">The reality, though, is that this was, in fact, an actual Monet.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/40e80747-9c48-4017-8822-059d4a84e5ff/image.png?t=1778843214"/></div><p class="paragraph" style="text-align:left;">Which begs the question: <b><i>are people against AI just for the sake of it?</i></b> </p><p class="paragraph" style="text-align:left;">To be fair, I think this is just the same reaction that we see with every industrial revolution. If you take a fabricated leather shoe that looks man-made, and ask people to critique it, they will find ways to explain why it’s not as great as a man-made version.</p><p class="paragraph" style="text-align:left;">Interestingly, currently, <b>AI is as unpopular as ever</b>. A new Gallup survey shows that people are strongly opposed to having data centers nearby <b>and would rather have a fricking nuclear reactor</b>.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/e36e7620-1807-42b0-b3fa-c2388de792d3/image.png?t=1778843398"/></div><p class="paragraph" style="text-align:left;">And while I’m the first to acknowledge that the industry incumbents are the first to blame for all of this, and data centers aren’t precisely fountains of job creation, this is just madness. <b>People hate AI and will cling to any reason for it</b>, point-blank.</p></div><div class="section" style="background-color:#222222;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/03892b19-a66b-4099-9947-88e7858efd7d/image.png?t=1759994450"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>CODING</b></span><br>You can now access Codex from your ChatGPT iPhone app.</h2><p class="paragraph" style="text-align:left;">If my dog park human colleagues already probably thought I was somewhat shy, being all the time listening to podcasts instead of making small talk (in case you’re wondering how I manage to write so much about AI, that’s the kind of thing I have to do), <a class="link" href="https://openai.com/index/work-with-codex-from-anywhere/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Codex might make me look like a sociopath</a> after their new release that lets you talk to your models and processes on your laptop from your iPhone.</p><p class="paragraph" style="text-align:left;">A feature available to Claude Code users <a class="link" href="https://www.oneusefulthing.org/p/claude-dispatch-and-the-power-of?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">for a while now</a>, it’s finally available in Codex, and it’s an amazing experience. In an hour-long dog walk, <b>I’ve pushed several new features to an app I’m building</b> (don’t worry, I am very aware of the tech debt and always take time to clean code afterward) simply by talking to my Codex agents.</p><p class="paragraph" style="text-align:left;">Truly incredible times we’re living in.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">A word has to be said on the “quality” of the code we’re seeing in the world right now (including probably mine).</p><p class="paragraph" style="text-align:left;">As some people are tracking commits to GitHub that are at least coauthored by Claude and Codex, <b>they have skyrocketed</b> (and I suspect the actual number is way higher because you can choose to exclude coauthors from the commit), <b><i>and the numbers below are just from February!</i></b></p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0052a6fc-bc0b-4963-9071-8f68616811e8/image.png?t=1778864947"/><div class="image__source"><span class="image__source_text"><p>Source: SemiAnalysis</p></span></div></div><p class="paragraph" style="text-align:left;">The scale at which these agents write code makes it very, very hard to believe <b>this code is being audited, like at all (unless it’s being done by other AIs).</b> For good measure, all the apps I’m building are for personal use and/or for my business; no public app accessible to other users has my name and Codex’s signature on it.</p><p class="paragraph" style="text-align:left;">Personally, I am very aware of tech debt and obsessively refactor my code almost daily after long sessions.</p><p class="paragraph" style="text-align:left;">Still, <b>I fear the quality might not be superb because it’s all AI-generated</b>. At this point, the only thing we can hope for is for future models to be excellent debuggers. Otherwise, it’s going to be a mess.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>ROBOTICS</b></span><br>Unitree’s Transformer robot</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8fe71da0-c2f1-48ba-9523-03a0e6f547bd/image.png?t=1778851164"/></div><p class="paragraph" style="text-align:left;">In what’s possibly one of the coolest robotics demos, or even <b>the coolest robotics demo ever</b>, <a class="link" href="https://x.com/UnitreeRobotics/status/2054067819634159622?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Unitree has presented what seems to be a real-life version of a Hollywood transformer</a>; yes, those that fly and fight and do all that crazy stuff.</p><p class="paragraph" style="text-align:left;">The video shows Unitree’s founder, Xingxing Wang, literally riding this thing, and boy must he feel like a Superhero.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;"><b>Cool demo that is obviously more about collecting bragging rights than actual utility</b>, but I could see this thing being used for dangerous construction work and other situations where we still need human motion but want the human out of harm’s way.</p><p class="paragraph" style="text-align:left;">Of course, some people will see this and imagine the end of humanity, but those people have probably seen too many Hollywood films.</p><p class="paragraph" style="text-align:left;">As far as we know, <b>this is fully human-controlled</b>.</p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>CONSUMER HARDWARE</b></span><br>Google’s GoogleBook and the Magic Pointer</h2><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/73832e92-299e-4ae4-93e1-818cb5b47302/image.png?t=1778842589"/></div><p class="paragraph" style="text-align:left;"><a class="link" href="https://googlebook.google/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Google has announced Googlebook</a>, a new premium laptop category designed around Gemini.</p><p class="paragraph" style="text-align:left;">The company says Googlebook will combine Android technologies with ChromeOS capabilities, <b>while putting Gemini directly into the core laptop experience</b> rather than leaving it as a separate chatbot or assistant.</p><p class="paragraph" style="text-align:left;">One of the main new features is <a class="link" href="https://x.com/GoogleDeepMind/status/2054246119635300451?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">Magic Pointer</a>, developed with Google DeepMind. <b>The idea is to turn the computer cursor into an AI interface</b>: users point at something on screen, such as a date, image, table, object, or location, <b>and Gemini can understand what is being referenced and suggest an action</b>.</p><p class="paragraph" style="text-align:left;">DeepMind describes this as a rethink of the mouse pointer, a 50-year-old interface that has mostly remained a tool for clicking, dragging, and selecting. <b>With Magic Pointer, the cursor becomes a way to tell the AI what “this” or “that” means</b>, keeping users in their current workflow rather than forcing them to copy text, upload files, or manually describe context.</p><p class="paragraph" style="text-align:left;">Google says the first Google Books will come from partners including Acer, ASUS, Dell, HP, and Lenovo, <b>with availability expected in the fall</b>. The company has not yet shared full specifications or pricing.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">We’ll see how far this idea goes. <b>Google&#39;s entry into the laptop market is interesting</b>. I’m particularly intrigued by the hardware; <i>will it be strong enough to run local models like Gemma?</i></p><p class="paragraph" style="text-align:left;">I hope it is, otherwise it’s just a waste of my time.</p><p class="paragraph" style="text-align:left;">The Magic Pointer idea looks incredible, but I’m not convinced it will become something I can’t live without. <i>Will it be another solution looking for a problem to solve?</i></p><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>FINANCE</b></span><br>ChatGPT’s New Personal Finance Feature</h2><p class="paragraph" style="text-align:left;">As always just when I’m closing the edit of my newsletter for sending, an AI Lab decides it’s the perfect time to drop a feature I’m compelled to talk about.</p><p class="paragraph" style="text-align:left;">Now, <a class="link" href="https://openai.com/index/personal-finance-chatgpt/?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more" target="_blank" rel="noopener noreferrer nofollow">OpenAI has launched a preview of a new personal finance experience in ChatGPT for Pro</a> users in the US.</p><p class="paragraph" style="text-align:left;">Users can connect financial accounts, view a dashboard of spending, investments, subscriptions, upcoming payments, and liabilities, and ask ChatGPT questions grounded in their own financial data. The rollout starts on web and iOS, with support for more than 12,000 financial institutions through Plaid, and Intuit support is planned.</p><h3 class="heading" style="text-align:left;">TheWhiteBox’s takeaway:</h3><p class="paragraph" style="text-align:left;">This is basically a copy of what I am building for myself; the functionality appears very similar. I’m interested in knowing how complex the solution is, whether it allows you to change categories, ask for very specific KPIs and data, and support other features that my app does.</p><p class="paragraph" style="text-align:left;">Considering it’s a first release, I would bet it will be somewhat limited, but it takes no genius to see the appeal of this: having a single place to manage all your personal finances.</p></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><h2 class="heading" style="text-align:left;" id="closing-thoughts">Closing Thoughts</h2><p class="paragraph" style="text-align:left;">The biggest news of the week is, of course, the Cerebras IPO, which I talk about in excruciating detail below. Beyond that, <b>we’ve seen a lot of “what the world might give us in the future”</b> and not much about the present, aside from a couple of OpenAI releases.</p><p class="paragraph" style="text-align:left;">For a week, it seems AI is going back to ignoring the present and trying to imagine an incredible future, the type of stuff you do when proof of reality is not available.</p><p class="paragraph" style="text-align:left;">However, personally, I think it’s about time we realize the size of this industry, and the huge debt and invested capital around it make all that future talk pointless. <b>We need results now.</b></p><p class="paragraph" style="text-align:left;">For what it’s worth, we’re getting them, but as Anthropic’s latest round shows, just when we thought AI had cracked the revenue code, we see that the corresponding losses accelerate even faster. Not a good sign.</p><p class="paragraph" style="text-align:left;">And without further ado, I leave you my analysis of one, if not the most, fascinating company I’ve covered in this newsletter, probably ever.</p><p class="paragraph" style="text-align:left;"><i>But is the stock market way too excited?</i> Let’s find out.</p><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="button" style="text-align:center;"><a target="_blank" rel="noopener nofollow noreferrer" class="button__link" style="background-color:#FF5632;" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=a-real-transformer-robot-a-historical-ipo-more"><span class="button__text" style=""> Upgrade to Full Premium to continue reading </span></a></div></div><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/53f8bf15-00e7-4c53-851b-306c6da49f9f/image.png?t=1759853205"/></div></div><div class="image"><img alt="" class="image__image" style="border-radius:0px 0px 0px 0px;border-style:solid;border-width:0px 0px 0px 0px;box-sizing:border-box;border-color:#E5E7EB;" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b67649ed-f616-4dd7-97fc-997f8e172571/image.png?t=1721747392"/></div><div class="section" style="background-color:#F9FAFB;border-color:#FF5632;border-radius:5px;border-style:solid;border-width:1px;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=c7747071-d564-4dc5-955e-912dbb47f293&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

      <item>
  <title>We finally uncovered what AIs REALLY think.</title>
  <description></description>
      <enclosure url="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aca78ed3-2c43-4b8d-971b-49bfdffb7fcc/image.png" length="1704419" type="image/png"/>
  <link>https://thewhitebox.beehiiv.com/p/we-finally-uncovered-what-ais-really-think</link>
  <guid isPermaLink="true">https://thewhitebox.beehiiv.com/p/we-finally-uncovered-what-ais-really-think</guid>
  <pubDate>Tue, 12 May 2026 21:00:00 +0000</pubDate>
  <atom:published>2026-05-12T21:00:00Z</atom:published>
    <dc:creator>Ignacio de Gregorio Noblejas</dc:creator>
  <content:encoded><![CDATA[
    <div class='beehiiv'><style>
  .bh__table, .bh__table_header, .bh__table_cell { border: 1px solid #BFBFBFFF; }
  .bh__table_cell { padding: 5px; background-color: #222222FF; }
  .bh__table_cell p { color: #2D2D2D; font-family: 'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
  .bh__table_header { padding: 5px; background-color:#FF5632; }
  .bh__table_header p { color: #2A2A2A; font-family:'Syne',Helvetica,Arial,sans-serif !important; overflow-wrap: break-word; }
</style><div class='beehiiv__body'><div class="section" style="background-color:#FF5632;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f3b30d64-0c96-4851-8a63-3c5907e5a84f/image.png?t=1760260675"/></div></div><div class="section" style="background-color:transparent;margin:0.0px 0.0px 0.0px 0.0px;padding:0.0px 0.0px 0.0px 0.0px;"><h2 class="heading" style="text-align:left;"><span style="color:rgb(255, 86, 50);font-size:0.8rem;"><b>THEWHITEBOX</b></span><br>We finally uncovered what AIs REALLY think.</h2><p class="paragraph" style="text-align:left;"><i>Have you ever wondered why this newsletter is called TheWhiteBox?</i> For the sake of my ego, I’m going to assume you ask yourself that question first thing in the morning, every day.</p><p class="paragraph" style="text-align:left;">Jokes aside,<b> the idea was that AI is a “black box”</b> (we don’t know how they work, like, at all), and my newsletter would try to demystify it. Rather pretentious, but what marketing giveth, marketing taketh.</p><p class="paragraph" style="text-align:left;">But let me tell you something: <i>Did you know that a model might be thinking bad things about you while sounding completely normal?</i></p><p class="paragraph" style="text-align:left;">Well, it’s true. Until now, we suspected it, <b>but we can now </b><b>see it </b>thanks to Anthropic’s new research,<b> Natural Language Autoencoders</b> <a class="link" href="https://transformer-circuits.pub/2026/nla/index.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-finally-uncovered-what-ais-really-think#introduction" target="_blank" rel="noopener noreferrer nofollow">(NLAs)</a>, which begin to answer the question of what lies beneath ChatGPT’s overly formal and professional demeanor.</p><p class="paragraph" style="text-align:left;">And the answer is a mixture of candidness, honesty, lies, anger, scheming, and many other behaviors and, for lack of a better term, “beliefs” and “sentiments” that are now emerging and that you should definitely be aware of.</p><p class="paragraph" style="text-align:left;"><i>The opportunity?</i> An entirely new AI industry worth hundreds of billions in AI assurance.</p><p class="paragraph" style="text-align:left;"><span style="color:#FF5632;"><b>Let’s dive in.</b></span></p><h2 class="heading" style="text-align:left;">From Neurons to Thoughts</h2><p class="paragraph" style="text-align:left;">To fully comprehend the huge implications of today’s research, it’s important that we transition ourselves from seeing AIs for what they look like to what they actually are, <b>so that we can understand the fundamentals of interpretability</b>, the field that aims to uncover AI’s biggest mysteries.</p><h3 class="heading" style="text-align:left;">The basic structures</h3><p class="paragraph" style="text-align:left;">On paper, AIs are just a bunch of elements called neurons that interact and combine to produce the output.</p><p class="paragraph" style="text-align:left;">We call them ‘neurons’ because they are tightly connected to each other, much like human brain neurons, and they also have the “fire or mute” behavior, too.</p><p class="paragraph" style="text-align:left;">It’s a “little bit” more complicated, but at the heart of every AI model you’ve interacted with lately, sits something like this:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/f87ac77c-363b-4de9-be3a-bc21b1358e18/image.png?t=1778569047"/></div><p class="paragraph" style="text-align:left;"><i>And how does it work?</i></p><p class="paragraph" style="text-align:left;">Say we have an animal classification model that takes in several images of a specific animal and outputs its name.</p><p class="paragraph" style="text-align:left;">We feed the three inputs to the model, which processes them through the hidden layers to identify common patterns and determine the likely animal, for example generating a text response, “horse,” to three inputs depicting horses.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0077331d-f88e-46b7-a71e-828554d65167/image.png?t=1778569556"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">This is something ChatGPT would do, though the actual architecture under the hood is way more complicated than this. But at heart, it’s just a “bunch of neurons.”</p><p class="paragraph" style="text-align:left;">In reality, what is going on under the hood looks something like the image below.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8ee4bb78-8ebb-475e-986b-41abde8a6ae8/image.png?t=1778569967"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">I know, this sounds awfully mathematical and complex, but don’t worry, I’m not going to bore you with maths. Instead, <b>what we’re going to do is reveal what these activation circuits, like the one depicted in red, actually mean</b>.</p><p class="paragraph" style="text-align:left;">That is, the issue with these neurons is that, taken at face value, they are gibberish; they are a bunch of numbers magically understanding that there were horses in the images. <b>Taken at face value, they reveal very little about their nature</b>, making AIs look like black boxes.</p><p class="paragraph" style="text-align:left;">However, neurons can be surprisingly revealing when viewed through the right lens. They aren’t just a bunch of numbers,<b> and they actually “encode” meaning</b>.</p><p class="paragraph" style="text-align:left;"><i>But what do I mean by that?</i></p><h3 class="heading" style="text-align:left;">The idea of representations</h3><p class="paragraph" style="text-align:left;">Machines only understand numbers, so every concept we present to them must be ‘represented’ as numbers. Therefore, an ‘AI representation’ is just a mathematical representation of a real concept.</p><p class="paragraph" style="text-align:left;">I’m not going to get into the weeds of this, <a class="link" href="https://thewhitebox.beehiiv.com/p/deepseek-the-start-of-the-chinese-era-of-ai?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-finally-uncovered-what-ais-really-think" target="_blank" rel="noopener noreferrer nofollow">which would warrant an entire additional piece I already wrote recently</a>. For today, it’s more than enough to simply internalize two things:</p><ol start="1"><li><p class="paragraph" style="text-align:left;">Every concept can be represented in numbers</p></li><li><p class="paragraph" style="text-align:left;">AIs take concepts and transform them into other concepts.</p></li></ol><p class="paragraph" style="text-align:left;">A way to understand <b>representations is to view them as a list of attributes</b>; a pelican will have a ‘1’ for ‘beak or no beak’, a ‘1’ for ‘feathered’, and a ‘0’ for mammal, among multiple others:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7cd59128-947e-42da-96d7-57c862408835/image.png?t=1778571676"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">But the crucial thing to understand about AI models is point two, the transformation part.</p><p class="paragraph" style="text-align:left;">For instance, if we put the word pelican in a sequence such as ‘The pink pelican,’ <b>we can ‘transform’ the representation of pelican into a pink one</b> by turning on the ‘pink’ attribute.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/8dd143ee-4f6e-43b8-a678-56720caeacaf/image.png?t=1778571723"/><div class="image__source"><span class="image__source_text"><p>Source: Author</p></span></div></div><p class="paragraph" style="text-align:left;">If you’ve understood this, <b>you can actually say you understand neural networks</b>, at least Transformers, the overwhelming majority of modern AIs today (including all the ones you know by name), <b>because this is “all” they do</b>.</p><p class="paragraph" style="text-align:left;">Simplified much, they take a bunch of concepts represented as numbers in, and they apply transformations to them, creating new representations that lead to the desired prediction; “Draw me a pink pelican” is something ChatGPT can do because it can transform a pelican, the word it knows, into a pink one by applying such transformations internally.</p><p class="paragraph" style="text-align:left;">But again, we run into the same problem: interpretability. We infer that this is what is going on, but we have no way to actually “see” it, as the model doesn’t tell us “I’m thinking about a pink pelican” or whether a specific number in a representation makes it a mammal.</p><p class="paragraph" style="text-align:left;">Thus, <i>how do we decode those numbers?</i></p><h3 class="heading" style="text-align:left;">Decoding the encoding</h3><p class="paragraph" style="text-align:left;">Probably one of the biggest contributions Anthropic has ever made to this industry, given they hardly publish anything, <a class="link" href="https://transformer-circuits.pub/2025/attribution-graphs/methods.html?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-finally-uncovered-what-ais-really-think" target="_blank" rel="noopener noreferrer nofollow">is the discovery of monosemantic circuits and attribution graphs</a>, which sound scarier than they are.</p><p class="paragraph" style="text-align:left;">In layman’s terms, AI models can represent and create concepts like the ones we’ve been describing by combining neurons into identifiable circuits.</p><p class="paragraph" style="text-align:left;">Once we can map certain neuron activation circuits to certain concepts (because they fire every time that concept appears in the output), <b>we can then trace these and find how they combine internally to create other concepts</b>, as seen below, where the neurons experts on sport, basketball, and Michael Jordan combine to help the model predict ‘basketball’.</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/6b127f06-cf51-4a17-9751-ba27b47874bd/image.png?t=1778616367"/></div><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">This led to famous examples like the Golden Bridge LLM; once Anthropic found one of its models had a Golden Gate Bridge concept mapped to certain neurons, <b>they clamped the values of these neurons, and the model essentially became “the embodiment” of the actual bridge</b>. Basically, it couldn’t stop talking about it.</p><p class="paragraph" style="text-align:left;">Or if they clamped the sycophantic praise feature, the model became excessively sycophantic:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/41e6ec26-bbb5-4843-9585-d3b7f5b37928/image.png?t=1778573419"/><div class="image__source"><span class="image__source_text"><p>Source: Anthropic</p></span></div></div><p class="paragraph" style="text-align:left;">This was great progress, <b>but it didn’t allow us to actually see what the model was thinking</b>. Put another way, while we can map certain parts of the human brain to certain behaviors or movements, we can’t read human minds.</p><p class="paragraph" style="text-align:left;"><i>And how do we visualize thoughts?</i> Enter Anthropic’s new proposal, the <b>NLAs</b>.</p><h2 class="heading" style="text-align:left;">Verbalizing the inner monologues</h2><p class="paragraph" style="text-align:left;"><b>Natural Language Autoencoders</b> represent a totally new approach to AI interpretability that aims to verbalize (turn into words) the model’s internal activations. </p><div class="blockquote"><blockquote class="blockquote__quote"></blockquote></div><p class="paragraph" style="text-align:left;">Simply put, NLAs analyze how models behave internally and describe those behaviors in text. <i>The hope?</i> <b>That we can “read into the minds” of AI models</b>, know what they are thinking, and hopefully steer them to our liking.</p><p class="paragraph" style="text-align:left;">This has led to some of the most surprising, exciting, and, in some cases, disturbing discoveries in AI in a long time, like the example below, where, even if the model wasn’t explicitly saying so, it was internally suspecting it was being evaluated:</p><div class="image"><img alt="" class="image__image" style="" src="https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/7f751841-b723-499a-95f9-8b38129280d1/image.png?t=1778574647"/></div><p class="paragraph" style="text-align:left;"><i>But how on Earth did Anthropic build this magical ‘AI mind reader’?</i></p><p class="paragraph" style="text-align:left;">Behind the paywall, we go down the rabbit hole of explaining the internal functioning of AI models in first principles, explaining reward hacking and agentic misbehaviors like scheming or blackmailing (and importantly why they actually occur), the difference between verbalized thoughts and a model’s internal thoughts and beliefs system, the architecture and principles underpinning this new model class called NLAs, what NLAs unlock, and implications beyond research.</p></div><div class="paywall"><hr class="paywall__break"/><div class="paywall__content"><h2 class="paywall__header"> Subscribe to Full Premium package to read the rest. </h2><p class="paywall__description"> Become a paying subscriber of Full Premium package to get access to this post and other subscriber-only content. </p><p class="paywall__links"><a class="paywall__upgrade_link" href="https://thewhitebox.beehiiv.com/upgrade?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-finally-uncovered-what-ais-really-think">Upgrade</a> Translation missing: en.app.shared.conjuction.or <a class="paywall__login_link" href="https://thewhitebox.beehiiv.com/login?utm_source=thewhitebox.beehiiv.com&utm_medium=newsletter&utm_campaign=we-finally-uncovered-what-ais-really-think">Sign In</a></p><div class="paywall__upsell"><div class="paywall__upsell_header"><h3> A subscription gets you </h3></div><ul class="paywall__upsell_features"><li class="paywall__upsell_feature"> NO ADS </li><li class="paywall__upsell_feature"> An additional insights email on Tuesdays </li><li class="paywall__upsell_feature"> Gain access to TheWhiteBox&#39;s knowledge base to access four times more content than the free version on markets, cutting-edge research, company deep dives, AI engineering tips, & more </li></ul></div></div></div></div><div class='beehiiv__footer'><br class='beehiiv__footer__break'><hr class='beehiiv__footer__line'><a target="_blank" class="beehiiv__footer_link" style="text-align: center;" href="https://www.beehiiv.com/?utm_campaign=ac268084-9098-4115-8f12-390606905876&utm_medium=post_rss&utm_source=thewhitebox_by_nacho_de_gregorio">Powered by beehiiv</a></div></div>
  ]]></content:encoded>
</item>

  </channel>
</rss>
