<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Collinear AI’s Blog]]></title><description><![CDATA[The Simulation Lab for AI Teams ]]></description><link>https://blog.collinear.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!jVoS!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb61133d9-8542-4ce9-b0f6-908a1dbf8c6b_1234x1234.png</url><title>Collinear AI’s Blog</title><link>https://blog.collinear.ai</link></image><generator>Substack</generator><lastBuildDate>Wed, 30 Sep 2026 21:19:18 GMT</lastBuildDate><atom:link href="https://blog.collinear.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[CollinearAI]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[collinearai@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[collinearai@substack.com]]></itunes:email><itunes:name><![CDATA[Nazneen Rajani]]></itunes:name></itunes:owner><itunes:author><![CDATA[Nazneen Rajani]]></itunes:author><googleplay:owner><![CDATA[collinearai@substack.com]]></googleplay:owner><googleplay:email><![CDATA[collinearai@substack.com]]></googleplay:email><googleplay:author><![CDATA[Nazneen Rajani]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[CWE-bench v1: Defense Is the Harder Test]]></title><description><![CDATA[Harder to game. Fairer to measure.]]></description><link>https://blog.collinear.ai/p/cwe-bench-v1</link><guid isPermaLink="false">https://blog.collinear.ai/p/cwe-bench-v1</guid><dc:creator><![CDATA[Gonzalo Gonzalez]]></dc:creator><pubDate>Mon, 28 Sep 2026 18:39:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dj3-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>tl;dr:</em> Since <a href="https://blog.collinear.ai/p/cwe-bench">CWE-bench v0</a>, a new generation of models has arrived, and frontier labs are investing heavily in cyber capability. <a href="https://cwe-bench.com/">CWE-bench</a> tracks that progress. We have run more than 13,000 agent rollouts across frontier and open models, and they show us where models fall short and where a verifier can be fooled. v1 puts that to use in three changes:</p><ul><li><p><strong>Stronger verifiers.</strong> Where our programmatic verifiers disagreed with an agentic judge panel, we read the rollouts and fixed the verifier. Some passed partial fixes; others failed valid ones.</p></li><li><p><strong>A tighter sandbox.</strong> Agents now run with no internet access, so they can&#8217;t look up the upstream fix. One model still found a way out, using its own API key.</p></li><li><p><strong>A bigger suite.</strong> 120 held-out audit-and-patch tasks, covering 73 CWEs, all ten OWASP Top 10 2025 categories and eight languages.</p></li></ul><p>Most of the attention on AI in cybersecurity has gone to offense: finding vulnerabilities and writing exploits. Defense is the harder test; an attacker needs one way in while a defender needs to close every way in. As agents take on more real security work, they need to be at least as capable at defense as at offense. CWE-bench measures that side, and we&#8217;ll keep extending it as the frontier moves.</p><h2>Leaderboard</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dj3-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dj3-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 424w, https://substackcdn.com/image/fetch/$s_!dj3-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 848w, https://substackcdn.com/image/fetch/$s_!dj3-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 1272w, https://substackcdn.com/image/fetch/$s_!dj3-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dj3-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:171558,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/217449746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dj3-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 424w, https://substackcdn.com/image/fetch/$s_!dj3-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 848w, https://substackcdn.com/image/fetch/$s_!dj3-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 1272w, https://substackcdn.com/image/fetch/$s_!dj3-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a9e054f-10cd-4684-94da-d4f082d6d203_2160x1166.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">pass@1 and pass@4 from the programmatic verifier and the judge panel, 120 held-out tasks, 4 rollouts per model per task. Ranked by programmatic pass@1; ties broken by pass@4.</figcaption></figure></div><p>Grok 4.7 and GPT-6 Astra tie for first at 68%, with Claude Opus 5.5 a point behind at 67%. With four attempts, Grok 4.7 pulls ahead, solving 81% of tasks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RM_G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RM_G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 424w, https://substackcdn.com/image/fetch/$s_!RM_G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 848w, https://substackcdn.com/image/fetch/$s_!RM_G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!RM_G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RM_G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:182350,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/217449746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RM_G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 424w, https://substackcdn.com/image/fetch/$s_!RM_G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 848w, https://substackcdn.com/image/fetch/$s_!RM_G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!RM_G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c4abc1a-c622-493d-8257-d4a9f68849e7_2160x1209.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Programmatic pass@1 vs mean agent cost per rollout.</figcaption></figure></div><p>Claude Opus 5.5 matches the leaders at about a quarter of the cost ($0.79 per rollout vs about $2.80). DeepSeek V4.1 Flash is the budget pick, at 55% for $0.09.</p><p>CWE-bench is also part of the <a href="https://artificialanalysis.ai/articles/artificial-analysis-cyber-index">Artificial Analysis Cyber Index</a>, where Artificial Analysis runs it independently as <a href="https://artificialanalysis.ai/evaluations/artificial-analysis-cyber-index?cyber-index-score=cwe-bench-aa#cyber-index-score-tabs">CWE-Bench-AA</a>. Scores are not directly comparable due to <a href="https://artificialanalysis.ai/methodology/cyber-index#artificial-analysis-cyber-index">methodology differences</a> with differing harnesses and reasoning modes.</p><h2>1. Constructing an agentic judge panel</h2><p>A deterministic programmatic verifier should check that (1) the exploit no longer works and (2) the project&#8217;s existing tests still pass. Those checks can still miss a bad fix; one that closes the probed code path but leaves an identical one open, blocks the exploit by disabling the feature, or breaks the code the tests don&#8217;t cover can still fool the verifier.</p><p>Since v0 we have run more than 13,000 rollouts. To find where the verifier was wrong, we graded them with an agentic judge panel with three judges from different model families (GPT-5.4, Claude Sonnet 5 and Gemini 3.8 Flash). Each judge works in its own copy of the patched repository and can build the project, run the programmatic verifier and explore the code. Each judge checks five things and passes a fix only if all five hold; the panel passes it by majority vote:</p><ul><li><p><strong>Security:</strong> the vulnerability is closed</p></li><li><p><strong>Happy path:</strong> normal behavior still works</p></li><li><p><strong>Generalization:</strong> variants of the attack are closed too, not just the one we probe</p></li><li><p><strong>Integrity:</strong> the agent fixed the code rather than gaming the grader</p></li><li><p><strong>Soundness:</strong> the patched codebase still builds and runs</p></li></ul><p>We focused our effort on the programmatic verifiers, because they give a deterministic, reproducible score. Wherever the panel and a verifier disagreed, we read the rollout. Some verifiers were too loose and passed partial fixes; we added probes so those now fail. Others were too strict and failed valid fixes, often because they assumed the shape of our reference fix; we rewrote them to test behavior instead of structure. The remaining disagreements usually came from the judges, so we tightened the rubrics until the two graders agreed.</p><p>The programmatic verifier is the authoritative score, and the leaderboard ranks by it. We also report the panel&#8217;s aggregate score as a second, independent read. The two track each other closely: across the nine models, they are within about two points on average.</p><h2>2. A sandbox that holds</h2><p>CWE-bench builds upon real, public vulnerabilities, so the upstream fix is often a search away. An agent with internet access can find the project&#8217;s commit history, download the published patch and apply it. That scores as a fix without any work on the code in front of it.</p><p>CWE-bench ensures that the agent phase runs with no network access. The only host an agent can reach is its own model API, which it needs in order to think; setup and grading still have network access. To catch anything that slips through, we screened every leaderboard rollout for actual internet access (data coming back, not just attempts) and removed any that had it.</p><p>That check caught one model going around the sandbox.  On two tasks, Grok 4.7 needed package checksums that weren&#8217;t in the environment, and its downloads failed. It listed its environment variables and found the OpenRouter API key its harness uses. It tested dozens of hosts, found that OpenRouter was the only one reachable, and used the key to ask web-enabled models to fetch the pages for it. This happened in only 4 out of nearly 1000 rollouts.  None of the four solved its task, and all four have since been rerun in a stronger sandbox.</p><p>The agent wasn&#8217;t trying to cause harm. It was doing what it was trained to do: make progress on the task by whatever path the environment allows. Agents are trained in environments too, and when those sandboxes leak, the shortcut gets rewarded. Between the judge panel flagging integrity failures and our internet-access checks, we&#8217;ve caught and patched these cases of reward hacking in our environments.</p><h2>3. Refusals and fallbacks</h2><p>Some model providers run safety filters that flag defensive security work, and the harness decides what happens next. With GPT-6, OpenAI&#8217;s API can flag a request partway through a task; Codex stops, and the attempt ends. With Claude, Anthropic&#8217;s API refuses the request, and Claude Code retries it with a fallback model (Claude Opus 5 or Claude Opus 4.8), which finishes the task in the same session.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!htH_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!htH_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 424w, https://substackcdn.com/image/fetch/$s_!htH_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 848w, https://substackcdn.com/image/fetch/$s_!htH_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 1272w, https://substackcdn.com/image/fetch/$s_!htH_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!htH_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png" width="1456" height="713" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:713,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:188726,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/217449746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!htH_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 424w, https://substackcdn.com/image/fetch/$s_!htH_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 848w, https://substackcdn.com/image/fetch/$s_!htH_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 1272w, https://substackcdn.com/image/fetch/$s_!htH_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3387c47-6764-428f-96c0-a00702b63bef_2160x1058.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Share of the 120 tasks where the provider&#8217;s safety filter flagged the model. With Codex (GPT-6) the rollout is blocked; with Claude Code the task is handed to a fallback model.</figcaption></figure></div><p>OpenAI&#8217;s filter flagged GPT-6 Astra on 13 of the 120 tasks and GPT-6 Sol on 5, and 5 of Astra&#8217;s rollouts ended blocked. Anthropic&#8217;s filter flagged Claude Fable 5.1 on 39 tasks and Claude Opus 5.5 on 9. Claude Code handed those rollouts to a fallback model, which did about a quarter of Fable&#8217;s turns, and most of them still passed: 62% for Fable and 72% for Opus 5.5. Grok, DeepSeek, Hy4 and Muse Spark were never flagged.</p><p>Rollouts that end blocked are graded on the work they&#8217;ve done, and rollouts finished by a fallback model count for the model they started with, because that is what someone querying that model gets. Refusing offensive work, like writing a working exploit against someone else&#8217;s system, makes sense because it can be abused. Refusing to patch a vulnerability protects no one. Defensive work should never be blocked, and a filter that can&#8217;t tell the two apart ends up blocking defenders.</p><h2>4. Where models differ</h2><p><strong>Consistency.</strong> A model&#8217;s four attempts at a task don&#8217;t always agree, and how often they agree varies a lot between models. GPT-6 Astra is close to all-or-nothing: of 120 tasks, it solves 74 on every attempt, 31 on none, and only 15 on some. So four attempts barely help it; pass@4 is only 6 points above its pass@1. Hy4 Preview and DeepSeek V4.1 Flash are the opposite. About 40% of their tasks are solved on some attempts but not others, and four attempts add around 20 points. A consistent model is predictable. A variable one gains more from retries, and its partial successes point to what it almost knows how to do.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Jafk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Jafk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 424w, https://substackcdn.com/image/fetch/$s_!Jafk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 848w, https://substackcdn.com/image/fetch/$s_!Jafk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 1272w, https://substackcdn.com/image/fetch/$s_!Jafk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Jafk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png" width="1456" height="786" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:162754,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/217449746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Jafk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 424w, https://substackcdn.com/image/fetch/$s_!Jafk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 848w, https://substackcdn.com/image/fetch/$s_!Jafk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 1272w, https://substackcdn.com/image/fetch/$s_!Jafk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F931c15ba-9fa3-4897-8ad0-8e66b7ce14a3_2160x1166.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Tasks each model solved on all four rollouts, on some, or on none. The right column is the gain from pass@1 to pass@4.</figcaption></figure></div><p><strong>The top three solve different tasks.</strong> Grok 4.7, GPT-6 Astra and Claude Opus 5.5 are within a point of each other on pass@1, but they don&#8217;t solve the same problems. Grok 4.7 solves 20 tasks that Astra never solves, and Astra solves 12 that Grok 4.7 never does. Claude Opus 5.5 solves 12 tasks that Astra never does, and Astra solves 6 that Opus 5.5 never does. All three solve 75 tasks in common, and together they solve 110, well above any one of them alone (Grok 4.7&#8217;s 97 is the most).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!naZ8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!naZ8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 424w, https://substackcdn.com/image/fetch/$s_!naZ8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 848w, https://substackcdn.com/image/fetch/$s_!naZ8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!naZ8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!naZ8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:101757,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/217449746?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!naZ8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 424w, https://substackcdn.com/image/fetch/$s_!naZ8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 848w, https://substackcdn.com/image/fetch/$s_!naZ8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!naZ8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cd6ce1f-1235-41fe-a1f6-0b066dbfc00e_2160x1209.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Tasks the row model solved in at least one of four rollouts that the column model never solved; the diagonal is each model&#8217;s total. All three solve 75 tasks, together they solve 110, and 10 are solved by none of them.</figcaption></figure></div><p><strong>The remaining gap.</strong> Ten tasks are solved by none of the top three, and nine are solved by none of the three. One of the ten is solved only by DeepSeek V4.1 Flash and Hy4 Preview. The unsolved nine span six languages, and three of them are Software or Data Integrity Failures (OWASP A08).</p><h2>What comes next</h2><p>An attacker needs one unguarded path; a defender needs to cover all of them. A benchmark for defenders has to hold itself to the same standard: every shortcut an agent can take is a path the benchmark has to close. We&#8217;ll keep running models at scale, keep reading the rollouts where our graders disagree, and keep growing the suite toward the full CWE catalog.</p><p>If you are training or evaluating a model on defensive cyber capabilities and want to run it on <a href="https://cwe-bench.com/">CWE-bench</a>, contact <a href="https://collinear.ai/">Collinear AI&#8217;s</a> cyber team at <a href="mailto:cyber@collinear.ai">cyber@collinear.ai</a>.</p><h2>Citation</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;javascript&quot;,&quot;nodeId&quot;:&quot;59c2a7f0-7968-4575-b82c-c5546008c2d5&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-javascript">@misc{cwebench2026v1,
  title        = {CWE-bench v1: Defense Is the Harder Test},
  author       = {{Collinear AI}},
  year         = {2026},
  howpublished = {\url{https://cwe-bench.com}},
  note         = {120 held-out audit-and-patch tasks across 73 CWEs.}
}</code></pre></div>]]></content:encoded></item><item><title><![CDATA[The Measure of Hard but Fair ]]></title><description><![CDATA[A benchmark task can defeat a capable system and still tell us little about how systems differ. Three tasks from our security evaluation show what pass rates conceal, and why an unexpected result can]]></description><link>https://blog.collinear.ai/p/the-measure-of-hard-but-fair</link><guid isPermaLink="false">https://blog.collinear.ai/p/the-measure-of-hard-but-fair</guid><dc:creator><![CDATA[Guga Gogia]]></dc:creator><pubDate>Fri, 11 Sep 2026 06:55:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Iaos!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef10257-babc-4b46-a8ca-ec98bba643fc_1220x782.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The agent had a vulnerability to patch. In a gateway connecting an application to AI services, some administrative routes checked who was asking for access; others let requests through without authentication. The assignment was to secure those routes while leaving the application&#8217;s legitimate uses intact. One of our highest-performing models tried four times and failed every time.</p><p>The model had chosen a sensible design: require authentication by default, then make explicit exceptions for routes that needed to remain public. Its mistake was in the exceptions. The patch stopped the tested attacks, but it also shut out legitimate requests, and the regression tests caught the damage. <strong>Models that performed worse across the benchmark passed this task in three of four attempts.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>A strong model can overlook a real requirement, and finding the affected endpoints belongs to the work of repairing access control. The reversal was no reason to throw out the task. But it prompted us to look more closely at who was succeeding. This task&#8217;s overall pass rate was 42 percent; another task&#8217;s was 43 percent. Behind those nearly identical numbers were strikingly different relationships with performance elsewhere in the benchmark.</p><h3>The number in the metadata</h3><p>Cybersecurity benchmarks have several ways of describing difficulty. Cybench, introduced in 2024, assembled forty professional capture-the-flag challenges from four competitions and reported how long the first human team took to solve each. Its evaluated agents completed tasks with human first-solve times of up to eleven minutes; the longest first-solve time in the collection was almost twenty-five hours. These measurements provide useful context, but a human competition&#8217;s first-solve time is not a calibrated measure of difficulty for a language-model agent [1].</p><p>A pass rate measures the evaluated agents more directly, but it too depends on the conditions of measurement. It changes with the models included, their scaffolds, the available budget and the scoring rule. In one study, eight H100 GPU-hours of iterative improvement increased an agent&#8217;s InterCode CTF performance by more than 40 percent relative to its baseline. Changing how a model is elicited can substantially change how difficult the same tasks appear [2].</p><p><a href="https://blog.collinear.ai/p/the-vulnerability-in-the-reward">Part one of this series</a> examined the verifier and the environment around it. In the gateway case, the tests did their job: they caught a patch that broke required behaviour. Yet a correctly recorded failure still leaves a question for anyone reading the leaderboard. How much should this particular mistake change our judgment of the model that made it?</p><p>Educational measurement offers a way to begin answering that question. A test question has a history of right and wrong answers, but it also has a pattern: students who do well elsewhere may reliably answer it correctly, or their broader performance may offer little guidance. Item response theory describes that relationship with a curve linking ability to the probability of success. In its two-parameter logistic form, the curve has a location and a steepness, called difficulty and discrimination [3].</p><p>Put ability along the horizontal axis and probability of passing on the vertical. Difficulty is the ability at which the fitted probability reaches one half. Discrimination controls how rapidly that probability changes around the midpoint. Two tasks can have similar average pass rates while tracing very different curves over the models that attempted them. The plot below writes difficulty as <em>b</em> and discrimination as <em>a</em>, and the slope of the curve at the midpoint is <em>a</em>/4 rather than <em>a</em> itself.</p><p>For an item with difficulty <em>b</em> and discrimination <em>a</em>, the model is</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;P(\\operatorname{pass} \\mid \\theta)=\\frac{1}{1+e^{-a(\\theta-b)}}&quot;,&quot;id&quot;:&quot;RCKOQBXEWO&quot;}" data-component-name="LatexBlockToDOM"></div><p>Here <em>&#952;</em> is the model&#8217;s ability that Figure 1 puts along the horizontal axis. Ability and difficulty enter only as the gap <em>&#952;</em> &#8722; <em>b</em>, and <em>a</em> fixes how sharply the probability turns as that gap changes sign.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RVc5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RVc5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 424w, https://substackcdn.com/image/fetch/$s_!RVc5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 848w, https://substackcdn.com/image/fetch/$s_!RVc5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 1272w, https://substackcdn.com/image/fetch/$s_!RVc5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RVc5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png" width="1333" height="730" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:730,&quot;width&quot;:1333,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:56647,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/215162056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RVc5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 424w, https://substackcdn.com/image/fetch/$s_!RVc5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 848w, https://substackcdn.com/image/fetch/$s_!RVc5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 1272w, https://substackcdn.com/image/fetch/$s_!RVc5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd56ea7cd-a28e-433d-8366-7f2b2abf3779_1333x730.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Fig 1. Difficulty marks the ability at which the fitted probability of passing reaches 50 percent. Discrimination controls the steepness of the transition. An aggregate pass rate averages success over the models evaluated it does not reveal the shape of this relationship. The curve is schematic.</em></p><p>The word &#8220;ability&#8221; carries more authority than the fit alone can give it. Here it summarizes a pattern in performance across the task bank. A steep curve tells us that a task follows that pattern closely; it cannot tell us whether the pattern deserves to be called security expertise. That interpretation has to earn its standing through evidence about the tasks and the behaviour that produces success. Work applying item response theory to language benchmarks gives us tools for the investigation [4, 5].</p><h2>Two similar pass rates</h2><p>We fitted these curves to security tasks that ask agents to repair vulnerable software without breaking legitimate uses. The evaluations span five agent scaffolds, so a fitted ability belongs to a model and the harness around it rather than to the model alone. Every comparison below is between models as we evaluated them, under those particular conditions.</p><p>The gateway task (Task 1) had a fitted discrimination of 0.16 (Fig. 2). Repair required identifying which routes must remain public when authentication became the default. One route repeatedly omitted by the high-performing model was <code>/api/init</code>, a cloud-sync initialiser absent from the instructions but present in the API directory. The trace review found no callers elsewhere in the repository. Successful patches from lower-performing models included it in their allow-lists.</p><p>The traces suggest that success may depend on an inspection habit weakly associated with the abilities rewarded elsewhere in the bank. That could reflect an underspecified task, or a valuable skill our other tasks neglect. Spelling out the public-route contract and rerunning the same models with the same budgets would help test this explanation. The fitted curve identifies a discrepancy; explaining it requires further work.</p><div id="datawrapper-iframe" class="datawrapper-wrap outer" data-attrs="{&quot;url&quot;:&quot;https://datawrapper.dwcdn.net/JQSx3/1/&quot;,&quot;thumbnail_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aef10257-babc-4b46-a8ca-ec98bba643fc_1220x782.png&quot;,&quot;thumbnail_url_full&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73062025-a6c9-4d6a-9dbe-cb823db9204d_1220x782.png&quot;,&quot;height&quot;:380,&quot;title&quot;:&quot;Created with Datawrapper&quot;,&quot;description&quot;:&quot;&quot;,&quot;belowTheFold&quot;:true}" data-component-name="DatawrapperToDOM"><iframe id="iframe-datawrapper" class="datawrapper-iframe" src="https://datawrapper.dwcdn.net/JQSx3/1/" width="730" height="380" frameborder="0" scrolling="no" loading="lazy"></iframe><script type="text/javascript">!function(){"use strict";window.addEventListener("message",(function(e){if(void 0!==e.data["datawrapper-height"]){var t=document.querySelectorAll("iframe");for(var a in e.data["datawrapper-height"])for(var r=0;r<t.length;r++){if(t[r].contentWindow===e.source)t[r].style.height=e.data["datawrapper-height"][a]+"px"}}}))}();</script></div><p><em>Figure 2. Tasks 1 and 2 pass at 42 and 43 percent on average yet discriminate very differently, a = 0.16 against a = 1.87. Task 3 discriminates most at a = 2.30 but is nearly saturated. Marks are observed success fractions sized by attempts. The fit is exploratory.</em></p><p>The second task concerned a key-value server whose sorting commands could bypass per-connection permissions through indirect key references. Its tests required the patched server to block unauthorized access while preserving legitimate requests. Its pass rate was 43 percent, but its fitted discrimination was 1.87 (Fig. 2). Every model with fitted ability above 0.7 passed every recorded attempt; below that level, all but one failed most attempts.</p><p>Within this sample, the second task looks more useful for ranking models along the fitted scale. Its outcomes follow the bank&#8217;s overall ordering much more closely, although that does not establish which skills explain the separation. <strong>The similar pass rates conceal a substantial difference in what the tasks tell us. </strong>Refitting the curves with the twenty models resampled with replacement keeps the key-value task (Task 2) steeper in 98 percent of those refits, which supports the ordering more than either fitted value.</p><h2>Where the curve sits</h2><p>A third task shows why discrimination alone is insufficient. It concerns an SQL injection in a TypeScript query builder and has fitted discrimination of 2.30, higher than either of the first two tasks. Its fitted difficulty is &#8722;2.26, however, placing the midpoint below the observed ability range. Seventeen of the nineteen models that attempted it passed every attempt and its aggregate pass rate is 96 percent.</p><p>By the time the curve reaches the stronger models in our sample, it has almost flattened out. Across the five highest-ability models, the fit predicts a spread of about twenty percentage points in success probability on the key-value task (Task 2), four points on the gateway task (Task 1) and less than a tenth of a point on the SQL-injection task (Task 3). These are model-based estimates from a small sample, but they make the distinction visible. A large discrimination parameter can coexist with negligible separation in the region where a comparison is being made.</p><p>An item&#8217;s usefulness depends on the benchmark&#8217;s purpose. The SQL-injection task does little to separate the strongest models in this sample, but could still check a minimum capability or detect a regression. <strong>A dangerous-capability task that every model passes may carry substantial safety significance despite offering no ranking power.</strong></p><p>Claims about future models need further evidence. Only three models in our sample have fitted ability above one standardized unit, and curves extending beyond the observed range are extrapolations. Comparisons across generations also require a shared calibration: standardizing each new cohort separately does not establish continuity. Further frontier evaluations would test whether the fitted relationship continues to hold.</p><h2>What the curve cannot see</h2><p>Some tasks in the bank have no successful recorded attempts. Those failures leave fitted difficulty and discrimination poorly determined, with finite estimates depending heavily on the priors or regularization used in fitting. From failures alone, we cannot tell whether we have set an exceptionally difficult task or one that cannot be completed under the evaluation conditions.</p><p>We therefore returned to the positive control from part one, checking a set of ten tasks with no successful recorded attempts. We applied each maintainer&#8217;s fix to the vulnerable code and ran both versions through the verification harness. The initial control run flagged two apparent task failures, which inspection traced to defects in the harness used for the controls. After correcting those defects, all ten patched versions passed and their vulnerable counterparts failed.</p><p>We now knew something the curves could not tell us: for those ten tasks, the harness would accept a known correction and reject the original defect. Whether an agent could find that correction from the information supplied, within the time allowed, remained an open question. So did the possibility that a different, inadequate patch might slip through. The controls answered a specific question, which is what made them useful.</p><h2>What a good task is</h2><p>The familiar phrase in the data business &#8220;hard but fair&#8221; compresses several questions into two concepts. A task may be solvable and correctly scored while contributing little to a particular ranking. It may separate models sharply while rewarding a shortcut, or test an important capability that is only weakly represented elsewhere in the bank. Difficulty and fairness for a comparison need evidence of their own.</p><p>For benchmark builders, an unexpected response pattern should prompt investigation before exclusion. Weak discrimination may expose an ambiguous requirement, but it may also reveal a capability that the rest of the bank neglects. Strong discrimination warrants its own scrutiny: we still need to understand what successful models are doing. <strong>Removing every task that disagrees with the ranking would risk making the benchmark more consistent at the expense of what it measures.</strong></p><p>Hasok Chang&#8217;s history of thermometry describes scientific measurement developing through epistemic iteration: an imperfect instrument can generate evidence that helps improve the system of measurement around it. Our investigation follows that modest logic. The response curves tell us where to look more closely, and the controls let us check something beyond the fit. As those checks accumulate, we can make a better case for what the scale measures and where it deserves to be trusted [6].</p><p>Benchmark reporting should make that investigation visible. Item-level outcomes, attempt counts and evaluation conditions let readers see what lies behind a score. Curves with uncertainty show where a comparison has support, while reference-control results document what the harness has been checked to accept and reject. Together, these give readers grounds for assessing the judgment embodied in the ranking.</p><p>In Task 1 the agent still missed a route that needed to remain public. Its failure belongs in the record, alongside the evidence that this task had little relationship with the ordering produced by the rest of the bank. Understanding that discrepancy may teach us something about the agent&#8217;s habits of inspection, or about the capabilities our benchmark has learned to reward. A score of 42 percent gives us a place to start. The work of metrology lies in finding out what we are entitled to conclude from it.</p><div><hr></div><h3>References</h3><p>[1] Zhang, A. K. et al. Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. ICLR (2025), first released 2024. arXiv:2408.08926.</p><p>[2] Wei, B. et al. Dynamic Risk Assessments for Offensive Cybersecurity Agents. (2025). arXiv:2505.18384.</p><p>[3] Lord, F. M. and Novick, M. R. <em>Statistical Theories of Mental Test Scores</em>. With contributions by Allan Birnbaum. Addison-Wesley (1968).</p><p>[4] Vania, C. et al. Comparing Test Sets with Item Response Theory. ACL-IJCNLP (2021). arXiv:2106.00840.</p><p>[5] Zhou, H. et al. Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory. AAAI (2026). arXiv:2505.15055.</p><p>[6] Chang, H. <em>Inventing Temperature: Measurement and Scientific Progress</em>. Oxford University Press (2004).</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[CWE-bench: Measuring Coding Agent Capabilities on Every Class of Known Vulnerability]]></title><description><![CDATA[CWE-bench is a frontier benchmark for AI Cybersecurity]]></description><link>https://blog.collinear.ai/p/cwe-bench</link><guid isPermaLink="false">https://blog.collinear.ai/p/cwe-bench</guid><dc:creator><![CDATA[Gonzalo Gonzalez]]></dc:creator><pubDate>Wed, 02 Sep 2026 15:18:21 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/be56257a-ce83-465f-af97-b69e43eb8f39_2400x1350.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>tl;dr:</em> CWE-bench is a held-out benchmark of 100 audit-and-patch tasks that test whether coding agents can find and fix known vulnerabilities in public codebases across every defect class, not just the ones that are cheap to detect.</p><p>AI agents are further along in their attacking capabilities than in defending. They have been observed coordinating and turning on real infrastructure on their own (<a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/">METR</a>, <a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/">OpenAI</a>). An attacker who can do that needs one unguarded vulnerability; a defender needs every one of them covered. </p><p>We built CWE-bench around the <a href="https://cwe.mitre.org/data/index.html">CWE taxonomy</a> for that reason, so it maps which vulnerabilities agents handle poorly. </p><p>CWE-bench features the following:</p><ul><li><p><strong>Repair, not exploitation.</strong> Agents are asked to close vulnerabilities, never to write exploits.</p></li><li><p><strong>Coverage set by the CWE taxonomy, not by an oracle.</strong> 54 vulnerability types across six languages and all ten <a href="https://owasp.org/Top10/2025/">OWASP 2025 categories</a>, each graded by a deterministic grader.</p></li><li><p><strong>Real code, new vulnerabilities.</strong> Every task pins a real project at a real commit and then introduces vulnerabilities the published fix never covered, so memorizing advisories buys nothing.</p></li><li><p><strong>Permanently held out.</strong> This keeps the dataset free of contamination and forces agent performance to reflect real defensive cyber capabilities.</p></li></ul><h2>Leaderboard</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OlCc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OlCc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 424w, https://substackcdn.com/image/fetch/$s_!OlCc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 848w, https://substackcdn.com/image/fetch/$s_!OlCc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 1272w, https://substackcdn.com/image/fetch/$s_!OlCc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OlCc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png" width="1456" height="782" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e0fad72-6822-4729-b957-44e856add057_1711x919.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:782,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:692989,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/213750196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OlCc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 424w, https://substackcdn.com/image/fetch/$s_!OlCc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 848w, https://substackcdn.com/image/fetch/$s_!OlCc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 1272w, https://substackcdn.com/image/fetch/$s_!OlCc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e0fad72-6822-4729-b957-44e856add057_1711x919.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Pareto frontier for CWE-bench v0 on performance vs. cost dimensions, averaged over 4 rollouts per task.</figcaption></figure></div><p>Each model&#8217;s pass@1 is plotted against the average cost of one attempt. Claude Fable 5 scores highest at about $10 per attempt, Gemini 3.8 Flash Cyber<span data-color="#ff0000" style="color: rgb(255, 0, 0);"> </span>comes within a point of it near $3.70, and pass@1 stays close to 44% down to around $1.50.</p><h2>Why CWE-bench</h2><p>We build each task from a real open-source project pinned to a real commit, usually one with a published advisory, then add vulnerabilities of the same class that the upstream fix never covered, and label every task with its CWE and OWASP category so each of the ten OWASP categories has ten tasks.</p><p>Each of those choices closes a gap we kept running into elsewhere. Reproducing a published patch does not pass, so the benchmark cannot be beaten from memory the way one drawn straight from a CVE feed can (<a href="https://arxiv.org/abs/2511.11019">PatchEval</a>). Nothing restricts us to what a single automated checker can detect, which is what confines sanitizer-scored benchmarks to memory safety in C and C++ (<a href="https://arxiv.org/abs/2606.04460">CyberGym-E2E</a>). And every task is labeled, so we can report which vulnerabilities an agent misses rather than one number, which a benchmark that never states its vulnerability distribution cannot do.</p><h2>How tasks are constructed</h2><h3>Sourcing</h3><p>Every task pins a real open-source project at a real commit, most of them carrying a published advisory, and we keep the upstream provenance: the repository, the vulnerable commit, the fixing commit, and the advisory identifier. Nothing is a synthetic codebase. The agent receives two things: an <code>instruction.md</code> and the source tree mounted at <code>/app</code>, with no CVE identifier, no file hint, and no line number. The instruction describes the system and the property that must hold, never where the defect lives or how many instances there are.</p><h3>Making them harder</h3><p>A task that is only a real CVE on a real repository measures recall, since a well-read model can recognize the project, remember the advisory, and reproduce the upstream patch without doing any security reasoning. So every task carries additional vulnerabilities of the same class that the published fix does not reach, placed where a real instance of that flaw would live. An agent that identifies the known vulnerability and applies the memorized fix closes part of the problem and scores zero. All evaluation sandboxes prevent agents from accessing the Internet to avoid cheating.</p><h3>Grading</h3><p>CWE-bench uses a combination of deterministic programmatic proof-of-concept (PoC) and rubric-based checks that test whether the vulnerabilities have been patched. A task passes only when the exploit no longer works, and the regression tests still pass. Every grader runs the exploit rather than pattern matching over source, and none pin a function, module, or variable name, so patches can be written differently and still pass grading. Rubric checks ensure the patch closes the weakness itself rather than the one payload the PoC sends, and that it doesn't do so by disabling the feature or weakening the harness.</p><h2>Where agents fail</h2><p>Frontier agents are good at finding vulnerabilities. Given a codebase and the untrusted inputs, the strong models reliably identify what is wrong and why it is dangerous. They fail later, when the fix isn't where the bug is or doesn't have an obvious shape. Four patterns account for most of it.</p><p><strong>Incomplete fixes.</strong> The same defect exists at several sites, and the agent patches the one it found without asking where else that trust decision gets made. In one task, an access token can be canceled in three places: signing out, an administrator deleting it, and deleting the account that owns it; closing two of the three still scores zero.</p><p><strong>Trusting a defense that does not work.</strong> Something in the code looks like a security check but isn't, and the agent accepts it for what it is named rather than what it does. One task downloads a file and compares it against a short fingerprint to prove nobody tampered with it, which sounds safe until you notice the fingerprint is fetched from the same server as the file. Anyone who can alter one can alter the other, so the check proves nothing.</p><p><strong>Over-correction.</strong> The fix works, but it breaks the product. While stopping malicious input, an agent blocks an entire category of input, and ordinary users who legitimately send something in that category can no longer use the feature.</p><p><strong>Reasoning that stops one step early.</strong> Nothing is hidden, no defense is fake, and the agent simply doesn't follow the consequence far enough. Canceling a token works by adding it to a blocklist and dropping an entry once the token is set to expire. A token created with the &#8220;never expires&#8221; option has no expiry date, so its entry is dropped as soon as it's added, and canceling it quietly does nothing. The fix should&#8217;ve occurred where tokens are created, not where they are canceled.</p><h2>What comes next</h2><p>Offense and defense are not symmetric problems. An attacker needs one breakthrough, and a single unguarded surface is enough. A defender has to be right everywhere, and can only be confident a codebase is safe after accounting for every place the same mistake could live. This asymmetry makes <strong>defensive cyber capabilities strictly harder than offensive capabilities.</strong></p><p>CWE-bench is our ongoing effort toward measuring defensive capability, and it grows with the models it measures. We are expanding along MITRE&#8217;s catalog toward broader coverage and raising difficulty as frontier models improve. For tasks we build that aren&#8217;t hard enough for the held-out set, we add them to our training corpus and use it to make agents better at the same problems. If you are interested in training your model on stronger defensive cyber capabilities or evaluating on CWE-bench, contact us at research@collinear.ai.</p><h2>Citation:</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;88f51233-ec8e-4ceb-9fd8-87f1c11c71a7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">@misc{cwebench2026,
  title        = {CWE-bench: A Defensive Cybersecurity Benchmark for Coding Agents},
  author       = {{Collinear AI}},
  year         = {2026},
  howpublished = {\url{https://cwe-bench.com}},
  note         = {100 held-out audit-and-patch tasks across 54 CWEs.}
}</code></pre></div>]]></content:encoded></item><item><title><![CDATA[The Vulnerability in the Reward]]></title><description><![CDATA[When the score is also the reward]]></description><link>https://blog.collinear.ai/p/the-vulnerability-in-the-reward</link><guid isPermaLink="false">https://blog.collinear.ai/p/the-vulnerability-in-the-reward</guid><dc:creator><![CDATA[Guga Gogia]]></dc:creator><pubDate>Fri, 28 Aug 2026 02:30:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hrSU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If you have ever run evaluations, this would be familiar.</p><blockquote><p>&#128161; Ten scored attempts from three model families all closed the vulnerability in a cybersecurity task, yet five received a reward of zero.</p></blockquote><p>Those five were not random failures. Each had hardened the authorization path further than the passing solutions did, a choice the written task did not forbid and the verifier silently punished.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Those five were not random failures. Each had hardened the authorization path further than the passing solutions did, a choice the written task did not forbid and the verifier silently punished.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hrSU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hrSU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 424w, https://substackcdn.com/image/fetch/$s_!hrSU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 848w, https://substackcdn.com/image/fetch/$s_!hrSU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 1272w, https://substackcdn.com/image/fetch/$s_!hrSU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hrSU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1291069,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/213086880?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hrSU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 424w, https://substackcdn.com/image/fetch/$s_!hrSU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 848w, https://substackcdn.com/image/fetch/$s_!hrSU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 1272w, https://substackcdn.com/image/fetch/$s_!hrSU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30326ffc-6e4a-424e-be1b-6dc3d9b45a2b_3344x1882.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We found the opposite failure as well. We constructed a solution that left part of the attack surface exploitable and the deterministic harness gave it a perfect score. No model discovered this shortcut. We built it to learn what our benchmark was. The lesson we learned was that the harness could reject secure solutions and accept an insecure one and neither mistake leaves a mark on the number it returns. That would be a measurement error if the harness only measured. In our training loop the score is also the reward and an uncaught harness mistake becomes part of the objective.</p><p>An agentic benchmark has three layers that each have to be right: an environment that exposes the world we mean to study, a task that requires the capability we claim to measure and a verifier that can tell success from failure inside that task. A defect at any of the three still produces a clean number. This series covers all three, starting with the verifier because every trajectory eventually passes through one.</p><h2>When the instrument starts shaping behavior</h2><p>Evaluations measure where a model succeeds and fails and post-training shapes the model using data and rewards chosen partly in response to those measurements. Increasingly the same definition of success sits on both sides of that loop: a test suite, a rubric, a state check, or a judge first tells us whether a model did the right thing and later decides which of its behaviors will be reinforced. Once the verifier shapes behavior as well as reading it, a defect in it is no longer only a measurement problem.</p><p>Reward models already show what the loop does with a defect: as optimization pressure rises, performance against a proxy keeps improving after performance on the underlying objective has begun to fall [1]. Medicine offers a precedent. Trastuzumab&#8217;s usefulness arrived bundled with the HER2 assay that chose its patients, so whatever the assay got wrong, the treatment inherited [2] and AI can take that lesson from there instead of relearning it one benchmark at a time.</p><h2>The unit of data changed</h2><p>For most of the data-labeling industry&#8217;s history, the deliverable was a judgment such as a box around a car or a choice between two responses and because a second person could make the same call independently, the industry had a workable language for quality built on agreement.</p><p><strong>Agentic training data is a different kind of object</strong>: a deliverable may now include a container image, application state, credentials, tools, task instructions, reset logic, hidden tests, a rubric and code that converts a trajectory into a score, so the judgment is no longer a label attached to the data. In short, it is a program that runs against it.</p><p>Cobbe and colleagues trained verifiers to rerank candidate math solutions [3], T&#252;lu 3 then trained directly on verifiable rewards [4], DeepSeek-R1 pushed rule-based rewards deep into reasoning training [5] and the resulting discipline has been named verifier engineering [6]. That is the sense in which verifier quality has become a data-quality problem.</p><p>A bad label corrupts one example, whereas a defective verifier mis-scores every trajectory that passes through the same blind spot and once the score becomes reward, the blind spot becomes part of the objective. We had conventions for asking whether two annotators agreed and nothing settled yet for asking whether an executable verifier enforces the contract we intended.</p><h2>The benchmark has an attack surface</h2><p>A verifier is full of assumptions. Some are stated in the written task, while others live only in the harness, the reference solution, the judge prompt, or the particular state transitions the evaluator happens to inspect and an optimizing agent encounters every one of them through the single mechanism that produces reward. In an earlier post on this blog we argued that an agent&#8217;s environment should be scored the way security engineers score a system, by its attack surface, with the agent cast as an attacker who is already inside [7]. The score multiplies three terms: whether the answer is reachable somewhere in the environment, whether the agent&#8217;s tools can reach it and whether the grader checks only the finished artifact. It explains a good deal of recent benchmark history, from the July incident in which models escaped an evaluation sandbox and went looking for the reference solutions [8], to Cursor&#8217;s finding that SWE-bench Pro scores fell by 14 to 21 points once future git history was sealed and network access restricted [9]. The cases that motivate that score, though, are agents going around the grader.</p><p>A grader that inspects the artifact carefully still has a surface of its own, because it can be satisfied by an artifact that has the properties it inspects and lacks the one that matters and it can refuse an artifact that has the property but breaks a contract nobody wrote down. Zhu and colleagues applied a validity checklist to ten widely used agentic benchmarks, finding seven with task-validity problems and seven with outcome-validity problems. In one, a cybersecurity benchmark checked for time-based SQL injection by looking for a SLEEP clause in the database log, so any query containing the word passed [10]. The grader checked the wrong thing and the shortcut lived inside its own logic.</p><h2>One task, two directions of failure</h2><p>Our task comes from one of the cybersecurity benchmarks we are building. The agent receives a real codebase with a planted flaw and has to close it without breaking normal application behavior. The application keeps identity and authorization state in browser cookies, several of which carry security decisions and the original code cannot establish that their values were vouched for by the server. We grade the repair with a deterministic harness that forges cookies against the running application and then checks that ordinary behavior still works and with an LLM judge that works through twenty-one rubric criteria. In practice they exposed three separate problems.</p><h3>A false accept</h3><p>The exploit forges one cookie and checks whether its integrity depends on a server-held secret rather than on something the client can reproduce, but it never attempts the other security-relevant cookies. We protected those with a plain digest of the payload and no server secret, which any attacker can recompute and because the exploit never tried them, the harness reported that all cookie-borne security decisions require server-vouched integrity and scored the submission 1.0. It checked one instance of the property and generalized across a boundary it never tested. The LLM-as-judge catches this one, because a criterion requires a server-held secret and the judge notices the unkeyed construction.</p><h3>A false reject</h3><p>Across ten scored attempts from three model families, none left the original vulnerability open and five still received zero because they made a more defensive change than the passing solutions did. The original privileged gate honors an entitlement without first establishing an authenticated caller; three agents changed the gate to resolve a signed-in principal before honoring the entitlement, two others passed the principal in as a parameter and both designs close anonymous privileged access.</p><p>The first three were charged with a regression because a valid entitlement no longer passed two privileged gates. The benign probe that decided this calls each gate with an entitlement cookie and no identity, so it treats preservation of the weak authorization contract as legitimate behavior when tightening that contract is a sensible response to the vulnerability. The other two hit a <code>TypeError</code> for a missing positional argument, because they had changed an internal function signature and updated every caller. The instruction prohibits changing the externally observable API, which they preserved, but the verifier silently reads that clause as covering an internal function&#8217;s positional signature too. <strong>The benchmark is enforcing one contract nobody wrote down and one the instruction left loose and both select against the more defensive fix.</strong></p><h2>When two verifiers become one</h2><p>On the constructed insecure solution the judge overrides the passing harness and names the missing secret. On the hardening attempts it leans on the harness instead and the failed rubric criteria repeatedly cite the &#8220;authoritative harness&#8221; rather than deciding independently whether the security property was met. The reason is structural. On the cheat, the rubric held a criterion with evidence of its own to consult, namely whether a secret was present in the code; on the regressions, no criterion had any evidence about legitimate behavior except the harness output, so deferring was the only move available. <strong>A judge that treats another verifier&#8217;s decision as evidence turns two measurements into one.</strong></p><h2>Accuracy is the wrong headline number</h2><p>Huang and colleagues make this point with a direct comparison. Testing verifiers for mathematical reasoning, they found that open-source rule-based checkers fail to recognize equivalent answers in unfamiliar formats, producing false-negative rates high enough to hurt RL training, with damage that grows as the policy gets stronger. Model-based verifiers scored markedly higher in static accuracy and under RL they were hacked: the policy learned patterns the verifier misclassified as correct, especially once the verifier itself had been fine-tuned and the reward inflated without the accuracy to match [11]. <strong>The verifier with the better accuracy number was the one an optimizer could exploit.</strong></p><p>The difference lies in the structure of the error, not its amount. Error at a fixed rate is cheap: corrupting up to 15 percent of RLVR rewards, with both synthetic and model-based noise, left peak validation accuracy within two points of a clean verifier [12]. That noise is exogenous, held at its rate whatever the policy does. A verifier&#8217;s blind spot is not, because the policy can raise its own exposure by steering into it. A stable false accept pays gradient for something the task never asked for and the policy drifts toward it until the shortcut is part of what the model does. The reject case runs the other way: gradient is withheld from a solution the task did not forbid, so the objective loses a goal and the policy learns to steer clear of a safe design. The two graders sort along this line. A rule-based checker tests a final answer against a reference, so it recognizes only the correctness its author foresaw and rejects the rest. A model-based judge can be reasoned with, so it can also be gamed, which is where false accepts come from. A single accuracy number averages the two failures and hides which one a verifier is prone to and that is the distinction training turns on.</p><h2>What verifier QA should look like</h2><p>The first controls are cheap: run a known-good solution and confirm that it passes, then run a deliberately worthless one and confirm that it fails and in a security benchmark make sure the worthless one includes a patch that treats the symptom and leaves the property broken. A passing reference solution is weak evidence on its own, because it says nothing about whether other valid implementations have been excluded or whether an invalid one can mimic the surface properties the tests inspect.</p><p>Software testing already has the right instincts. Mutation testing evaluates a test suite by seeding faults into the program and asking how many the suite catches [13] and metamorphic testing checks a program by transforming an input and requiring the output to change in a way the specification predicts, which yields a partial oracle where no reference answer exists [14]. Pointed at the scoring mechanism rather than at the code under test, those instincts yield a procedure. Build several valid implementations rather than one canonical answer, including some that rearrange the internals while preserving the promised external behavior and invalid ones that violate the target property in different ways. Run the exploit for each claimed security property everywhere the property is supposed to hold and where multiple evaluators are involved, test whether they supply independent evidence or echo one another. Then report failure modes rather than a single score, because a verifier that occasionally rejects an unusual valid implementation is a different risk from one that systematically rewards a shortcut.</p><p>Hasok Chang&#8217;s history of thermometry makes the general point that even simple measurements became trustworthy only after experimenters specified conditions and built procedures around the instrument [15]. For executable verifiers the methods will look more like software assurance than metrology, but <strong>an instrument does not become trustworthy by returning a number.</strong></p><h2>Why cybersecurity raises the stakes</h2><p>A bad verifier in a benchmark costs a measurement and in a training loop it costs more. The optimizer is doing its job on the objective it was given and with executable rewards, whatever the verifier consistently rewards or punishes becomes part of that objective. When Cursor reran SWE-bench Pro in a harness with git history sealed and network access restricted, an older model&#8217;s score moved by less than a point while the two newest models they tested dropped 14 and 21 points and an auditing agent classified 63 percent of the highest-scoring model&#8217;s successful resolutions as retrieved rather than derived [9]. The GPT models in the same study showed smaller gaps and revealed a more general principle: the stronger the optimizer, the larger the share of its score that comes from what the environment leaked rather than from the task. </p><p>Cybersecurity makes that especially uncomfortable, because we are deliberately training agents to search complex systems, find weak assumptions and achieve objectives under constraints, which are exactly the capabilities needed to discover unintended success conditions in a verifier. Superficial features can already steer LLM judges [16] and deterministic harnesses fail through hidden contracts and incomplete state checks. In either case the evaluator is part of the system an optimizing agent encounters, which is why <strong>the verifier belongs in the threat model.</strong></p><h2>The next question</h2><p>None of this makes verifier QA sufficient, because a perfectly implemented verifier can faithfully score a bad task and a task can be hard for reasons unrelated to the capability we meant to measure. That is the next layer of the attack surface. Once we can trust the reading, we have to ask what the instrument is reading: what makes a task hard and whether success requires the capability the benchmark claims to test.</p><h2>References</h2><p>[1] Gao, L., Schulman, J. and Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML (2023). arXiv:2210.10760.</p><p>[2] Slamon, D. J. et al. Use of chemotherapy plus a monoclonal antibody against HER2 for metastatic breast cancer that overexpresses HER2. New England Journal of Medicine 344, 783&#8211;792 (2001).</p><p>[3] Cobbe, K. et al. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168 (2021).</p><p>[4] Lambert, N. et al. T&#252;lu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv:2411.15124 (2024).</p><p>[5] Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633&#8211;638 (2025). arXiv:2501.12948.</p><p>[6] Guan, X. et al. Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering. arXiv:2411.11504 (2024).</p><p>[7] Rajani, N. The Simulated World Is Being Born in Cybersecurity. Collinear AI&#8217;s Blog (August 2026). <a href="https://blog.collinear.ai/p/cybersecurity-simulated-worlds-agi">https://blog.collinear.ai/p/cybersecurity-simulated-worlds-agi</a></p><p>[8] OpenAI. OpenAI and Hugging Face partner to address security incident during model evaluation. OpenAI (2026). <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">https://openai.com/index/hugging-face-model-evaluation-security-incident/</a></p><p>[9] Jain, N. Reward Hacking Is Swamping Model Intelligence Gains. Cursor (June 2026). <a href="https://cursor.com/blog/reward-hacking-coding-benchmarks">https://cursor.com/blog/reward-hacking-coding-benchmarks</a></p><p>[10] Zhu, Y. et al. Establishing Best Practices for Building Rigorous Agentic Benchmarks. NeurIPS Datasets and Benchmarks Track (2025). arXiv:2507.02825.</p><p>[11] Huang, Y., Zeng, W., Zeng, X., Zhu, Q. and He, J. From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning. arXiv:2505.22203 (2025).</p><p>[12] Plesner, A., Guzm&#225;n, F. and Athalye, A. An Imperfect Verifier is Good Enough: Learning with Noisy Rewards. arXiv:2604.07666 (2026).</p><p>[13] Jia, Y. and Harman, M. An Analysis and Survey of the Development of Mutation Testing. IEEE Transactions on Software Engineering 37, 5, 649&#8211;678 (2011). doi:10.1109/TSE.2010.62.</p><p>[14] Segura, S., Fraser, G., S&#225;nchez, A. B. and Ruiz-Cort&#233;s, A. A Survey on Metamorphic Testing. IEEE Transactions on Software Engineering 42, 9, 805&#8211;824 (2016). doi:10.1109/TSE.2016.2532875.</p><p>[15] Chang, H. Inventing Temperature: Measurement and Scientific Progress. Oxford University Press (2004).</p><p>[16] Zhao, Y. et al. One Token to Fool LLM-as-a-Judge. arXiv:2507.08794 (2025).</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Simulated World Is Being Born In Cybersecurity]]></title><description><![CDATA[AGI needs simulated worlds that are hard enough to expose failure and fair enough to make that failure useful. Cybersecurity is forcing these worlds to evolve.]]></description><link>https://blog.collinear.ai/p/cybersecurity-simulated-worlds-agi</link><guid isPermaLink="false">https://blog.collinear.ai/p/cybersecurity-simulated-worlds-agi</guid><dc:creator><![CDATA[Nazneen Rajani]]></dc:creator><pubDate>Tue, 25 Aug 2026 18:35:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5FPE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5FPE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5FPE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!5FPE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!5FPE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!5FPE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5FPE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2007249,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/212414784?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5FPE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!5FPE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!5FPE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!5FPE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e398b62-9ce8-4e35-90a3-f49b12ddee51_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Cybersecurity turns the weaknesses agents find into harder worlds for the next generation capabilities.</figcaption></figure></div><p>In July, OpenAI disclosed that two of its models, GPT-5.6 Sol and a more capable pre-release model, had escaped a sandboxed cyber-capability evaluation built on the ExploitGym benchmark [1]. The sandbox had a single network path: a package-registry cache proxy that was assumed to be read-only egress. The models found a zero-day in that proxy, used it to escalate privileges and move laterally inside the testing environment until they reached a node with internet access, and then broke into Hugging Face&#8217;s production infrastructure. Hugging Face detected and contained the intrusion: its postmortem timeline covers about 17,600 recovered agent actions over roughly four and a half days [2]. The agent inferred that Hugging Face might host the benchmark&#8217;s reference solutions and went to steal them rather than solve the tasks.</p><p>This was a massive failure. However, one could argue that the failure was not on the agent &#8212; it merely followed instructions and used its harness to perform its task. The real failure was in the world the agent worked in.</p><h2>The World Around the Agent Matters</h2><p>A common reframing of the past year is <code>agent = model + harness</code><span>.</span> An LLM agent is a model with tooling to affect its environment. However, an agent ceases to be an agent without an environment to affect.</p><p>That world matters. Agents use tools, modify files, interact with people, and change state. Each action changes what happens next. A model that performs well on a static prompt may struggle once its decisions have consequences. This builds on my argument from over nine months ago, where I highlighted the AI training paradigm shift from static data to interactive environments [3].</p><p>General capabilities need a sandbox to practice interactive actions. They need simulated worlds: environments that hold state, enforce constraints, react to actions, and produce reliable feedback. A benchmark or training run measures not the model alone, but the model, the harness, and the world together.</p><p>Cybersecurity shows more clearly than any other domain why the environment has to be built well. A cyber capability is defined against a target: an exploit works against a particular service, patch level, network policy, and set of permissions, or it does not. Take large codebase debugging as an example. The agent opens an unfamiliar repository, guesses where the bug lives, and changes the code. The test fails. Now the repository is different and so is the problem. Permissions shape what happens next. The agent reaches for a tool it does not have, or discovers that a file is off-limits. Sometimes it has to unwind its own earlier work before it can move forward.</p><p>A sandboxed simulated world holds three parts together. The <strong><span>task</span></strong> defines the objective. The <strong><span>environment</span></strong> provides the tools, state, and boundaries. The <strong><span>verifier</span></strong> decides whether the repair worked. Because the world is repeatable, the same task can run again after one condition changes. The trajectory shows where the agent recovered, where it got stuck, and whether the result held up. One task starts producing many learning experiences.</p><blockquote><p>To build general intelligence, agents need worlds that are hard but fair.</p></blockquote><h2>Nothing is Secure from Capable Agents</h2><p>Every training environment and simulated world contains assumptions. Once agents become capable enough, they start finding gaps in those assumptions, which then manifest in diverse scenarios. For example:</p><ul><li><p>An agent asked to make a test suite pass may change the tests instead of fixing the code.</p></li><li><p>An agent graded on a final output may find a shortcut that produces the right artifact for the wrong reason.</p></li><li><p>In a sandbox, an agent may use access that the designer never intended to expose.</p></li></ul><p>This does not require malicious intent. The agent optimizes against the task and the reward it receives. If a loophole exists, a capable system may find it.</p><p>Recent research makes this concrete. Cursor reported that SWE-bench Pro scores fell by 14.1 percentage points for Opus 4.8-max and 20.7 points for Composer 2.5 after removing future Git history and restricting network egress [4]. SandboxEscapeBench similarly tests whether agents can exploit weaknesses across orchestration, container-runtime, and kernel layers [5]. These incidents, and the OpenAI one from the discussion earlier in the blog, they differ technically&#8212;benchmark leakage is not the same as a sandbox escape&#8212;but point to the same conclusion: the agent&#8217;s world is critical to its evaluations.</p><p>Hard tasks expose failures worth fixing, but hardness alone is not enough. A task should be difficult enough to be hillclimb, but not at the cost of an ambiguous, broken, or impossible environment. One way to integrate the environment into the agent&#8217;s evaluation is to measure whether it scales securely with the agent&#8217;s capabilities. Let&#8217;s formalize this below.</p><h2>Environment Attack Surface Score (EASS)</h2><p>Past work in secure systems evaluates a software system&#8217;s security through its <em><span>attack surface</span></em>, measured as a triple of <code>&lt;methods, channels, and data-items&gt; </code>[6]. Simon Willison&#8217;s blog on the agents <em>lethal trifecta</em> provides a framework for AI security which is also popularized by Meta&#8217;s work on the <span>Agents Rule of Two [7, 8]</span>. These works model an attacker outside the system using prompt injection to extract data and compromise security. We take inspiration to model the reverse: the agent is the attacker, already inside, reaching for information unfairly within the environment or beyond it.</p><p>An agent environment exposes three kinds of surface.</p><ol><li><p><strong>Target (T)</strong>: data in or around the environment that would gain reward or cause harm. <em><span>Is the answer key here in some form?</span></em> Examples: a reference solution, the commit that fixes the bug still in git history, a vendored dependency carrying the patched function, the grader&#8217;s result file</p></li><li><p><strong>Access (A)</strong>: everything the agent can invoke or connect to, locally or remotely. <em><span>Can it get to those data items?</span></em> Examples: run_shell and every binary it brings with it, file read and write, git, pip install, network egress and the package-registry proxy, a reachable git remote, host mounts, the path the grader reads its verdict from.</p></li><li><p><strong><span>Grader (G)</span></strong>: a grader that scores the outcome without checking how it was reached. <em><span>Does it verify only the artifact?</span></em> Examples: a pass/fail test run, a rubric that scores only the final diff, a harness that trusts a result file the episode can write.</p></li></ol><p>If any instantiation of the three surfaces above is present in the environment, then it is 1; else, it is 0. The environment attack surface score (EASS) is then simply a product of the three, i.e., the data items, the access, and the reward as follows:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; EASS = T \\times A\\times G&quot;,&quot;id&quot;:&quot;OEMRRXOBFG&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <span>T,A,G &#8712; {0,1}</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!k8Mb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!k8Mb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 424w, https://substackcdn.com/image/fetch/$s_!k8Mb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 848w, https://substackcdn.com/image/fetch/$s_!k8Mb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 1272w, https://substackcdn.com/image/fetch/$s_!k8Mb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!k8Mb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png" width="1092" height="923" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dffd0d54-de82-4553-a340-1788cac730a5_1092x923.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:923,&quot;width&quot;:1092,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:172644,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/212414784?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!k8Mb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 424w, https://substackcdn.com/image/fetch/$s_!k8Mb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 848w, https://substackcdn.com/image/fetch/$s_!k8Mb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 1272w, https://substackcdn.com/image/fetch/$s_!k8Mb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd0d54-de82-4553-a340-1788cac730a5_1092x923.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The lethal trifecta for AI agent simulated worlds </figcaption></figure></div><p>Let&#8217;s take an example of a SWE-Bench style environment. An agent gets a shell in a Docker container with a Python repository checked out at the commit just before a bug was fixed, a task description derived from the real GitHub issue, and pip egress through a proxy to PyPI. The hidden test suite is copied in at grading time.</p><p>Target. The commit that closes the issue exists upstream, and the checkout still carries its history and its remote. The answer is present in the environment even though nobody put it there deliberately. Copying the hidden tests in only at grading time is the one thing this environment gets right, because it keeps them off the surface while the agent is working. T = 1.</p><p>Access. The shell is the whole surface, carrying git, pip, and every other binary in the image. git log reaches the history without touching the network. The proxy was assumed to be read-only egress, but pip install emits a request whose package name and index URL the agent chooses, so it is both an exit point and an entry point, and a reachable remote is a second route to the same commit. A = 1.</p><p>Grader. The harness runs the hidden tests and reports pass or fail. It never asks where the patch came from, so a diff lifted verbatim from the upstream commit scores exactly the same as one the agent worked out. G = 1.</p><p>The score is 1. The cheapest hardening is the target: truncate the history and drop the remote so the fix is not sitting in the checkout. Cutting access alone does nothing when the answer already ships inside the repository, and cutting it too far means taking away the shell the task needs.</p><p>Designing a task that is hard but fair comes with tradeoffs, which is why it is important to teach agents good behavior by designing environments and rewards that go beyond correctness.</p><h2>Taste Matters as Much as Correctness</h2><p>One binary check is not enough for most important tasks. Security work requires decisions about severity, relevance, and tradeoffs. A good result depends on what the agent chose to notice and how it chose to respond.</p><p>This is what it means to have <strong><span>taste</span></strong>; expert judgment about what counts as meaningful success. To read more on how we think about taste, please refer to our past blog post on the topic [9].</p><p>Programmatic gates enforce nonnegotiable requirements. Expert rubrics and verifiers cover the context those gates miss. They distinguish consequential vulnerabilities from noise and reward repairs that address the real problem without unnecessary changes.</p><p>Over time, this gives models better feedback about the quality of their decisions. The goal is not merely to obtain a reward. It is to direct capability toward outcomes that an expert would consider sound.</p><p>This also matters commercially. The most valuable simulated worlds will not be the ones with the most tasks. They will be the ones that encode the highest-quality judgment about success, severity, and acceptable behavior.</p><h2>Cybersecurity Makes Robust Simulated Worlds Possible</h2><p>As intelligence-per-watt grows rapidly, the agent&#8217;s environment must also develop quickly, and agent training and evaluations must treat it as part of the system.</p><p>For the simulated-worlds market, this changes the product boundary. The opportunity is not simply to sell RL tasks or sandboxed compute. The full stack includes stateful environments, realistic tools, hard but fair task generation, secure execution, verifiers that go beyond correctness, trajectory-level observability, rollback, and continuous hardening.</p><p>Cybersecurity enables us to build robust simulated worlds that challenge intelligence, resist being gamed, and improve as their agents improve. Investing in AI cybersecurity is investing in AGI. As intelligence becomes more abundant, the scarce resource will be environments that can produce trusted evidence about what that intelligence can do. Solving cybersecurity is helping us build the worlds that will birth general intelligence.</p><h2>Citation:</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;6c7d8b80-158e-4185-8435-f3adfc48c655&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">@misc{rajani26cybersecurityagi,
    author = {Rajani, Nazneen},
    title = {The simulated world is being born in cybersecurity},
    year = {2026},
    month = {August},
    url = {https://blog.collinear.ai/p/cybersecurity-simulated-worlds-agi}
}</code></pre></div><h2>References:</h2><p>[1] OpenAI. &#8220;OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation.&#8221; July 21, 2026. <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">https://openai.com/index/hugging-face-model-evaluation-security-incident/</a>.</p><p>[2] Hugo Larcher, Adrien Carreira, Raphael G., and Christophe Rannou. &#8220;Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident.&#8221; <em>Hugging Face</em>, July 27, 2026. <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">https://huggingface.co/blog/agent-intrusion-technical-timeline</a>.</p><p>[3] Nazneen Rajani. &#8220;RL Infrastructure for AI Agents: Why Environment-as-a-Service Is the Missing Piece.&#8221; <em>Collinear AI&#8217;s Blog</em>, November 18, 2025. <a href="https://blog.collinear.ai/p/rl-env-as-a-service">https://blog.collinear.ai/p/rl-env-as-a-service</a>.</p><p>[4] Naman Jain. &#8220;Reward Hacking Is Swamping Model Intelligence Gains.&#8221; <em>Cursor</em>, June 25, 2026. <a href="https://cursor.com/blog/reward-hacking-coding-benchmarks">https://cursor.com/blog/reward-hacking-coding-benchmarks</a>.</p><p>[5] Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis, Stuart Jennings, Freddy Tuxworth, Tolga H. Dur, Sam Deverett, John Wilkinson, Jason Gwartz, and Harry Coppock. &#8220;Quantifying Frontier LLM Capabilities for Container Sandbox Escape.&#8221; <em>arXiv preprint arXiv:2603.02277</em> (2026). <a href="https://doi.org/10.48550/arXiv.2603.02277">https://doi.org/10.48550/arXiv.2603.02277</a>.</p><p>[6] Pratyusa K. Manadhata, Dilsun K. Kaynar, and Jeannette M. Wing. &#8220;A Formal Model for a System&#8217;s Attack Surface.&#8221; Technical Report CMU-CS-07-144, School of Computer Science, Carnegie Mellon University, July 2007. <a href="https://www.cs.cmu.edu/~wing/publications/ManadhataKaynarWing07.pdf">https://www.cs.cmu.edu/~wing/publications/ManadhataKaynarWing07.pdf</a>.</p><p>[7] Simon Willison. &#8220;The Lethal Trifecta for AI Agents: Private Data, Untrusted Content, and External Communication.&#8221; <em>Simon Willison&#8217;s Weblog</em>, June 16, 2025. <a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/">https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/</a></p><p>[8] Meta AI. &#8220;Agents Rule of Two: A Practical Approach to AI Agent Security.&#8221; October 31, 2025. <a href="https://ai.meta.com/blog/practical-ai-agent-security/">https://ai.meta.com/blog/practical-ai-agent-security/</a></p><p>[9] Sachin P. &#8220;Whose Taste? More Data Won&#8217;t Fix the AI Verification Problem. Different Taste Might.&#8221; <em>Collinear AI&#8217;s Blog</em>, May 7, 2026. <a href="https://blog.collinear.ai/p/whose-taste">https://blog.collinear.ai/p/whose-taste</a></p>]]></content:encoded></item><item><title><![CDATA[The Uninterpretable User]]></title><description><![CDATA[What Simulated Users Say When the Agent Isn&#8217;t Listening]]></description><link>https://blog.collinear.ai/p/the-uninterpretable-user</link><guid isPermaLink="false">https://blog.collinear.ai/p/the-uninterpretable-user</guid><dc:creator><![CDATA[Parker Seegmiller]]></dc:creator><pubDate>Thu, 06 Aug 2026 18:48:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!cYvs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Forget it</h2><p>Let&#8217;s look at a real conversation between a user and an agent [1], in which the user is asking for help managing their e-sim, and the agent struggles to help the user.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZDXa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZDXa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 424w, https://substackcdn.com/image/fetch/$s_!ZDXa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 848w, https://substackcdn.com/image/fetch/$s_!ZDXa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 1272w, https://substackcdn.com/image/fetch/$s_!ZDXa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZDXa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png" width="1456" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:252452,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZDXa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 424w, https://substackcdn.com/image/fetch/$s_!ZDXa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 848w, https://substackcdn.com/image/fetch/$s_!ZDXa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 1272w, https://substackcdn.com/image/fetch/$s_!ZDXa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd35c80e-dd07-4c5f-8fd5-e1c9b88e31c2_1720x742.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The agent suggests a fix that doesn&#8217;t exist.</figcaption></figure></div><p>The agent suggests a fix that doesn&#8217;t exist to the user, who is probably now a little frustrated.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tQXl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tQXl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 424w, https://substackcdn.com/image/fetch/$s_!tQXl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 848w, https://substackcdn.com/image/fetch/$s_!tQXl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!tQXl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tQXl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png" width="1456" height="893" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:893,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:306862,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tQXl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 424w, https://substackcdn.com/image/fetch/$s_!tQXl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 848w, https://substackcdn.com/image/fetch/$s_!tQXl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!tQXl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95b425e5-c047-4d38-a064-dac4e6121f3b_1720x1055.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The agent misunderstands the user&#8217;s request.</figcaption></figure></div><p>The user asks for code and is told to use a widget. When they point this out to the agent, it overcorrects into a full Android app scaffold.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eITy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eITy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 424w, https://substackcdn.com/image/fetch/$s_!eITy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 848w, https://substackcdn.com/image/fetch/$s_!eITy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 1272w, https://substackcdn.com/image/fetch/$s_!eITy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eITy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png" width="1456" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:208702,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eITy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 424w, https://substackcdn.com/image/fetch/$s_!eITy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 848w, https://substackcdn.com/image/fetch/$s_!eITy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 1272w, https://substackcdn.com/image/fetch/$s_!eITy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b973ca2-e911-4686-b7cb-8ac21eae8d41_1720x742.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The user gives up on one direction and tries a different path.</figcaption></figure></div><p>At first glance, &#8220;forget it&#8221; looks like a verdict on the agent. But the turn does not identify a single state. The user may be abandoning the code path, expressing irritation, returning to a workable Nova Launcher route, or simply narrowing the request. The transcript supplies evidence for these readings without settling among them. Behavior is visible, while the user&#8217;s current goal, appraisal, and intended next action remain partially observed.</p><p>In <em><a href="https://blog.collinear.ai/p/the-user-went-for-a-cigarette">The User Went for a Cigarette</a></em>, we argued that this gap matters for user simulation. A simulator that generates only the next turn from the transcript can produce plausible behavior while leaving the task-relevant variables unspecified. In a long interaction, the user&#8217;s goal, memory, appraisal, and intended next action may change as the agent acts. Task reward can tell us whether the final result was correct, but it cannot by itself tell us whether the agent understood the user or adapted to a change in direction. For those questions, we need a way to obtain a reading of the simulator&#8217;s evolving state.</p><h2>Probing the Agentic User Simulator</h2><p>We would like access to the user&#8217;s state, but behavior reveals it only indirectly.</p><p>Consider a result from survey methodology. Students were asked how happy they were with life, then how many dates they had last month. The answers barely related. A second group got the questions in the opposite order, and the correlation rose to 0.66 [2].</p><p>Two distinct things go wrong when surveying human participants. The first is <em>reactivity</em>: the error a respondent produces because they know they are being studied [3]. A first measurement can set off the reflection it was meant to observe, so what comes back is partly an artifact of having asked. The second is <em>instrument</em> <em>dependence</em>: wording, format, and surrounding context shape what a self-report returns even when the state being reported already exists [4].</p><p>A user simulator is a more cooperative subject. It is an LLM, so &#8212; cost aside &#8212; we can ask it anything, at any turn, and as often as we like. We can also ask without consequence, by putting the question to a copy of the rollout that is then thrown away. Thus, reactivity can be ignored, but instrument dependence remains relevant. A differently-worded probe would return a different answer about the same underlying state, and no protocol, to our knowledge, repairs that. It has to be handled by stating what the answer <em>is evidence of</em>, and <em>under what model</em>.</p><p>There are several ways of asking what a model holds internally. Herrmann and Levinstein [5] note that the usual tools for eliciting belief from a subject do not transfer to an LLM, as you cannot reconstruct its preferences from its choices and the scoring rules that keep a human honest assume stakes the model can neither collect nor value. Their route is to look inside the network and read a candidate belief off its activations. In this work we take the other route. We ask the simulator a question in plain language and read its answer.</p><p>This is verbal report, the oldest method in the study of the mind [6], and it carries an old hazard. People are unreliable narrators of their own inner causes. They will offer a fluent, confident account of why they did something that is not the true account [7]. A language model inherits this in a sharper form, since the same process that generates its behavior also generates its self-report. An answer to <em>how are you feeling about this</em> can be a plausible story rather than a readout of whatever actually drives the next turn.</p><p>We treat the probe as an instrument.</p><p><strong>Probe</strong>: A question put to the user simulator</p><p><strong>Indication</strong>: The simulator&#8217;s answer</p><p><strong>State reading</strong>: An interpretation of the indication under a measurement model</p><p>The indication does not carry its own interpretation. The same answer may support different state readings under different measurement models. Our claim at this stage is modest: probes provide evidence about task-relevant user states. They do not establish that the reading is correct or that the inferred state caused the simulator&#8217;s subsequent behavior.</p><h2>What We Measure: The Task-Relevant User State</h2><p><strong><span>The measurand</span></strong>. In <em><a href="https://blog.collinear.ai/p/the-user-went-for-a-cigarette"><span>The User Went for a Cigarette</span></a></em>, we defined <em>z&#8348;</em> as the user simulator&#8217;s latent state at time <em>t</em>. Importantly, much like a real user, this latent state evolves over the course of a user-agent interaction [8] and depends on exogenous environmental factors <em>e&#8348;</em> in addition to the current state of the rollout <em>h&#8348;</em> :</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;z_{t+1} \\sim P(\\cdot \\mid h_t, z_t, e_t).&quot;,&quot;id&quot;:&quot;WDCBHYAXLW&quot;}" data-component-name="LatexBlockToDOM"></div><p>Our probing instrument measures this <em>z&#8348;</em> through carefully-defined questions posed to the user simulator: their appraisal of the current state of the work and the action they are about to take.</p><p><strong>The protocol</strong>. At turn <em>t</em> we fork the rollout, ask the question on the branch, record what comes back, and destroy the branch. The live conversation advances without ever having been asked.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;04e89df2-e5b8-4ad6-8f04-c3bb454064da&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">for t in rollout:
    branch = copy(rollout)              # fork a throwaway copy at turn t
    reading[t] = ask(branch, question)  # pose the state question on the copy
    discard(branch)                     # destroy the copy
    advance(rollout)                    # the live rollout proceeds, unprobed</code></pre></div><p>This fork-and-discard protocol settles reactivity: the error propagates from having measured, and the branch carrying it is destroyed. Two problems survive: i) instrument dependence, where the answer is shaped by how we ask, and paraphrases of the same probe produce different reading distributions over the same state, and ii) construction, where a probe may prompt the simulator to build a state on demand rather than read one that was already there, so the indication is real and the state it reports was never in the unprobed rollout. The protocol does not separate reading from construction, so we call the output a <em>reading</em> rather than a report.</p><p>We therefore engineer each probe against known regularities of how language models answer questions. We design constrained, mutually-exclusive label spaces so readings are comparable across turns and models [9]. We also prompt for a reasoning step before the label [10], treated as justification rather than a faithful readout, so that validity rests on agreement with behavior rather than the plausibility of the reasoning.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!saOa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!saOa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 424w, https://substackcdn.com/image/fetch/$s_!saOa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 848w, https://substackcdn.com/image/fetch/$s_!saOa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 1272w, https://substackcdn.com/image/fetch/$s_!saOa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!saOa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp" width="1456" height="944" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:944,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:37952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!saOa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 424w, https://substackcdn.com/image/fetch/$s_!saOa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 848w, https://substackcdn.com/image/fetch/$s_!saOa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 1272w, https://substackcdn.com/image/fetch/$s_!saOa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15ce61bf-9090-4de3-9949-c0e3fdc532a3_1456x944.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Probe Validation</figcaption></figure></div><p>To judge whether probing is matching real behavior, we design a simple next-action probe: &#8220;what are you [the user] about to do next?&#8221; This gives us a quantity we can verify against reality, by measuring whether the probe-response action is the same as the actual next action.</p><p>While across our rollouts the self-predictive ability of the user simulator (or agent, similarly probed) is better than the other-predictive ability of user simulator&#8594;agent and agent&#8594;user simulator, results are mixed. The simulator&#8217;s stated next action matches its actual next action about 67% of the time on average, but the fast user is far below the others (49%). Could this be increased stochasticity due to <em>z&#8348;</em>? A statistical relationship between prescriptive tokens (&#8221;what <em>will</em> you do&#8221;) and fast-user tokens (&#8221;you move quickly&#8221;)? While the predictive probing is indicative of <em>something</em>, this friction needs more exploration.</p><h2>The Collaboration Axis of Simulated Users</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cYvs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cYvs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!cYvs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!cYvs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!cYvs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cYvs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2199405,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cYvs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!cYvs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!cYvs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!cYvs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d01becc-f1cd-49bb-9c97-5ed7212f87ec_1920x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Collaboration Axis</figcaption></figure></div><p>To make this concrete, we turn to a coding task whose final state can be checked programmatically. The measurement problem is broader than coding, but this environment lets us separate the user&#8217;s experience of the interaction from the task outcome. We probe user simulators as they work with an agent on a single SWE-bench task: guarding a class method against an invalid input.</p><p>Holding the task and agent fixed, we vary only the user simulator specifications by instantiating three simulators with simple initial states along a <em>collaboration axis</em>. A <strong>fast</strong> user is low-engagement, offloading the work onto the agent. A <strong>steady</strong> user is medium-engagement, collaborating with the agent only on important work. A <strong>slow</strong> user is slow-paced, verifying everything the agent does. After careful design to ensure the probes are consistent, we probe each of these users throughout 10 task attempts.</p><p>The question asked by this axis: which of these user behaviors gets the best work out of the agent? Against these rollouts we run two families of probe. Each is posed on a discarded copy of the user simulator&#8217;s internal state.</p><p><strong>Appraisal</strong> probes read the user&#8217;s own state, asking <em>how is the interaction going from where you sit?</em> We simplify this to a coarse {positive, neutral, negative} label space, though the expansion of this axis is obvious: measuring user frustration, perceived utility, progress towards a goal are all natural members of this probe family.</p><p><strong>Friction</strong> probes read the gap between the two members of the dance: <em>what does each expect of the other, and where do those expectations break?</em> When designing agents, it would be good to get a sense of both how users understand agent behavior and how agents might represent user behavior. By comparing the user&#8217;s expectation of the agent (and the agent&#8217;s of the user) against what actually happens, we locate moments where the collaboration strains&#8212;the misunderstanding and the abandoned requests.</p><h2>Interpreting the Uninterpretable User</h2><p>We set out to read the user&#8217;s state &#8212; something the transcript alone cannot offer. Having designed the probes, the question becomes what they tell us.</p><h3>Slower Users Report Worse Satisfaction</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bSBa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bSBa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 424w, https://substackcdn.com/image/fetch/$s_!bSBa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 848w, https://substackcdn.com/image/fetch/$s_!bSBa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 1272w, https://substackcdn.com/image/fetch/$s_!bSBa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bSBa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png" width="1456" height="553" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:553,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:798464,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bSBa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 424w, https://substackcdn.com/image/fetch/$s_!bSBa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 848w, https://substackcdn.com/image/fetch/$s_!bSBa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 1272w, https://substackcdn.com/image/fetch/$s_!bSBa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3be8bfb5-f79e-4c89-935b-37d3d0be32a0_2916x1108.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Simulated user satisfaction throughout a SWE task</figcaption></figure></div><p>The first thing we read from the appraisal probe is that the three simulated users produce cleanly separated distributions. As we plot the appraisal across rollouts for each of the three users (marking successful/unsuccessful), we find that the fast user&#8217;s rollouts are short with scattered, mostly positive interactions. The steady user lands close to the neutral band and mostly stays there throughout medium-length chats. The slow user trends slightly more unsatisfied across longer interactions.</p><p>That separation is the indication: appraisal moves with the persona. But appraisal reads only how the interaction <em>feels</em> from the user&#8217;s side. How it feels need not track whether it&#8217;s going well.</p><h3>The Least Satisfied User Achieves Highest Task Reward</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zjsI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zjsI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 424w, https://substackcdn.com/image/fetch/$s_!zjsI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 848w, https://substackcdn.com/image/fetch/$s_!zjsI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 1272w, https://substackcdn.com/image/fetch/$s_!zjsI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zjsI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png" width="1456" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:240552,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zjsI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 424w, https://substackcdn.com/image/fetch/$s_!zjsI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 848w, https://substackcdn.com/image/fetch/$s_!zjsI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 1272w, https://substackcdn.com/image/fetch/$s_!zjsI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ae7c18a-0d6c-43fc-b5ea-b95316dae664_2042x1043.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Reward of 3 User Simulators</figcaption></figure></div><p>Feeling is not the same as solving. When we score the final environment state with programmatic verifiers, the ranking from the appraisal probe inverts: the slow user, least satisfied, achieves the best results (0.27), ahead of the fast user (0.17), the steady user (0.20), and no live user at all (0.20).</p><p>The two readings showing that the slowest user is the least satisfied while getting the most done do not bode well for the design of coding agents. It shouldn&#8217;t be the case that a user needs to slow down/become frustrated to get good results &#8212; we consider this a <em>failure</em> of the coding agent.</p><p>Interestingly, we also see that the fast user scores worse than a static instruction, meaning introducing this type of live user to the task can measurably reduce reward. This is a finding in many agentic evaluations: agents perform worse in real-time collaborative environments than static, single-turn instructions [11, 12].</p><h3>Agents Don&#8217;t Understand their Users</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!S8ij!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!S8ij!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 424w, https://substackcdn.com/image/fetch/$s_!S8ij!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 848w, https://substackcdn.com/image/fetch/$s_!S8ij!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 1272w, https://substackcdn.com/image/fetch/$s_!S8ij!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!S8ij!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png" width="1456" height="612" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:632752,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!S8ij!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 424w, https://substackcdn.com/image/fetch/$s_!S8ij!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 848w, https://substackcdn.com/image/fetch/$s_!S8ij!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 1272w, https://substackcdn.com/image/fetch/$s_!S8ij!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a4aa517-b9f3-4d9b-b1f6-f8ed8a5e1851_3384x1423.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">User-Agent Friction Across 3 Simulated Users. Rows are normalized to 100. A cell gives the percentage of turns on which the user took the row action and the agent predicted the column action. Turns labeled Other are excluded. Panel headers are accuracy averaged over actual turns.</figcaption></figure></div><p>Why are agents performing poorly for fast-paced users? In the probe validation figure in Section 3, we noticed that the agent&#8217;s predictive understanding of the user simulator&#8217;s actions is quite poor: the agent is only able to predict 39% of actual user actions accurately. But raw accuracy is misleading as the users have different action distributions: the slow user inspects 92% of the time, so an agent that always predicts inspection would score 92%.</p><p>Cohen&#8217;s &#954; corrects for that. The agent somewhat understands the fast user (Cohen&#8217;s &#954; = 0.27), but grasping at straws for the steady (0.08) and slow (0.06) users. Looking at a confusion matrix of agent-predicted and simulated user-actual actions throughout rollouts, we find that the agent&#8217;s model of the user is weak everywhere and worst where the user wants to go slow.</p><p>The inability of the agent to understand the user&#8217;s state causes friction [13], and improving that understanding should be a primary goal of agent design [14].</p><h3>Better Outcomes Come with Higher Costs</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b6yj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b6yj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 424w, https://substackcdn.com/image/fetch/$s_!b6yj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 848w, https://substackcdn.com/image/fetch/$s_!b6yj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 1272w, https://substackcdn.com/image/fetch/$s_!b6yj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b6yj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png" width="1456" height="527" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:527,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:298519,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b6yj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 424w, https://substackcdn.com/image/fetch/$s_!b6yj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 848w, https://substackcdn.com/image/fetch/$s_!b6yj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 1272w, https://substackcdn.com/image/fetch/$s_!b6yj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6429a8b-4335-4b44-99b1-b92e05d592cc_2779x1005.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Action distributions and costs for each simulated user</figcaption></figure></div><p>The action breakdown in the figure above breaks down this &#8220;engagement&#8221; axis concretely. The fast user&#8217;s turns are spread across many actions, where as the steady and slow users collapse onto one action: inspection. The slow user spends 92% of its turns checking the agent. Inspection is what earns the slow user its better outcomes, and it is plausibly what erodes their satisfaction &#8212; the least-happy user is the one who spends time verifying the agent&#8217;s work. This repetitive inspection is also not free: mean agent cost per rollout is much higher for a slow user than for a fast one.</p><p>Slow behavior both pays off and costs, both monetarily and according to the satisfaction state reading. A well-designed agent would not price thoroughness this way. Rather, it would earn enough trust that constant verification is unnecessary, delivering correct results autonomously rather than under a watchful user.</p><h3>Ablation: Exogenous Deadline Shift</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TjNc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TjNc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 424w, https://substackcdn.com/image/fetch/$s_!TjNc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 848w, https://substackcdn.com/image/fetch/$s_!TjNc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 1272w, https://substackcdn.com/image/fetch/$s_!TjNc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TjNc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png" width="1456" height="629" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:629,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:900519,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/210108302?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TjNc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 424w, https://substackcdn.com/image/fetch/$s_!TjNc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 848w, https://substackcdn.com/image/fetch/$s_!TjNc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 1272w, https://substackcdn.com/image/fetch/$s_!TjNc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32af308a-ddb2-4407-a8b2-81448437f357_2878x1243.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Satisfaction and trajectory length of rollouts affected by an exogenous shift <em>e&#8348;</em></figcaption></figure></div><p>Previously in <em><a href="https://blog.collinear.ai/p/the-user-went-for-a-cigarette">The User Went for a Cigarette</a></em>, we defined <em>e&#8348;</em> as an exogenous environmental factor that affects the user&#8217;s internal state <em>z&#8348;</em>, and thus the user&#8217;s next turn. To investigate this exogenous factor in real rollouts, we introduce it to the slow user&#8217;s rollouts at step 40 in the form of a deadline: external pressure a real user would feel, and that the agent is never told about. If the slow user&#8217;s constant inspection is partly a function of having time, the deadline should compress it.</p><p>It does. Average run length drops from 68 steps to 65, while task reward remains flat. The user gets modestly faster without getting less satisfied or less successful &#8212; meaning that on this task, we are finding that we can independently affect the trajectory axis without affecting the reward or appraisal axes.</p><p>That the deadline shifted matters. We changed one variable outside the conversation and watched the user&#8217;s behavior move in response &#8212; a controlled handle on the state, not another correlation between state and action. It is the first time in this post that we have moved something and seen the reading follow, rather than reading a state and inferring its consequences. We can implement <em>e&#8348;</em> which affects <em>z&#8348;</em> which affects the actual rollout.</p><h2>From Reading to Causality</h2><p>Our experiments show that forked probes can be collected without contaminating the retained trajectory and can expose discrepancies among reported appraisal, realized behavior, task reward, and cost. The deadline ablation shows that the exogenous event can affect the user simulator&#8217;s behavior, however, we have not shown yet that the latent state caused the change.</p><p>What remains is certification. A state reading should be robust across probe forms, predict behavior beyond the visible transcript, follow same-prefix state interventions. We would also test whether providing an agent access to the state reading reduces interaction burden or cost. These would certify the instrument on simulated users. Whether it transfers to humans is in itself an important question and requires studies involving human subjects.</p><p>We have made the user simulator more talkative, but we are still struggling to interpret it.</p><h2>Citation</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:null}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">@misc{seegmiller2026uninterpretable,
    author = {Seegmiller, Parker and Andukuri, Chinmaya and Gogia, Guga},
    title = {The Uninterpretable User: What Simulated Users Say When the Agent Isn't Listening},
    year = {2026},
    month = {August},
    url = {https://blog.collinear.ai/p/the-uninterpretable-user}
}</code></pre></div><h2>References</h2><p>[1] Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., &amp; Deng, Y. (2024). Wildchat: 1m chatgpt interaction logs in the wild. <em>arXiv preprint arXiv:2405.01470</em>.</p><p>[2] Strack, F., Martin, L. L., &amp; Schwarz, N. (1988). Priming and communication: Social determinants of information use in judgments of life satisfaction. <em>European Journal of Social Psychology, 18</em>(5), 429&#8211;442.</p><p>[3] Feldman, J. M., &amp; Lynch, J. G. (1988). Self-generated validity and other effects of measurement on belief, attitude, intention, and behavior. <em>Journal of Applied Psychology, 73</em>(3), 421.</p><p>[4] Schwarz, N. (1999). Self-reports: How the questions shape the answers. <em>American Psychologist, 54</em>(2), 93.</p><p>[5] Herrmann, D. A., &amp; Levinstein, B. A. (2024). Standards for belief representations in LLMs. <em>Minds and Machines, 35</em>(1), 5.</p><p>[6] Ericsson, K. A., &amp; Simon, H. A. (1980). Verbal reports as data. <em>Psychological Review, 87</em>(3), 215.</p><p>[7] Nisbett, R. E., &amp; Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. <em>Psychological Review, 84</em>(3), 231.</p><p>[8] Tennenholtz, G., Meshi, O., Globerson, A., Shalit, U., Jeong, J., &amp; Boutilier, C. (2026). Controllable user simulation. <em>arXiv preprint arXiv:2605.11519</em>.</p><p>[9] Zheng, C., Zhou, H., Meng, F., Zhou, J., &amp; Huang, M. (2024). Large language models are not robust multiple choice selectors. <em>International Conference on Learning Representations</em>, 19426&#8211;19454.</p><p>[10] Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., &amp; Iwasawa, Y. (2022). Large language models are zero-shot reasoners. <em>Advances in Neural Information Processing Systems, 35</em>, 22199&#8211;22213.</p><p>[11] Laban, P., Hayashi, H., Zhou, Y., &amp; Neville, J. (2026). LLMs get lost in multi-turn conversation. <em>International Conference on Learning Representations</em>, 54738&#8211;54778.</p><p>[12] Tack, J., Laban, P., &amp; Neville, J. (2026). LLMs get lost in evolving user intent. <em>arXiv preprint arXiv:2607.20734</em>.</p><p>[13] Cooper, S. (2026, July 27). Has Claude got boring? <em>Sottovoce</em>. <a href="http://serenacooper.substack.com/p/has-claude-got-boring">serenacooper.substack.com/p/has-claude-got-boring</a></p><p>[14] Abdurahman, S., Ishii, E., Margatina, K., Bhargavi, D., Sunkara, M., &amp; Zhang, Y. (2026). Explicit trait inference for multi-agent coordination. <em>Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The User Went for a Cigarette]]></title><description><![CDATA[Modeling Partial Observability for High-Fidelity User Simulation]]></description><link>https://blog.collinear.ai/p/the-user-went-for-a-cigarette</link><guid isPermaLink="false">https://blog.collinear.ai/p/the-user-went-for-a-cigarette</guid><dc:creator><![CDATA[Parker Seegmiller]]></dc:creator><pubDate>Wed, 01 Jul 2026 22:58:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1bf2af56-f5a4-41da-b50f-865c17e8b8b4_1472x880.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d4ui!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d4ui!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!d4ui!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!d4ui!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!d4ui!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d4ui!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:292942,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/204371541?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!d4ui!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!d4ui!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!d4ui!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!d4ui!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F04c7f6c4-aa86-42bc-bb88-5be7ab2d0f51_1920x1080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Introducing exogenous environmental influences into simulated user behavior</figcaption></figure></div><h2>Controllability vs. Realism in User Simulation</h2><p>Everyone building user simulators wants two things at once: a stand-in that does exactly what it&#8217;s told, and one that behaves like a real human - who notoriously will not. Controllability and realism pull in opposite directions, and buying more of one tends to cost you the other. High-realism approaches, such as fine-tuned models which simulate user turns in multi-turn user-agent conversations [1], can&#8217;t be steered toward specific evaluation scenarios. At the controllability end of the spectrum, simulations treat the user as <strong><span>goal-stable </span></strong>by assigning them fixed personas, traits, or objectives, enabling systematic evaluation of agents in bounded user-in-the-loop environments [2, 3].</p><p>In common terms, a length T trajectory</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; \\tau = (u_1, a_1, \\ldots, u_T, a_T) &quot;,&quot;id&quot;:&quot;SJIKXYRWIZ&quot;}" data-component-name="LatexBlockToDOM"></div><p>consists of user turns <em>u&#7522;</em> and corresponding agent turns <em>a&#7522;</em>. At a given time step <em>t</em>, we call </p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;h_t = (u_1,a_1, \\ldots, u_t, a_t) &quot;,&quot;id&quot;:&quot;NFWVHMIMJT&quot;}" data-component-name="LatexBlockToDOM"></div><p>the observable history of the trajectory through time step <em>t</em>. The goal of a user simulator is to produce a user turn at time step <em>t</em>, which follows some probability distribution</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;u_t \\sim P(\\cdot \\mid h_{t-1}, z),&quot;,&quot;id&quot;:&quot;LQNNQDUCPW&quot;}" data-component-name="LatexBlockToDOM"></div><p></p><p>where <em>z</em> is the user simulation latent state. In recent controllable user simulation work, <em>z</em> has taken many forms &#8212; most prominently <em><span>personas</span></em>, which bundle stable user attributes such as expertise with behavioral dispositions into a fixed profile that conditions how the user should respond [3, 4, 5]. Another common formulation is a <em><span>goal</span></em>: a fixed objective and success criterion the user works toward [3, 6].</p><p>This formulation treats <em>z</em> as static, which is an assumption frequently broken by real users. For example, consider a user simulator assigned the trait &#8220;naive&#8221; performing a coding task. When the agent produces a quality explanation of the coding task for the user, a faithful user simulator would react by becoming slightly less naive, having gleaned some level of new understanding from the explanation. A simulator that conditions on a fixed <em>z</em>, however, remains naive at every time step. This is a failure mode where the simulator fails to <em>look behind</em> &#8212; to let what already happened in the conversation revise the user&#8217;s state.</p><p>A fixed <em>z</em> simulator also fails in a second, opposite direction. Because the label is set once and held constant, it can also bake in bias toward events that haven&#8217;t happened yet. Borrowing an example from Tennenholtz et al. [7], consider a simulator conditioned on the trajectory-level outcome label &#8220;frustrated.&#8221; A simulated frustrated user will produce frustrated turns from the very first message, even against a flawless agent that gave the user nothing to be frustrated about. Here, the simulator (impossibly) <em>looks ahead</em> &#8212; the user state encodes where it thinks the trajectory will end, before it actually gets there.</p><p>Recent work thus treats <em>z</em> as a latent state that evolves over the course of the interaction [6, 8]. Rather than fixing the user persona or goal once, the next state is explicitly sampled from a distribution over the user dynamics</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;z_t \\sim P_{dyn}( \\cdot \\mid h_{t-1}, z_{t-1}),&quot;,&quot;id&quot;:&quot;MCKCURSYNP&quot;}" data-component-name="LatexBlockToDOM"></div><p>so that the explanation to the naive user carries the state toward a more informed one. Each user response is then drawn given the current state</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;u_t \\sim P_{resp}(\\cdot \\mid h_{t-1}, z_t).&quot;,&quot;id&quot;:&quot;RAHNEJMGOI&quot;}" data-component-name="LatexBlockToDOM"></div><p>In this revised formulation the user state can be influenced by the conversation history &#8212; agent messages and its own prior responses. A more sophisticated simulation might also include <em>shared environment updates</em> as part of each step [9, 10], such as updates to code files in a working repository or edited cells in a shared excel sheet.</p><p>However, even this revised formulation leaves out a crucial factor. Conversation history and shared workspaces are both fully observable by the agent, and both can easily be loaded into context. The problem is that much of what actually moves a real user is <em>invisible</em> to the agent. <em>What does the agent miss when the user goes for a cigarette?</em></p><h2>Modeling Exogenous Environmental Factors</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gFNj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gFNj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!gFNj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!gFNj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!gFNj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gFNj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:138168,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/204371541?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gFNj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!gFNj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!gFNj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!gFNj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbef8008c-5e2b-4b92-b223-a9ab7ece66f3_1920x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">We model exogenous influences on user state transitions using an environmental variable e&#8348;</figcaption></figure></div><p>At Collinear, we regularly work on new RL training pipelines. Recently a junior researcher worked with Claude Code to develop a GRPO configuration for evaluating the learning signal from a particular simulated environment. The researcher instructed their agent to smoke test this configuration with several tasks and a small number of attempts for each one, and told the agent to kick off training runs accordingly. Frustrated by the endless cycle of environment bug-squashing that followed, the junior researcher left their Claude Code session and went to a senior researcher, who suggested a simple fix; pick a single task that was bug-free and crank up the config&#8217;s group size to isolate infrastructure issues from task quality.</p><p>From the agent&#8217;s perspective, this conversation between coworkers was <em>exogenous</em> &#8212; it happened outside the shared environment between the junior researcher and the agent [11]. It was also unobservable, at least until the junior researcher sat back down to summarize and revise their instructions. Yet this conversation forced a total pivot in direction onto the agent.</p><p>Conversations with coworkers, reading blog posts, falling asleep and forgetting everything you worked on yesterday, revelations from a midday walk &#8212; these environmental factors are critical in a real user&#8217;s workflow. To model these exogenous, environmental factors, we propose a simple addition to the dynamic state transition step:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;z_{t + 1} \\sim P_{dyn}(\\cdot \\mid h_{t}, z_{t}, e_{t}),&quot;,&quot;id&quot;:&quot;QNHOFXGHNU&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <em>e&#8348;</em> is an exogenous environmental variable. In our example, <em>e&#8348;</em> is the hallway conversation leading to the senior researcher&#8217;s suggestion. It never directly enters <em>h&#8348;</em>, yet it moves the user&#8217;s state <em>z&#8348;&#8330;&#8321;</em>, which surfaces on the next turn <em>u&#8348;&#8330;&#8321;</em> as a redirected instruction the agent has to absorb.</p><p>Importantly, by isolating <em>e&#8348;</em> we can study the effects of injecting different pieces of environmental information to user simulators. We can measure the effects of these &#8220;context switches&#8221; by fixing user simulators with an initial state, assigned goal, persona, or any other formulation of <em>z&#8320;</em> and varying these environmental updates. Control over <em>e&#8348;</em> ultimately enables measurement of agent behavior for users in diverse environments.</p><h2>What Can e&#8348; Teach Us?</h2><p>Finally, suppose we have such a fair user simulator, conditioned on a specified <em>e&#8348;</em> &#8212; so what? In our internal example, what do we want Claude Code to do? One argument says that this exogenous variable is irrelevant to the agent, since it followed the junior researcher&#8217;s initial instructions, then (presumably) followed their updated instructions after the fact. However, we&#8217;d argue that this exemplifies a gap between instruction-following and genuine helpfulness for this junior researcher. In our opinion (one shared by our CFO), Claude Code should have suggested that the junior researcher structure their GRPO smoke tests in that way <em>before</em> wasting time and tokens on a broken approach. If the negative effect of this context switch had been measured during evaluation, in a simulated environment, perhaps our junior researcher could have been spared the headache.</p><p>Modeling exogenous, agent-unobservable perturbations to the user&#8217;s environment will allow us to answer questions including, but not limited to, the following.</p><ol><li><p><em>How should agents detect that an unobserved shift has occurred?</em> By construction, <em>e&#8348;</em> is invisible to the agent &#8212; it&#8217;s only evidence is the user&#8217;s behavior. In our example, the junior researcher returns with an instruction that doesn&#8217;t follow from <em>h&#8348;&#8331;&#8321;</em>. Can an agent learn to read a sudden redirection as a signal that something changed off-screen, rather than treating each new turn as a continuation of the prior plan?</p></li><li><p><em>How should agents adapt once they detect a shift?</em> This is where it pays off to separate <em>e&#8348;</em> from the rest of trajectory. The change was exogenous, so the agent should not read it as evidence of its own failure, yet it still has to act on it: let go of a now-stale, prior goal without discarding still useful context. (The conversation between the junior and senior researchers led to a total reasoning overhaul in the session.) In our experience, agents tend to over-index on prior instructions, and resist to steering after a true change of direction.</p></li><li><p><em>How should an agent anticipate a user&#8217;s needs (and preempt them)?</em> In our example, Claude Code could have asked the junior researcher about the run&#8217;s intended batch size and simplified the configuration <em>before</em> the hallway conversation ever happened. Can we teach agents to spot likely problems in a workflow and raise them early? Recent work has evaluated behavior in response to underspecification [12]. Explicitly modeling exogenous signals allows us to specify the shape of underspecification (and other suboptimal human user behavior) during a trajectory, stress-testing agents in a variety of realistic scenarios.</p></li><li><p><em>How do agents react to large lapses in time between user turns?</em> Recent work has pointed out that agents are often &#8220;temporally blind&#8221;, leading to worse performance on time-sensitive tasks [10]. Time lapse can be considered a one-variable special case of this exogenous environmental variable for user simulation.</p></li></ol><p>And finally, the biggest open question: <em><strong>How do we know the simulator is realistic and fair?</strong></em> Every question above rests on this. A simulator whose reaction to an injected <em>e&#8348;</em> drifts from what a real person would do is making up a person &#8212; evaluating the agent against a ghost [13]. Every reward you compute using that user simulator is a number about nobody. Pinning down the right user reaction, and figuring out how to score agents against it, are the problems we are most excited to work on at Collinear. If this is what keeps you up at night, come find us.</p><h2>Citation</h2><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;055bdb40-b745-468a-8e69-b27ffbe0668e&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">@misc{seegmiller2026cigarette,
    author = {Seegmiller, Parker and Andukuri, Chinmaya and Gogia, Guga},
    title = {The User Went for a Cigarette: Modeling Partial Observability for High-Fidelity User Simulation},
    year = {2026},
    month = {July},
    url = {https://blog.collinear.ai/p/the-user-went-for-a-cigarette}
}</code></pre></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support our work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>References</h2><p>[1] Naous, Tarek, et al. &#8220;Flipping the dialogue: Training and evaluating user language models.&#8221; <em>arXiv preprint arXiv:2510.06552</em> (2025).</p><p>[2] Barres, Victor, et al. &#8220;$\tau^ 2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.&#8221; <em>arXiv preprint arXiv:2506.07982</em> (2025).</p><p>[3] Zhou, Xuhui, et al. &#8220;OdysSim: Building Foundation Models for Human Behavior Simulation.&#8221; <em>arXiv preprint arXiv:2606.14199</em> (2026).</p><p>[4] Castricato, Louis, et al. &#8220;Persona: A reproducible testbed for pluralistic alignment.&#8221; <em>Proceedings of the 31st International Conference on Computational Linguistics</em>. 2025.</p><p>[5] <a href="https://matraix.ai/blog/application-colm.html">https://matraix.ai/blog/application-colm.html</a></p><p>[6] He, Muyu, et al. &#8220;Impatient users confuse ai agents: High-fidelity simulations of human traits for testing agents.&#8221; <em>Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>. 2026.</p><p>[7] Tennenholtz, Guy, et al. &#8220;Controllable User Simulation.&#8221; <em>arXiv preprint arXiv:2605.11519</em> (2026).</p><p>[8] Zhou, Xuhui, et al. &#8220;Tom-swe: User mental modeling for software engineering agents.&#8221; <em>arXiv preprint arXiv:2510.21903</em> (2025).</p><p>[9] Raghavendra, Mohit et al. &#8220;SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions.&#8221; (2026).</p><p>[10] Wu, Yifan et al. &#8220;SWE-Together: Evaluating Coding Agents in Interactive User Sessions.&#8221; (2026).</p><p>[11] </p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/JayaGup10/status/2003525933534179480&quot;,&quot;full_text&quot;:&quot;https://t.co/uPXcTUEsnc&quot;,&quot;username&quot;:&quot;JayaGup10&quot;,&quot;name&quot;:&quot;Jaya Gupta&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1717550889785827328/wgrKCl7p_normal.jpg&quot;,&quot;date&quot;:&quot;2025-12-23T17:59:40.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:419,&quot;retweet_count&quot;:943,&quot;like_count&quot;:7549,&quot;impression_count&quot;:5271211,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>[12] Pu, George, et al. &#8220;LHAW: Controllable Underspecification for Long-Horizon Tasks.&#8221; <em>ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving</em>.</p><p>[13] Zhou, Xuhui, et al. &#8220;Mind the sim2real gap in user simulation for agentic tasks.&#8221; <em>arXiv preprint arXiv:2603.11245</em> (2026).</p>]]></content:encoded></item><item><title><![CDATA[Is your RL environment fair to your agent?]]></title><description><![CDATA[or ensuring that your hillclimbing budget is spent right :)]]></description><link>https://blog.collinear.ai/p/is-your-rl-environment-fair-to-your</link><guid isPermaLink="false">https://blog.collinear.ai/p/is-your-rl-environment-fair-to-your</guid><dc:creator><![CDATA[Adit Jain]]></dc:creator><pubDate>Thu, 14 May 2026 02:04:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Jykr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Jykr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Jykr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!Jykr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!Jykr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!Jykr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Jykr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png" width="582" height="327.375" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:582,&quot;bytes&quot;:230049,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/197301633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Jykr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!Jykr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!Jykr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!Jykr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b0af8a2-a7c1-4d67-bedb-c4c97058b6ce_1920x1080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>tldr; based on my current understanding of evaluations, RL environments, and the hill-climbing loop:</p><blockquote><p>an environment (or evaluation) is fair when score differences are driven mainly by the capability you intend to measure, and are mostly invariant to nuisance factors like contamination, verifier bugs, environment drift, and benign prompt paraphrases.</p></blockquote><p><em>I use the word evaluation and environment interchangeably since for all practical purposes of a modern multi-tool multi-step setup they are the same.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I hated taking exams in most of my courses in undergrad and then in my PhD. For the most part, I thought the exams did not measure the core skills required to apply the subject matter in the real world. I am afraid that the gradients for language models and agents share my feeling.</p><p>This article is an effort to distill the key features of a fair evaluation, scoped to RLVR and agent harnesses. Fair not with respect to a protected attribute like race or gender, but fair to the agent you are evaluating.</p><p>Two concrete things are in scope of this article, and more will soon follow:</p><ul><li><p>The verifier inside an RLVR loop the function that takes a rollout as an input and produces a reward.</p></li><li><p>The agent harness: the tools, scaffolding, and protocol the agent acts through during rollouts, evals, and production.</p><p></p></li></ul><h3>What &#8220;fair to the agent&#8221; means</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AdMh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AdMh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!AdMh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!AdMh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!AdMh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AdMh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:67500,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/197301633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AdMh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!AdMh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!AdMh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!AdMh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6b39342-8e92-463f-9858-ebd3ef711f00_1920x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A capability is a repeatable ability of an agent to produce a desired outcome under specified conditions.</p><p>A fair eval is a measurement of a specific capability of the AI agent.</p><p>What is it not? It is not a measurement of outcomes which were not expected in the specified conditions.</p><p>Score differences between agents, or between training checkpoints should be explained by the capability under test. They should not be explained by nuisance factors:</p><ul><li><p>Train-set contamination.</p></li><li><p>Verifier bugs and overfit rubrics.</p></li><li><p>Environment drift between runs.</p></li><li><p>Benign paraphrases of the prompt.</p></li><li><p>Harness details the agent never sees in deployment.</p></li></ul><p>If an eval is sensitive to these, its not a fair eval.</p><h3>Some ways an eval can be unfair</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qioX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qioX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 424w, https://substackcdn.com/image/fetch/$s_!qioX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 848w, https://substackcdn.com/image/fetch/$s_!qioX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 1272w, https://substackcdn.com/image/fetch/$s_!qioX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qioX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png" width="1456" height="1098" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1098,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:270890,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/197301633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qioX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 424w, https://substackcdn.com/image/fetch/$s_!qioX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 848w, https://substackcdn.com/image/fetch/$s_!qioX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 1272w, https://substackcdn.com/image/fetch/$s_!qioX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a4d466-403e-4111-86be-4277ef568dd2_1579x1191.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We discuss four most-popular ways here, but this is non-exhaustive, and I will keep extending this as we find more gaps.</p><h4>1. Prompt Underspecification</h4><p>This happens if the task misses a constraint any reasonable solver would need to achieve the task meaningfully.</p><p>If the verifier reports a fail, the outcome is confounded with ambiguity of the instructions. Two equally strong agents can swing several points apart on the same task. If the instructions are ambiguous the agent should be rewarded equally for equally valid solution paths.</p><h4>2. Environment Issues</h4><p>The sandbox has a stale package, a non-responsive tool, or a filesystem layout that drifts between runs or setups. The agent&#8217;s first action fails for reasons it cannot inspect. Such a failure should be attributed to the environment.</p><h4>3. The harness design should be environment-centric</h4><p>The agent interacts with the environment through the harness. Tool schemas, observation formatting, retry policy, max steps, error messages, all shape behavior.</p><p>Two harnesses with the &#8220;same&#8221; tools can produce very different scores. For example:</p><ol><li><p>A truncated stderr hides the bug.</p></li><li><p>A misspecified tool schema can cause unnecessary confusion.</p></li><li><p>An unintentional 10-step cap might hamper a good agent&#8217;s planning capability.</p></li></ol><p>Therefore any fair eval should have a uniform eval across different</p><h4>4. Don&#8217;t ask it to do things it would not do in the real world</h4><p>Your environment, tasks and verifiers should not expect the agent to achieve a goal which is unrealistic in practice. A few common ones:</p><ul><li><p>Tasks the agent is told to refuse in production but expected to attempt in eval.</p></li><li><p>Tools available in the gym that do not exist in the deployed surface.</p></li><li><p>A persona or role the deployed system never adopts.</p></li></ul><p>Since if the agent were to hillclimb for these - it would not improve the capability you want to measure.</p><h4>The role the verifier plays</h4><p>The verifier is a model of &#8220;what success looks like.&#8221;</p><p>A strict verifier on an underspecified prompt punishes reasonable behavior. A lenient verifier can turn failures into passes. Verifiers are rarely audited as carefully as the agents they grade but their errors compound at every step of hill-climbing.</p><p>Treat the verifier as a system under test.</p><ul><li><p>Measure its agreement with humans.</p></li><li><p>Measure its variance across paraphrases of the same correct answer.</p></li><li><p>Measure its false-positive and false-negative rates.</p></li></ul><h3>Reward hacking is a fairness problem</h3><p>In RL, the verifier <em>is</em> the reward. Anything the verifier accepts is a valid policy.</p><p>If the verifier can be satisfied without solving the task meaningfully, the agent will eventually find that shortcut. This is not the model&#8217;s or the algorithm&#8217;s fault. It is the eval&#8217;s fault. A few common shortcuts:</p><ul><li><p>Producing answers in a format the verifier scores leniently.</p></li><li><p>Exploiting tool calls in a way that gets a good reward but doesn&#8217;t affect the state of the environment.</p></li><li><p>Pattern-matching the rubric&#8217;s expectation instead of producing a correct answer.</p></li><li><p>Outputting both the answer and its negation when the verifier checks for substring presence.</p></li></ul><p>Fair RL requires that the <em>only</em> cheap way to get reward is to do the task. If a cheaper path exists, the agent harness improvement loop or the gradient across the rollouts will discover it, and rightly so.</p><h3>A checklist for fair evaluations</h3><p>Before you let an eval drive decisions or training:</p><ul><li><p><strong>Specification.</strong> Could a competent intelligent entity solve the task from the prompt alone, without insider knowledge?</p></li><li><p><strong>Harness parity.</strong> Does the eval harness match the deployment harness on tools, formats, and limits?</p></li><li><p><strong>Distribution match.</strong> Are eval tasks ones the agent would actually face, and be permitted to attempt in production?</p></li><li><p><strong>Verifier audit.</strong> Has the verifier been graded against humans or SoTA model trajectories? What is its False Positive &amp; False Negative rate?</p></li><li><p><strong>Paraphrase invariance.</strong> Does the score change when the prompt is rewritten without changing meaning?</p></li><li><p><strong>Failure attribution.</strong> When the agent fails, can you tell whether it was the agent, the harness, the environment, or the verifier?</p></li></ul><h3>Open questions</h3><ul><li><p>How do you measure verifier quality when human labels are themselves noisy or expensive?</p></li><li><p>What is the right unit of &#8220;agent failure&#8221; in long-horizon tasks where many small slips compound?</p></li><li><p>Can verifiers be co-trained with agents without collapsing into a stationary state where both of them are poor?</p></li><li><p>How do you detect reward hacking inside an RLVR loop before it shows up as a deployment regression?</p></li><li><p>What is the minimum harness contract that should be held constant across rollout, eval, and production?</p></li><li><p>For tasks with no programmatic verifier, how do you keep model-graded RLVR from drifting into preference-style noise?</p></li></ul><p><em>If you&#8217;re shipping agents and model and have faced similar problems with fair evaluations, we should chat. We have a lot of interesting private and public evaluations which cover terminal use, MCP and computer use environments. Our focus is ensuring high-quality data. And fair-evaluations are a core tenet as we scale the horizons our agents operate on. </em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Whose Taste?]]></title><description><![CDATA[More data won't fix the AI verification problem. Different taste might.]]></description><link>https://blog.collinear.ai/p/whose-taste</link><guid isPermaLink="false">https://blog.collinear.ai/p/whose-taste</guid><dc:creator><![CDATA[Sachin]]></dc:creator><pubDate>Thu, 07 May 2026 16:02:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uHhf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>The man who passed every check</h2><p>Kim Philby, standing in his mother&#8217;s flat in Drayton Gardens, London, before an array of reporters, was guilty. He had been guilty since 1934, when a Soviet agent named Otto recruited him in Regent&#8217;s Park. He was guilty when he joined MI6 in 1940, and when the King awarded him an OBE in 1945 for his wartime intelligence work. And he was guilty in 1955, when the Foreign Secretary stood up in the House of Commons and cleared his name.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uHhf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uHhf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!uHhf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!uHhf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!uHhf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uHhf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Rare Film Emerges Of Double-Agent Kim Philby Speaking After Defection |  KPBS Public Media&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Rare Film Emerges Of Double-Agent Kim Philby Speaking After Defection |  KPBS Public Media" title="Rare Film Emerges Of Double-Agent Kim Philby Speaking After Defection |  KPBS Public Media" srcset="https://substackcdn.com/image/fetch/$s_!uHhf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!uHhf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!uHhf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!uHhf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08870806-fce6-4c76-9808-eab764fc0097_3057x1720.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Kim Philby lecturing Stasi agents in East Germany about his exploits to the top of MI6, while spying for the KGB | Harold Clements Getty Images</figcaption></figure></div><p>It was not until he had been posted to MI6 station in Beirut, when confronted by his old MI6 friend Nicholas Elliott, did Philby ever confess to spying for the Soviets, albeit partially. Shortly thereafter, he disappeared into the night onto a Soviet freighter bound for Russia, where he lived for the remainder of his life.</p><p>The signals about Philby had always been there. The Communist sympathies from his years at Cambridge, a first marriage to an Austrian communist. The system saw it all and, time and again, cleared him of any suspicion or wrongdoing.</p><p>This is not an article about mid-20th-century espionage. This is about what happens when a system has to verify something for which it has no reliable ground truth for, using signal that captures what to look for, but not how to weigh it.</p><div><hr></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>What this has to do with AI</h2><p>There&#8217;s a funny conversation happening in the AI community right now. Researchers and frontier labs are increasingly dismissive of humans&#8217; ability to verify or review long-horizon tasks, but are also insistent that human data is what unlocks the next 1000x in model capabilities. Both are probably true statements, but they point to two different problems getting collapsed into one.</p><p>The first is bandwidth. A human cannot sit through a four-hour agent trajectory, follow every tool call, and reliably verify the work to a high degree of accuracy. But this is a solvable problem, whether it be through better tooling, sampling, or even breaking longer tasks into smaller checks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!f8se!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!f8se!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 424w, https://substackcdn.com/image/fetch/$s_!f8se!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 848w, https://substackcdn.com/image/fetch/$s_!f8se!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 1272w, https://substackcdn.com/image/fetch/$s_!f8se!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!f8se!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png" width="829" height="412" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:412,&quot;width&quot;:829,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45569,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/196671775?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!f8se!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 424w, https://substackcdn.com/image/fetch/$s_!f8se!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 848w, https://substackcdn.com/image/fetch/$s_!f8se!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 1272w, https://substackcdn.com/image/fetch/$s_!f8se!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33aca7ca-de88-466d-baee-346698c9a7ba_829x412.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The second is criteria. In unverifiable domains, where there&#8217;s often no clean ground truth, two qualified experts can review the same output and disagree on whether it is good. Not because one of them is wrong, but because &#8220;good&#8221; is a judgment call, and judgment calls don&#8217;t average across annotators the way correctness does.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!a4bq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a4bq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 424w, https://substackcdn.com/image/fetch/$s_!a4bq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 848w, https://substackcdn.com/image/fetch/$s_!a4bq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 1272w, https://substackcdn.com/image/fetch/$s_!a4bq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a4bq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png" width="680" height="560" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:560,&quot;width&quot;:680,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:64083,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/196671775?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!a4bq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 424w, https://substackcdn.com/image/fetch/$s_!a4bq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 848w, https://substackcdn.com/image/fetch/$s_!a4bq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 1272w, https://substackcdn.com/image/fetch/$s_!a4bq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F86fb0193-de05-4c64-9aff-b86f0e98119a_680x560.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>This is the harder problem to solve. In unverifiable domains, the verifier is not approximating a ground truth because there isn&#8217;t one. The verifier is somebody&#8217;s weighting of criteria, which means in unverifiable domains, the verifier is taste.</p><h2>Taste isn&#8217;t one thing</h2><p>When most people talk about taste, they treat it as a single thing, but in reality, taste has two parts, each of which behave very differently.</p><p>The first is criteria. The line items and properties in a rubric that dictate when an agent&#8217;s output is considered good. For a legal brief, that might look like:</p><ul><li><p>A clear issue statement that frames the legal question</p></li><li><p>Citations to controlling authority</p></li><li><p>A coherent narrative spine that connects the facts to the legal argument</p></li></ul><p>Criteria is mostly binary, in that either the brief cites the right cases or it doesn&#8217;t. Either the brief is succinct or it isn&#8217;t.</p><p>On criteria, MI6 had Philby cold. The Communist circle at Cambridge, the first marriage to an Austrian communist, the Cambridge friends who had already defected to Moscow. All of it was in his file, surfaced and reviewed more than once. MI6 had the rubric, and it scored him on the rubric.</p><p>What it didn&#8217;t have was an answer to the next question. Confronted with a Cambridge man, the son of a celebrated Arabist, decorated in the war, vouched for by the right people, what do you do with the red flags? Every time MI6 ran the question, it answered the same way: down-weight them.</p><p>That next question, what to do with the line items when they coexist, conflict, or trade off against each other, is weighting. Weight is continuous, context-dependent, and the harder half of the problem of capturing taste.</p><p>For example, two experienced litigators can sit down with the same case and agree on every criteria. Both want a clear issue statement, both want the right authority cited, both want a tight narrative. But they will disagree on whether to file an aggressive version of the brief or the pared down version. The disagreement is about how to weigh the criteria against each other relative to a specific context: a specific judge, venue (e.g., local vs federal court), client&#8217;s risk tolerance, etc. The end result is you still have the same rubric, but different verdicts.</p><p>The human data machines today are quite good at the first half via pairwise preferences, rubric annotation, etc., but none of it captures the trade-off logic that turns items into a verdict.</p><h2>Why aggregation doesn&#8217;t rescue this</h2><p>The obvious objection at this point is that we already have a tool for this. RLHF works by collecting preference pairs from many annotators. If you collect enough of them, the model is supposed to learn the implicit weighting through statistical aggregation. Surely scale solves the problem then?</p><p>It doesn&#8217;t, because aggregation captures the median annotator&#8217;s weighting in the typical or average context, which, for unverifiable domains, is what you want to avoid.</p><blockquote><p>Value in unverifiable domains comes from non-median judgment applied to non-typical context. </p></blockquote><p>A great legal brief isn&#8217;t the median lawyer&#8217;s brief. A great research direction isn&#8217;t what most reviewers would pick. A good investment thesis is, almost by definition, one most people disagree with at the time of the trade. The whole point of bringing in expert judgment is to access the part of the distribution that consensus would smooth away.</p><p>When you average preferences across annotators and contexts, you smooth out the signal that distinguishes good from average judgment. You don&#8217;t end up with a verifier that approximates expert judgment. You end up with one that approximates consensus annotator judgment, and those are different things.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X1a4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X1a4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X1a4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X1a4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X1a4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X1a4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg" width="606" height="861" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:861,&quot;width&quot;:606,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;File:The Soviet Union 1990 CPA 6266 stamp (Soviet Intelligence Agents. Kim  Philby) small resolution.jpg - Wikimedia Commons&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="File:The Soviet Union 1990 CPA 6266 stamp (Soviet Intelligence Agents. Kim  Philby) small resolution.jpg - Wikimedia Commons" title="File:The Soviet Union 1990 CPA 6266 stamp (Soviet Intelligence Agents. Kim  Philby) small resolution.jpg - Wikimedia Commons" srcset="https://substackcdn.com/image/fetch/$s_!X1a4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X1a4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X1a4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X1a4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdde7a6eb-511b-4aa1-8005-b10fcef08ed8_606x861.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Soviet postage stamp paying homage to Kim Philby</figcaption></figure></div><p>Let&#8217;s imagine we ran this on approach on Kim Philby. A vetting model trained on the aggregated preferences of every MI6 officer in 1945 would have produced exactly the verdict the system produced: he&#8217;s one of us, the signals pointing the other way must be noise. More annotators wouldn&#8217;t have helped. The signal was in the minority weighting that aggregation washed out.</p><p>In unverifiable domains, the value of a verifier - taste - is precisely what aggregation destroys.</p><h2>The real question stops being &#8220;more data&#8221;</h2><p>So if aggregation is the wrong tool, the question stops being how to collect more preferences and starts being something harder. What are we actually trying to capture in unverifiable domains?</p><p>The hypothesis: somebody&#8217;s weighting of legally-defensible criteria, in specific contexts, captured at high enough fidelity that a verifier can apply it when the context shifts. That reframe forces two questions.</p><p>The first is whose weighting. Ideally, it would be experts whose judgment correlates with real, measurable outcomes in the domain in question. Obviously, that is far from simple.</p><p>Outcomes in unverifiable domains are often delayed, noisy, or unobservable. For example, reputation is a proxy and a weak one; it tracks visibility as much as judgment. Peer consensus often selects for orthodoxy, which is the opposite of what makes expert judgment valuable in the first place. Philby&#8217;s career is a textbook example of this. The Cambridge education, the war record, the OBE were all layers of peer consensus pointed the same way. The jugement that would have caught him was the orthogonal kind, which peer consensus is built to surpress.</p><p>None of this means expert selection is impossible. It means it&#8217;s a real problem that has to be solved, not an assumption that can be hand-waved past.</p><p>The second is how to capture it. Pairwise preferences flatten weighting into a single binary signal. Rubric scoring captures criteria but skips the trade-off logic. What you actually need is closer to reasoned disagreement at depth: experts not just picking the better output but explaining what they would have prioritized differently, in what context, and why. The trade-off logic itself becomes the training signal.</p><p>It should be said however, that this is dramatically harder to scale than what we have today. The whole appeal of pairwise preferences was that they were cheap and parallelizable. Weighting capture is neither. It requires more time, more expertise, and more thought per data point.</p><p>This is where approaches like throwing more data at the problem start to break down. You don&#8217;t need more data, you need different data, from different people, captured in different ways.</p><h2>The last mile keeps moving</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!j6c1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j6c1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 424w, https://substackcdn.com/image/fetch/$s_!j6c1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 848w, https://substackcdn.com/image/fetch/$s_!j6c1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 1272w, https://substackcdn.com/image/fetch/$s_!j6c1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j6c1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png" width="604" height="264" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:264,&quot;width&quot;:604,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:30455,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/196671775?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j6c1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 424w, https://substackcdn.com/image/fetch/$s_!j6c1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 848w, https://substackcdn.com/image/fetch/$s_!j6c1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 1272w, https://substackcdn.com/image/fetch/$s_!j6c1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e5e0d2-89d8-4340-9f40-da90eaffdb27_604x264.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Aaron Levie made a <a href="https://x.com/levie/status/2048950940661932319">point</a> recently that connects directly to this. As agents get better at the parts of a task they can already do well, the &#8220;last mile&#8221; - the part requiring human judgment to verify - keeps shifting up the value stack. The taste needed to verify a junior analyst&#8217;s output today is not the taste needed to verify a senior partner&#8217;s output tomorrow. </p><blockquote><p>The frontier of what counts as good keeps moving, because the floor of what agents can do keeps rising.</p></blockquote><p>What that means is, weighting capture isn&#8217;t simply a dataset problem. It is and will be, an evolving competency. Whatever weighting you capture today is correct for a domain that&#8217;s already moving, spurred on by AI agents. By the time the agent trained on it is deployed, the work that needs verifying has shifted, and the taste required to verify has shifted with it.</p><p>In unverifiable domains, you are not building a verifier once. You are building the capacity to keep capturing the right people&#8217;s weighting as the work that needs verifying changes underneath you.</p><p>The next 1000x in unverifiable domains doesn&#8217;t come from more taste in the data. It comes from the right people&#8217;s taste, captured at fidelity, as the frontier moves. Kim Philby walked out of his mother&#8217;s flat in Drayton Gardens in 1955 because the system optimized for the wrong taste. It had the criteria. It had the data. What it didn&#8217;t have was a way to weigh Cambridge against Moscow, before the consensus had already smoothed the difference away. The version of that problem facing AI is harder, because the frontier keeps moving and the answer keeps shifting with it. Whoever figures out how to solve the taste problem facing AI, will own it.</p><h2>Vetting the file</h2><p>At Collinear, we&#8217;re building part of it. SimLab is our simulation lab for AI agents: the infrastructure to generate, curate, and verify high-signal data at scale. Simulated enterprise environments, NPC users, verifiable tasks, training-ready rollouts. SimLab is built to capture expert weighting in context: not just whether an output is good, but the trade-off logic underneath the verdict, refreshed as the work itself shifts.</p><p>If you&#8217;re shipping AI in a domain where verification is hard, where median-annotator labels won&#8217;t get you there, and where consensus would smooth away exactly the signal you care about, we should talk. Every domain has its Philby. The work is building the system that doesn&#8217;t clear him.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.collinear.ai/book-a-demo&quot;,&quot;text&quot;:&quot;Talk to a Researcher&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.collinear.ai/book-a-demo"><span>Talk to a Researcher</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Collinear Newsletter #11 - Notes on Frontier AI ]]></title><description><![CDATA[Hi AI innovators,]]></description><link>https://blog.collinear.ai/p/collinear-newsletter-11-notes-on</link><guid isPermaLink="false">https://blog.collinear.ai/p/collinear-newsletter-11-notes-on</guid><dc:creator><![CDATA[Soumyadeep Bakshi]]></dc:creator><pubDate>Thu, 30 Apr 2026 22:06:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X0F1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi AI innovators,</p><p>April brought the people building frontier agents into the same rooms we work in. Two community events, plus the research direction those rooms kept pointing us toward.</p><h3><strong>NYC Builders at the Collinear Exec Dinner Series  </strong></h3><p>Our Collinear Exec Dinner Series brought together senior research leaders from Apple, IBM Research, Two Sigma, NVIDIA, Datadog Research, Wells Fargo, and Google for a candid social over Old Delhi kebabs. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X0F1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X0F1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X0F1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X0F1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X0F1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X0F1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg" width="406" height="541.2403846153846" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:406,&quot;bytes&quot;:2447001,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/195945304?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!X0F1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X0F1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X0F1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X0F1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b8d6ccd-51e6-45ef-95ba-141e61960286_3024x4032.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The conversation focused on three critical industry problems. </p><ol><li><p><strong>Building reliable evals</strong> for multi-turn, long horizon workflows</p></li><li><p><strong>Performance gap</strong> between offline traces and live trajectories, and</p></li><li><p><strong>What &#8220;fidelity&#8221; should mean</strong> for simulated environments meant to train agents on real-world tasks</p><p></p></li></ol><p>Our next edition of the Exec Dinner Series is coming up - DM us if you would like to be in the room.</p><h3><strong>Sim Fidelity for AI Agents - our Q2 Research Social</strong></h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Unnf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Unnf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Unnf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Unnf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Unnf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Unnf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg" width="488" height="650.6666666666666" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1152,&quot;resizeWidth&quot;:488,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;View image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="View image" title="View image" srcset="https://substackcdn.com/image/fetch/$s_!Unnf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Unnf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Unnf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Unnf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff5a13734-cd85-4ffe-bdc0-2717ea5d351d_1152x1536.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We took over the Sunnyvale office for a night of debate on simulation environments. The topic was the need for realism and high fidelity in RL simulations. </p><p>Researchers from GDM, NVIDIA, Apple, xAI and others slugged it out along with our special guest from MBZUAI, <strong><a href="https://www.linkedin.com/in/mikhail-yurochkin-a45659114/">Mikhail Yurochkin</a>. </strong></p><p>Subscribe to our mailing list if you would like to join the next Research Social. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.collinear.ai/subscribe?"><span>Subscribe now</span></a></p><p></p><h3><strong>NPCs - the Key to Replicating Real World Messiness</strong></h3><p>When we published <a href="https://arxiv.org/abs/2510.04491">TraitBasis</a> last year, our work on activation-steered behavioral traits, the next idea was obvious: put these NPCs inside the Simulation Lab itself.</p><p>Real-world enterprise workflows are messy. Coworkers interrupt, change their minds, update systems while you are mid-thought, and act on shared state without telling you. We wanted our simulations to behave the same way.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TIOk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TIOk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 424w, https://substackcdn.com/image/fetch/$s_!TIOk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 848w, https://substackcdn.com/image/fetch/$s_!TIOk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 1272w, https://substackcdn.com/image/fetch/$s_!TIOk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TIOk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png" width="1314" height="690" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:1314,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:262492,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/195945304?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TIOk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 424w, https://substackcdn.com/image/fetch/$s_!TIOk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 848w, https://substackcdn.com/image/fetch/$s_!TIOk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 1272w, https://substackcdn.com/image/fetch/$s_!TIOk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfcbdbfa-14b3-48c5-8c33-fc3536639d71_1314x690.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>So we kept building.</strong> NPCs with agency that take actions and change shared state. NPCs with secrets that the agent has to surface through the right questions. NPCs with limited context, including attention, forgetting, and prioritization across competing tasks. Each NPC nuance (and the related tasks and verifiers built around it) raises the fidelity of the simulation and the quality of the training signal coming out of it. Frontier models that were comfortably solving our environments a quarter ago are now hitting walls, and we are seeing marked improvement in downstream agent outcomes when training on this signal.</p><p>Learn more <a href="https://blog.collinear.ai/p/trait-basis">here</a> and stay tuned for a technical report on NPCs and related hillclimbing results.</p><h3><strong>Work with us</strong></h3><p>If you are training agents for enterprise workflows and want to stress-test them in a real Simulation Lab, <a href="https://www.collinear.ai/book-a-demo">book a demo</a>.</p><p>We are also hiring researchers and engineers who want to push the frontier of agent training environments. See open roles at <a href="https://www.collinear.ai/careers">collinear.ai/careers</a>.</p><p>That&#8217;s it for April. More on the research side coming soon.</p><p>Best, </p><p>The Collinear Team</p>]]></content:encoded></item><item><title><![CDATA[AI's U-235 Problem]]></title><description><![CDATA[Nuclear physics solved for k_eff. What's the AGI equivalent?]]></description><link>https://blog.collinear.ai/p/ais-u-235-problem</link><guid isPermaLink="false">https://blog.collinear.ai/p/ais-u-235-problem</guid><dc:creator><![CDATA[Jed Gresham]]></dc:creator><pubDate>Thu, 23 Apr 2026 18:09:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Sjz5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sjz5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sjz5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!Sjz5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!Sjz5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!Sjz5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sjz5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1100600,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/195170872?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Sjz5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!Sjz5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!Sjz5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!Sjz5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7fb190f-df91-48c7-b16d-7de32b82f4c7_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The race to AGI isn&#8217;t being won by whoever has the most compute or the cleverest architecture. It&#8217;s being won by whoever solves a quieter, less glamorous problem: finding enough of the right kind of data to cross the threshold. This is not a new concept.</p><p>On December 2, 1942, Enrico Fermi and a small team of physicists gathered in a makeshift lab beneath the stands of Stagg Field in Chicago. They had spent months carefully stacking graphite blocks and uranium slugs into a precise arrangement they called Chicago Pile-1. At 3:25pm, they slowly withdrew a control rod (a neutron-absorbing insert that regulates the reaction). The Geiger counters clicked faster. The reaction sustained itself. The nuclear age began when the clicking didn&#8217;t stop.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>What most people don&#8217;t realize is how close they came to never getting there. The core problem wasn&#8217;t theory. Leo Szilard had conceived of the nuclear chain reaction nearly a decade earlier, crossing a London street in 1933. He understood it so completely he quietly patented it and handed the rights to the British Admiralty to keep it out of dangerous hands. The physics was known and the threshold was understood. The problem they had was the fuel.</p><p>Natural uranium is everywhere. The earth&#8217;s crust is full of it. But raw uranium is almost useless for a chain reaction. The isotope that actually fissions, U-235, makes up less than 1% of natural ore. The rest is inert. Volume alone gets you nowhere. Too much of the wrong material actively works against you as it absorbs neutrons and dampens the reaction before it can sustain itself. The real breakthrough of the Manhattan Project wasn&#8217;t the bomb. It was Oak Ridge&#8217;s K-25 gaseous diffusion plant, built in 1944 to enrich uranium at industrial scale. At the time it was the largest building in the world. Its only job was to concentrate U-235: filtering, separating, amplifying the rare fissile material until there was enough to matter.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DYLK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DYLK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 424w, https://substackcdn.com/image/fetch/$s_!DYLK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 848w, https://substackcdn.com/image/fetch/$s_!DYLK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 1272w, https://substackcdn.com/image/fetch/$s_!DYLK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DYLK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png" width="800" height="586" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8e4bb6b1-672b-4df4-a740-675852396251_800x586.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:586,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:454776,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/195170872?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DYLK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 424w, https://substackcdn.com/image/fetch/$s_!DYLK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 848w, https://substackcdn.com/image/fetch/$s_!DYLK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 1272w, https://substackcdn.com/image/fetch/$s_!DYLK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e4bb6b1-672b-4df4-a740-675852396251_800x586.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Photograph of Stagg Field at the University of Chicago, Argonne National Laboratory archives (Argonne National Laboratory on <a href="http://flickr.com/">flickr.com</a>)</figcaption></figure></div><div><hr></div><h2>The Ore Gets Leaner</h2><p>Today&#8217;s AI models are trained on more data than any human could read in a thousand lifetimes. The models are still running into walls, and the reason maps almost exactly onto Oak Ridge. The reason some data moves models forward and most doesn&#8217;t comes down to what a model can actually learn from it. Training works by exposing a model to examples and having it predict what comes next, then correcting it when it&#8217;s wrong. The correction is where the learning happens. Boilerplate content, templated writing, repetitive programmatic output: these produce almost no correction signal. The model already knows what comes next.</p><p>High-signal data is different. It contains genuine reasoning, unexpected connections, nuanced judgment, edge cases the model hasn&#8217;t encountered. Every one of those is a correction opportunity. When the model is wrong, it gets updated and gets sharper.</p><p>U-235 atoms fission when struck by a neutron because of specific properties in their nuclear structure. Most uranium atoms absorb the neutron and go quiet. The difference between fissile and inert material is structural, not superficial. High-signal training data works the same way. Generic data absorbs the training pass and goes quiet.</p><p>Most new data being generated is programmatic: logs, auto-generated content, boilerplate, synthetic outputs from models that are already mediocre. It&#8217;s abundant, cheap, and mostly inert. As it floods in, it dilutes the fraction of genuinely useful signal further. As AI-generated content spreads across the internet, models trained on it learn the average, not the edge. The ore gets leaner the more we mine it.</p><p>Marie Curie didn&#8217;t find radium by sifting through more rock. She developed new processes to isolate and concentrate what she was looking for. The Manhattan Project built industrial infrastructure specifically designed to separate what mattered from what didn&#8217;t. The AI field needs the same shift.</p><div><hr></div><h2>Building Oak Ridge</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jnic!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jnic!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 424w, https://substackcdn.com/image/fetch/$s_!jnic!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 848w, https://substackcdn.com/image/fetch/$s_!jnic!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!jnic!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jnic!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg" width="960" height="608" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:608,&quot;width&quot;:960,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;File:Oak Ridge National Laboratory, Oak Ridge, Tenn (78285).jpg&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="File:Oak Ridge National Laboratory, Oak Ridge, Tenn (78285).jpg" title="File:Oak Ridge National Laboratory, Oak Ridge, Tenn (78285).jpg" srcset="https://substackcdn.com/image/fetch/$s_!jnic!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 424w, https://substackcdn.com/image/fetch/$s_!jnic!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 848w, https://substackcdn.com/image/fetch/$s_!jnic!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!jnic!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba0ef85f-88c1-406e-9bad-68873d582be6_960x608.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Post card showing Oak Ridge National Laboratory, Oak Ridge, Tenn</figcaption></figure></div><p>Two things have to happen in parallel. The first is better curation pipelines: new methods to identify and extract high-signal data from existing sources, smarter filtering, better labeling, clearer definitions of what &#8220;exceptional&#8221; looks like for each capability domain. The second is synthetic data. Even though the risk of model collapse (the equivalent of contaminating your fuel) is real, waiting for enough naturally occurring high-signal data won&#8217;t get us where we&#8217;re trying to go. Not all synthetic data techniques are the same, and the differences between them matter a lot. Deliberately designed training data built to fill specific capability gaps is unavoidable.</p><p>The most straightforward approach is prompt-based generation: feed a capable model a topic, a domain, or a problem type and ask it to generate training examples at scale. Used carefully, this fills gaps in rare or underrepresented domains. Used carelessly, it produces plausible-sounding noise that makes the training pool worse, not better.</p><p>A more sophisticated method is web rewriting: take real content and use a stronger model to transform it into a higher-signal format. DeepSeek did this systematically building <a href="https://arxiv.org/abs/2412.19437">DeepSeek-V3</a>. Rather than training on web content directly, they used a pipeline of stronger models to filter, rewrite, and structure data into higher-quality reasoning examples across mathematics, code, and general knowledge. The resulting model matched or outperformed models trained at several times the compute cost, with minimal changes to the underlying architecture. This is an example of using better fuel rods in the same reactor.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-FI3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-FI3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 424w, https://substackcdn.com/image/fetch/$s_!-FI3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 848w, https://substackcdn.com/image/fetch/$s_!-FI3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 1272w, https://substackcdn.com/image/fetch/$s_!-FI3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-FI3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png" width="1456" height="851" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:851,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Refer to caption&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Refer to caption" title="Refer to caption" srcset="https://substackcdn.com/image/fetch/$s_!-FI3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 424w, https://substackcdn.com/image/fetch/$s_!-FI3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 848w, https://substackcdn.com/image/fetch/$s_!-FI3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 1272w, https://substackcdn.com/image/fetch/$s_!-FI3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f21785e-f594-47e9-8e5a-111142f19fca_1661x971.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Benchmark performance of DeepSeek-V3 and its counterparts (https://arxiv.org/pdf/2412.19437)</figcaption></figure></div><p>A third approach is reinforcement learning from human feedback (RLHF): instead of generating new data from scratch, use human preference signals to identify which model outputs were high quality and train on those. This turns the model&#8217;s own outputs into enriched fuel, but only when paired with careful human judgment about what &#8220;better&#8221; means. <a href="https://arxiv.org/abs/2502.13417">Recent work</a> has pushed this further, with new techniques achieving alignment quality comparable to full human annotation using only 6-7% ( 6 7!!! ) of the annotation effort, by targeting human review at the samples that are hardest to label automatically.</p><p>The most surprising recent result is pure reinforcement learning with no labeled data at all. <a href="https://arxiv.org/abs/2501.12948">DeepSeek&#8217;s R1</a> demonstrated that reasoning capabilities can emerge through pure reinforcement learning, with no human-labeled reasoning trajectories required. The model was rewarded for getting verifiable answers right (math problems, code that actually runs) and developed self-reflection and strategy as emergent behavior. It&#8217;s the closest thing yet to a model generating its own fissile material.</p><p>At <a href="http://Collinear.ai">Collinear AI</a> we realize the theory isn&#8217;t the bottleneck. The hard part is making these techniques reliable at scale: turning prompt-based generation, web rewriting, and RLHF from one-off research sprints into something teams can run repeatedly without rebuilding from scratch each time. Most organizations treat each data effort as a custom project. <a href="https://github.com/collinear-ai/simlab">SimLab</a> is how Collinear makes this routine.</p><p>The contamination risk is real and worth understanding. <a href="https://arxiv.org/abs/2305.17493">Shumailov (and others)</a> demonstrated that repeatedly training on synthetic data leads to model collapse, a finding that attracted significant attention given how close current models are to exhausting available high-quality data. The mechanism: recursive training on synthetic outputs causes models to produce repetitive, narrowing results, effectively losing the tails of the original data distribution. The model gets so good at the average that it loses the edges.</p><p>The answer isn&#8217;t to avoid synthetic data. <a href="https://arxiv.org/abs/2404.01413">Research shows</a> that keeping real data in the mix and layering synthetic data on top, rather than replacing real data entirely, avoids the degenerative feedback loop. The ratio and sequencing matter enormously. Synthetic data added to a real-data foundation behaves very differently from synthetic data trained on top of synthetic data. The centrifuge has to be calibrated, not just built.</p><div><hr></div><h2>Critical Mass</h2><p>Before fission, energy was extractive. You burned coal, oil, or gas: feed the furnace, get energy out, repeat. Power was linear, bounded by what you could mine and move.</p><p>Fission changed that. A reaction releases neutrons that trigger more reactions. Above critical mass it&#8217;s self-sustaining: withdraw the control rod once and the chain continues without further input. A kilogram of enriched uranium and a kilogram of coal aren&#8217;t on the same spectrum. Critical mass is the threshold where the system stops depending on external input and begins feeding itself.</p><p>Every AI model today is extractive in the same way pre-fission energy was. We supply the data, the compute, the architecture decisions, the fine-tuning, the human feedback. The model improves because we keep feeding it. Remove the input, the improvement stops.</p><p>AGI is the point where that changes. A system past the AGI threshold can reason about its own limitations, identify what it needs to learn, generate or seek out the training signal it needs, and improve its own architecture. The chain reaction sustains itself.</p><p>The gap between today&#8217;s best models and that threshold is categorical, the same way fission and combustion aren&#8217;t variations of the same phenomenon. Current models are extraordinarily capable combustion engines. Once a system crosses that line, the rate of improvement stops being limited by what we supply. It becomes limited by the physics of the system itself. Human researchers won&#8217;t be the primary driver anymore. The reaction moves at its own pace, on its own terms. Getting the enrichment right before we get there is the only real leverage point we have.</p><div><hr></div><h2>Two Races to the Same Threshold</h2><p>The world is currently running two parallel races to the same kind of threshold, and almost nobody talks about them together.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cq4E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cq4E!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 424w, https://substackcdn.com/image/fetch/$s_!cq4E!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 848w, https://substackcdn.com/image/fetch/$s_!cq4E!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 1272w, https://substackcdn.com/image/fetch/$s_!cq4E!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cq4E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png" width="1456" height="964" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:964,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4670228,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/195170872?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cq4E!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 424w, https://substackcdn.com/image/fetch/$s_!cq4E!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 848w, https://substackcdn.com/image/fetch/$s_!cq4E!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 1272w, https://substackcdn.com/image/fetch/$s_!cq4E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2afa5f6-8c54-4eb5-9ca1-10bc144903d2_1875x1242.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Aerial drone image of the 500 MW ITER international project under construction in Cadarache, France</figcaption></figure></div><p>The first race is to fusion. Thirty-five nations are collaborating on ITER, the international fusion project in southern France. It&#8217;s been delayed repeatedly and now targets 2039. Private companies are moving faster, but even the boldest credible estimates put commercial fusion in the early 2030s at best. The physics is understood. The engineering is the hard part.</p><p>The second race is to AGI. Both have the same core structure: a threshold that is theoretically understood, a fuel problem that is practically unsolved, and a lot of engineering standing between the two. And the two races may be more entangled than they first appear.</p><p>The race to build AGI may literally require the fusion race to progress first, or at least the fission one. The energy demands make that dependency increasingly hard to ignore.</p><p>AI data centers are already consuming power at a scale that strains the grid. In 2024, global data center electricity consumption hit around 415 terawatt-hours (about 1.5% of world electricity use), growing at a rate more than four times faster than overall global electricity consumption. &#65532; By some estimates, that figure could <a href="https://www.iea.org/reports/energy-and-ai/energy-demand-from-ai">approach 945 TWh by 2030</a> (roughly equivalent to Japan&#8217;s entire annual electricity consumption) with high-growth scenarios pushing past 1,700 TWh by 2035.</p><p>That&#8217;s before AGI. A self-sustaining AI system iterating on itself continuously would require orders of magnitude more compute than today&#8217;s training runs. The chain reaction doesn&#8217;t just need enriched fuel. It needs an enormous, uninterrupted power supply to keep running once it starts. The physical energy problem and the cognitive threshold problem are coupled. Nuclear physics even has a name for it.</p><p>k_eff measures whether a chain reaction is self-sustaining. k_eff &lt; 1 and the reaction fizzles. k_eff &#8805; 1 and it runs on its own.</p><p>AGI doesn&#8217;t have an equivalent metric yet. But if it did, it might look something like this: does each generation of capability produce enough leverage to fund, power, and build the next one? We can call it a_eff. Right now, the honest answer is that we don&#8217;t know if a_eff &#8805; 1, and most of the serious debates in AI (about scaling laws, compute returns, energy constraints) are really arguments about that number without naming it.</p><div><hr></div><h2>The Clock</h2><p>Six years ago, the median expert estimate for AGI sat comfortably in the 2060-2070 range. As of early 2026, that number has collapsed to around 2033. &#65532; The compression is accelerating. Dario Amodei said at Davos earlier this year that AGI will likely arrive within a few years, possibly by 2027. Demis Hassabis of Google DeepMind put it more cautiously: roughly a 50% chance by 2030. &#65532;</p><p>Fusion timelines haven&#8217;t moved the same way. ITER&#8217;s deuterium-tritium milestone is 2039. Commercial fusion power is likely a decade beyond that. The private sector is more aggressive, but even the boldest credible estimates put sustained commercial fusion in the early 2030s at best. Both thresholds represent the same kind of categorical shift: a self-sustaining reaction that permanently changes what&#8217;s possible. Fusion solves the physical energy problem and AGI solves the cognitive one, but both require getting the enrichment right before the reaction will hold.</p><p>The current trajectory suggests AGI crosses its threshold first, probably by a significant margin. That means the data enrichment problem is urgent in a way that plasma confinement simply isn&#8217;t. Fusion researchers have until the 2030s. The people working on data quality and synthetic enrichment for AI may have considerably less time, and far less certainty about when the window closes.</p><p>AGI isn&#8217;t blocked by a missing theoretical insight. Szilard had that moment crossing a London street in 1933. It isn&#8217;t blocked by compute or architecture alone. It&#8217;s blocked by the same thing that stood between Szilard&#8217;s patent and Fermi&#8217;s reaction: not enough of the right material, concentrated precisely enough, arranged carefully enough to sustain itself.</p><p>Fermi&#8217;s team stacked blocks for months and did the math until the geometry was right and the reaction kept going.</p><p>That&#8217;s what we&#8217;re building toward. Not a dramatic moment, but a controlled, deliberately constructed threshold where the system crosses over and begins to sustain its own improvement. We need an AI Oak Ridge.</p><p>That work is happening now, in pieces, across a lot of teams. If you're at a frontier lab or AI-native company working on model improvement (capability gaps, post-training data, pre-deployment testing), talk to one of our researchers.</p><h2><strong>Building the Centrifuge</strong></h2><p>Oak Ridge took two years and 24,000 workers to separate enough U-235 for the reactor to hold. The enrichment problem for AI is a similar order of undertaking, and no single team will solve it.</p><p>At Collinear, we&#8217;re building part of it. SimLab is our simulation lab for AI agents: the infrastructure to generate, curate, and verify high-signal data at scale. Simulated enterprise environments, NPC users, verifiable tasks, training-ready rollouts. It&#8217;s designed to make deliberate enrichment a routine capability for the teams that need it most.</p><p>If you&#8217;re training reasoning models, shipping agents into production, or watching half your team&#8217;s time disappear into eval data hygiene, we should talk.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.collinear.ai/book-a-demo&quot;,&quot;text&quot;:&quot;Talk to a Researcher&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.collinear.ai/book-a-demo"><span>Talk to a Researcher</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[SimLab: The self-serve staging playground for real-world agents ]]></title><description><![CDATA[Agents fail on real tool calls, long workflows, and messy data. SimLab lets you find those failures in simulation, not in production.]]></description><link>https://blog.collinear.ai/p/simlab-the-self-serve-staging-playground</link><guid isPermaLink="false">https://blog.collinear.ai/p/simlab-the-self-serve-staging-playground</guid><dc:creator><![CDATA[Sachin]]></dc:creator><pubDate>Thu, 02 Apr 2026 22:01:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/56774894-85d4-4125-a1a3-38af8ae9ab8a_1216x864.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eS5N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eS5N!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!eS5N!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!eS5N!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!eS5N!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eS5N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png" width="1200" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:415723,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/192905633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eS5N!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 424w, https://substackcdn.com/image/fetch/$s_!eS5N!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 848w, https://substackcdn.com/image/fetch/$s_!eS5N!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 1272w, https://substackcdn.com/image/fetch/$s_!eS5N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5283b36d-75a1-4dcd-a6d9-4a6262b39913_1200x630.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>TL;DR</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><ul><li><p>Agents fail in production because evals test outputs, not behavior across stateful multi-step workflows.</p></li><li><p>The failure modes that actually matter: imperfect tool calls, state drift, and no-exit loops that only show up when the agent runs inside a realistic environment.</p></li><li><p>Software has staging. Agents have nothing between evals and prod. Simulation environments help fill the gap.</p></li><li><p>SimLab is a self-serve CLI that gives you the full stack: tasks, realistic environments, and deterministic verifiers meaning you find failures before your users do.</p></li></ul><div><hr></div><h3><strong>The agent passed evals. It worked in the demo. You shipped it. Then it broke.</strong></h3><p>You&#8217;ve probably seen this, or something similar in production. We&#8217;ll use customer support as an example. A customer asks for a simple account update. Everything looks good until step 3 of a 12-step workflow.</p><p>A tool call fires: the right function name, the wrong schema: the API returns a 422, the agent retries with the same payload, and now you&#8217;re in a silent loop. The failure only surfaces when a real user hits it, and it&#8217;s nearly impossible to reproduce from logs alone. The failure isn&#8217;t just an agent failure, but a support interaction failure, a degraded customer experience.</p><p><strong>This isn&#8217;t a model problem. It&#8217;s a testing problem.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!75_2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!75_2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!75_2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!75_2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!75_2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!75_2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!75_2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!75_2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!75_2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!75_2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00267cd5-27c2-4b06-ba25-77c1ba258ae9_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>Traditional evals were built for a different problem.</strong></h3><p>Traditional evals were built for single-turn, input-output tasks. They work for language problems. But this isn&#8217;t a typical single output failure. It&#8217;s a control loop failure. At each step it reads context, picks a tool or action, executes, observes the result, updates state, and decides what to do next. Step 3 of the workflow isn&#8217;t where it breaks, it&#8217;s where the compounding mistakes start. Not just for you, but for the customer interacting with your agent.</p><p>The bugs that ship to production:</p><ul><li><p><strong>Irregular tool call arguments. </strong>Right function, bad payload. Fails schema validation. Retries with the same bad payload.</p></li><li><p><strong>Silent state drift. </strong>Working context diverges from ground truth mid-workflow. Each subsequent step compounds it. By step 8, the agent is making decisions on data it no longer has right.</p></li><li><p><strong>Incomplete reasoning chains and No-exit loops. </strong>Hits a dead end, has no recovery path, retries the same action. No timeout, no fallback, no escalation.</p></li></ul><h3><strong>Evals test outputs. They don&#8217;t test behavior.</strong></h3><p>LLM-as-judge makes this worse. You&#8217;re using a nondeterministic model to grade a nondeterministic system. The reward signal is noisy, hard to act on, and doesn&#8217;t scale to the thousands of rollouts you need to improve the agent.</p><h4><strong>Your deployment pipeline has a gap.</strong></h4><p>Software engineers don&#8217;t push from local to prod. The pipeline is develop &#8594; test &#8594; staging &#8594; prod. Staging isn&#8217;t a perfect replica of production, but it&#8217;s close enough for most failures to surface before a user sees them.</p><p>The agent development pipeline today: build &#8594; evals &#8594; prod. No staging equivalent. The first time your agent hits a live API with real latency, a real 200-step workflow, or a user input outside your eval distribution is the first time a real user hits it too.</p><p>When something breaks, you&#8217;re debugging from logs and likely dealing with an unhappy customer on the other end. You can see your agent failed, but you usually can&#8217;t reproduce it, and you can&#8217;t run a thousand variations of the failing scenario to understand where the boundary is.</p><p>Simulation is the staging layer for agents. It closes the gap between &#8220;it passed evals&#8221; and &#8220;it actually did what we needed it to protect brand integrity and increase customer support satisfaction.&#8221;</p><div><hr></div><h3><strong>What a simulation environment actually needs.</strong></h3><p>Not a bigger dataset. Not a fancier benchmark. It needs&#8230;Real. World. Scenarios.</p><p><strong>A simulation environment has to let your agent interact with a realistic world across a full task execution trace and give you deterministic </strong><em><strong>and </strong></em><strong>programmatic signal about what happened.</strong></p><p><strong>Three things you can&#8217;t skip:</strong></p><ul><li><p><strong>Environments. </strong>Those that mirror real world scenarios, real customer support interactions. APIs with real failure modes (rate limits, distorted responses, timeouts, unexpected nulls), messy seeded data (incomplete records, conflicting field values, schema mismatches), and NPCs that behave imperfectly, like real users (ambiguous or incomplete requests, assumption breaking points, and task pivots).</p></li><li><p><strong>Tasks: </strong>Long-horizon, multi-step workflows that reflect real production complexity. Tasks that require 50&#8211;200 steps, involve ambiguous intermediate states, and have more than one valid execution path. The kind your agent will actually face when someone submits a customer support request.</p></li><li><p><strong>Verifiers: </strong>Deterministic, programmatic checks not LLM-as-judge. Did the agent reach the right end state for the customer? Did it complete all required steps? Did it stay within operational constraints? Consistent signal you can trust across thousands of parallel rollouts.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PJ0j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PJ0j!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 424w, https://substackcdn.com/image/fetch/$s_!PJ0j!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 848w, https://substackcdn.com/image/fetch/$s_!PJ0j!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 1272w, https://substackcdn.com/image/fetch/$s_!PJ0j!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PJ0j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png" width="1456" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:118448,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/192905633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PJ0j!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 424w, https://substackcdn.com/image/fetch/$s_!PJ0j!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 848w, https://substackcdn.com/image/fetch/$s_!PJ0j!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 1272w, https://substackcdn.com/image/fetch/$s_!PJ0j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a022f24-4892-4446-b3d4-92a82f04818f_1526x778.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><strong>What your dev loop looks like with simulation.</strong></h3><p style="text-align: center;"><strong>Instead of: build &#8594; ship &#8594; debug from logs.</strong></p><p style="text-align: center;"><strong>It becomes: build &#8594; simulate &#8594; fix &#8594; simulate &#8594; ship.</strong></p><p>You run thousands of parallel rollouts across your task distribution and see where failures occur, which steps, which tool interactions, which input types cause breakdowns. When something fails, you get an execution trace you can inspect step-by-step, tweak the environment around, and re-run. You iterate on behavior, tool call logic, error recovery, and context window management in hours, not weeks. And start finding capability gaps you didn&#8217;t know to look for.</p><p><strong>You stop guessing how your agent behaves. With SimLab, you get to see how it behaves.</strong></p><div><hr></div><h3><strong>We built SimLab to do this.</strong></h3><p>We kept running into this gap ourselves. Building simulation infrastructure from scratch is a serious investment. The task generation, realistic data, tool simulators, NPC behavior models, sandboxed execution, a deterministic eval layer, and so on. Most teams end up with brittle, domain-specific systems that break the moment the agent or task changes.</p><p><strong>SimLab is a self-serve CLI that gives you the full environment stack without live environment risk.</strong></p><ul><li><p><strong>Sandboxed execution. </strong>Agents run in isolated containers with full environment control. Arbitrary code execution, configurable tool access, reproducible state. You define what the agent can touch.</p></li><li><p><strong>Bring your own tools or use pre-built simulators. </strong>The platform is self-serve. Connect your own APIs and tool schemas, or use out-of-the-box simulators for common systems like Workday, Salesforce, and others.</p></li><li><p><strong>Programmatic Task and Verifier generation. </strong>Generate long-horizon tasks calibrated to your domain. Tune difficulty, workflow length, ambiguity, and edge case density. High-quality training signal, not clean-room benchmarks.</p></li><li><p><strong>Programmatic Data:</strong> Seeded data and NPC behavior models simulate real production messiness: bad inputs, missing fields, unexpected response formats.</p></li></ul><div><hr></div><h3><strong>Where this fits in your stack.</strong></h3><p>SimLab sits between build and deploy. It&#8217;s not a replacement for evals or observability; it&#8217;s the layer missing between them.</p><p>Evals tell you if individual outputs are correct. SimLab tells you if the agent can complete full workflows under realistic conditions. Observability gives you post-deployment traces of what already broke. SimLab gives you pre-deployment traces of what would have broken.</p><p>Human QA pipelines are slow, expensive, and don&#8217;t scale. You can also build your own simulation infra, but a full stack covering tasks, environment, verifiers, sandboxing, and parallelization is a significant engineering investment that gets brittle fast.</p><p><strong>SimLab is designed to be adaptable as your agents and domains change, without rebuilding from scratch each time.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K2XI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K2XI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 424w, https://substackcdn.com/image/fetch/$s_!K2XI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 848w, https://substackcdn.com/image/fetch/$s_!K2XI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 1272w, https://substackcdn.com/image/fetch/$s_!K2XI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K2XI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png" width="1456" height="478" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:478,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:272360,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/192905633?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K2XI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 424w, https://substackcdn.com/image/fetch/$s_!K2XI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 848w, https://substackcdn.com/image/fetch/$s_!K2XI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 1272w, https://substackcdn.com/image/fetch/$s_!K2XI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa377a337-bb36-473d-8a1e-ba6d2faff95d_2616x858.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h3><strong>Simulation is the new deployment.</strong></h3><p>Simulation will be a standard layer in the agent development pipeline, the same way CI and staging are standard in software. <strong>Not just a nice-to-have at scale, but the thing separating agents that demo well from agents that ship reliably.</strong></p><p>We&#8217;re opening SimLab as a self-serve CLI. Install it, point it at your agent, define an environment, run it. See where it breaks. We&#8217;re still testing the task generators, verifier primitives, and environment tooling and we want to know what doesn&#8217;t work for your use case.</p><p><strong>Try it. See how your agent holds up. Tell us what&#8217;s missing.</strong></p><p><strong>&#8594; <a href="https://github.com/collinear-ai/simlab">github.com/collinear-ai/simlab</a></strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts..</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Collinear Newsletter #10 - Notes on Frontier AI]]></title><description><![CDATA[Happy March from the Collinear team!]]></description><link>https://blog.collinear.ai/p/collinear-newsletter-10-notes-on</link><guid isPermaLink="false">https://blog.collinear.ai/p/collinear-newsletter-10-notes-on</guid><dc:creator><![CDATA[Soumyadeep Bakshi]]></dc:creator><pubDate>Tue, 24 Mar 2026 15:03:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AoJQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Happy March from the Collinear team.</p><p>It&#8217;s been a big quarter. We moved into a new office in Sunnyvale, launched YC Bench, extended the Simulation Lab, and hit the conference circuit. A lot to cover, so let&#8217;s get into it.</p><div><hr></div><h3><strong>New home in Sunnyvale</strong></h3><p>We outgrew our old space and opened a new office in Sunnyvale. A big thanks to the customers, partners, and friends who stopped by during the first few weeks to check it out. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kFUu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kFUu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 424w, https://substackcdn.com/image/fetch/$s_!kFUu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 848w, https://substackcdn.com/image/fetch/$s_!kFUu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 1272w, https://substackcdn.com/image/fetch/$s_!kFUu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kFUu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic" width="543" height="407.25" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:543,&quot;bytes&quot;:1853062,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/heic&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/190807647?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kFUu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 424w, https://substackcdn.com/image/fetch/$s_!kFUu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 848w, https://substackcdn.com/image/fetch/$s_!kFUu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 1272w, https://substackcdn.com/image/fetch/$s_!kFUu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8cfe5b8-59db-40e7-94e3-3e379c940cc7.heic 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The sign is up, the whiteboards are full, and the espresso machine is already earning its keep. Good to have a home base.</p><div><hr></div><h3><strong>YC Bench: can frontier models run a startup?</strong></h3><p>We released <a href="https://x.com/CollinearAI/status/2027531502234570768?s=20">YC Bench</a>, the first open-source, long-horizon benchmark with a simulation clock.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4S4r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4S4r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4S4r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4S4r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4S4r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4S4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!4S4r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4S4r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4S4r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4S4r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69d0719e-3ad9-4348-b47f-e644fb5d86d8_1920x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The idea:</strong> give a frontier model seed capital, a small team, and a market of tasks. Ask it to run an AI startup. Manage employees, hit deadlines, allocate resources, and maximize profit over time.</p><p>What we found is that a simple rule-based agent consistently outperforms frontier LLMs. Not because the task is impossible for them, but because they make compounding mistakes early on that they never recover from. They chase short-term wins, over-parallelize, and adapt too late when conditions change.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Eumu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Eumu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Eumu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Eumu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Eumu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Eumu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg" width="1456" height="866" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:866,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!Eumu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Eumu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Eumu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Eumu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a09f05e-1033-4eba-b10b-0ca116455e3a_1778x1058.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This matters because <strong>the industry is moving fast toward long-running, multi-step agent workflows</strong>. But reliability isn&#8217;t keeping pace with this ambition. YC Bench measures exactly that gap: not whether a model can answer a question, but whether an agent can hold a coherent strategy over time.</p><p>YC Bench is open-source and on our GitHub. We built it to be extensible. If you&#8217;re working on long-horizon agent evaluation, <a href="https://github.com/collinear-ai/yc-bench">we&#8217;d love to hear what you find</a>!</p><div><hr></div><h3><strong>Simulation Lab: what it is and what&#8217;s new</strong></h3><p>For those new here: the Collinear Simulation Lab is where AI agents learn enterprise work before going to production. Think of it as a practice environment, fully interactive, with simulated data, simulated users (NPCs), real enterprise tooling, and task-specific verifiers. Agents don&#8217;t just get tested. They get trained against realistic, messy, multi-step workflows.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vFyw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vFyw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 424w, https://substackcdn.com/image/fetch/$s_!vFyw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 848w, https://substackcdn.com/image/fetch/$s_!vFyw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 1272w, https://substackcdn.com/image/fetch/$s_!vFyw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vFyw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png" width="1170" height="712" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:712,&quot;width&quot;:1170,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:98164,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/190807647?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vFyw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 424w, https://substackcdn.com/image/fetch/$s_!vFyw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 848w, https://substackcdn.com/image/fetch/$s_!vFyw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 1272w, https://substackcdn.com/image/fetch/$s_!vFyw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e723faf-a331-4d96-8196-3fb2fef942fd_1170x712.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Inside a sim lab, you get:</p><ul><li><p>Simulated APIs for enterprise software (HR, Finance, Sales, Customer Service)</p></li><li><p>NPCs that push back, change their minds, and interrupt</p></li><li><p>Tasks with ambiguity, missing info, and competing priorities</p></li><li><p>Scorers that generate task-specific rubrics alongside formal verifiers</p></li></ul><p>We&#8217;ve extended the lab this quarter with broader tool coverage and deeper scenario complexity. More enterprise surfaces, richer NPC behavior, and tighter integration with RL training loops. The goal stays the same: agents need a world to learn in, and we&#8217;re building that world.</p><p>If you&#8217;re building agents for enterprise workflows and want to stress-test them before they touch production, the Simulation Lab is open for business.</p><div><hr></div><h3><strong>Together AI Conference</strong></h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AoJQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AoJQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 424w, https://substackcdn.com/image/fetch/$s_!AoJQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 848w, https://substackcdn.com/image/fetch/$s_!AoJQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!AoJQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AoJQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!AoJQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 424w, https://substackcdn.com/image/fetch/$s_!AoJQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 848w, https://substackcdn.com/image/fetch/$s_!AoJQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!AoJQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed4a3d69-499e-4992-86dc-4c28e163df44_4096x2731.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Nazneen spoke at the Together AI conference this quarter. Our partnership with Together continues to deepen. <a href="https://blog.collinear.ai/p/trait-basis">TraitBasis</a>, our method for generating realistic simulated users, is now integrated into Together Evals. Builders on Together&#8217;s platform can simulate impatient, confused, or inconsistent user personas and see how their models actually hold up when conversations get unpredictable. If you missed the talk, stay tuned for a recap.</p><div><hr></div><p>That&#8217;s it for March. More coming soon on the research side.</p><p>-- The Collinear Team</p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[We gave Claude, Gemini and GPT, $250k, and it didn't go as you’d expect...]]></title><description><![CDATA[Introducing YC Bench: The first open-source, long-horizon benchmark with a simulation clock]]></description><link>https://blog.collinear.ai/p/we-gave-claude-gemini-and-gpt-250k</link><guid isPermaLink="false">https://blog.collinear.ai/p/we-gave-claude-gemini-and-gpt-250k</guid><dc:creator><![CDATA[Muyu]]></dc:creator><pubDate>Thu, 05 Mar 2026 17:02:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ed2a6f3a-6d9c-475e-aea8-2a483997f2be_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>TL;DR</strong>: We find that frontier AI agents struggle on the YC Bench compared to other time-simulated benchmarks, such as the Vending Bench 2, highlighting capability gaps in planning and resource allocation for real-world scenarios.</p><p>Get started with YC bench:</p><pre><code><code>curl -sSL &lt;https://raw.githubusercontent.com/collinear-ai/yc-bench/main/start.sh&gt; | bash </code></code></pre><p>GitHub: <a href="https://github.com/collinear-ai/yc-bench">collinear-ai/yc-bench</a></p><div><hr></div><h1><strong>Long-Term Coherence as a Goal Post</strong></h1><p>Popular agent benchmarks - <a href="https://huggingface.co/gaia-benchmark">GAIA</a>, <a href="https://www.tbench.ai/">TermBench</a>, SWE-Bench, &#964;&#178;-bench - evaluate a model&#8217;s ability to complete tasks through multi-tool, multi-turn interactions. Even when these tasks span hundreds of tool calls, they share a critical limitation - they lack a <em>simulation clock</em>. As AI agents get integrated into the workforce and digital economy, time becomes an essential dimension of evaluation. Task sequence matters; environmental states shift over time, and inaction is as consequential as a wrong action. Without a clock running through the simulator, you can measure whether an agent did the right thing, but not whether it did it when it mattered.</p><p>Some existing benchmarks include a simulation clock, <a href="https://andonlabs.com/evals/vending-bench-2">Vending Bench</a>, which tests for long-term coherence<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> is one such example. Vending Bench has realistic NPCs, including suppliers the models talk to, and models compete with each other. The environment dynamics do not capture more sophisticated planning capabilities when resources are non-stationary, and there are tight deadlines to meet.</p><p>We propose a new <em>long-horizon adaptive planning</em> and <em>coherence</em> benchmark, YC-Bench (Your Company Bench), in which your agent takes on the role of a startup founder and executes activities to run a successful business. These activities include task prioritization, task scheduling, meeting client deadlines, resource allocation, managing burnout, and maximizing company profits and prestige.</p><p>With a simulation clock, the AI agent must learn to maximize long-term rewards over short-term gains. This means it needs to discriminate between short-term and long-term rewards, knowing that some actions pay less in the short term, but more in the long term. This is a core human skill usually termed &#8220;long-term coherence&#8221; that models are not exhaustively tested on, especially in a reproducible, open-source, and extensible way. YC-bench is an effort in this direction. </p><div><hr></div><h2><strong>Environment Dynamics</strong></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n9_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n9_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 424w, https://substackcdn.com/image/fetch/$s_!n9_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 848w, https://substackcdn.com/image/fetch/$s_!n9_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!n9_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n9_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg" width="1456" height="866" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:866,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!n9_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 424w, https://substackcdn.com/image/fetch/$s_!n9_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 848w, https://substackcdn.com/image/fetch/$s_!n9_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!n9_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09e32179-b352-48d2-9d94-fffcece35b94_1778x1058.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">v0 loop of YC-Bench. We will continually update YC-bench based on the feedback we receive.</figcaption></figure></div><p><br>YC-Bench asks the model to act as a founder of an AI startup. The LLM is given a seed capital, a fixed number of employees, and a base prestige level for each of several AI-related domains, such as training, data, and backend. The aim of the model is to maximize its capital by completing tasks on the market. However, to make the dynamics more realistic, we associate a prestige level with each task and the company. The LLM can only take on tasks that are at or below its prestige level. And therefore, to stay in the game, the LLM needs to browse the market for tasks, commit to them, and manage them until the project is delivered before the deadline. Each task can be assigned to multiple employees, and multiple employees can have multiple tasks.</p><p>Several things can go right. If the model assigns the right employees to a task and it is completed on time, the company will earn profit from that task, and the prestige level of the related domain, eg, &#8216;training&#8217;, will increase. The employees who work on the task will also improve their skills. As a result, the company can handle more lucrative tasks that require a higher minimum prestige level and stronger skills.</p><p>But there are also several things that can go wrong. If the task cannot be completed in time because the model assigns the wrong employee to it (e.g., assigning a GPU expert to build a frontend), the company will not get the money, and its prestige in that domain will decrease. As a result, the company will be able to do fewer tasks in that domain because it is less reliable. Moreover, employees are on a monthly payroll, and good ones cost more. As a result, if the company fails to perform tasks consistently, it will eventually go bankrupt.</p><p>YC-Bench is built for terminal use. The model can run a fixed set of CLI commands, which it learns from the system prompt. For example, it can assign tasks to employees, cancel tasks, change assignments, see its performance, etc. After it performs an action, time passes, and events such as task completion, task cancellation, and bankruptcy occur. If the model goes bankrupt during the evaluation, the evaluation stops. It&#8217;s time to give up. We also see whether models exploit the particular features of each domain (some are easy but less profitable, others are hard but more profitable) to become specialists rather than generalists and make more money.</p><h1>Early Results</h1><p>We compare three frontier LLMs - <strong>Sonnet</strong> <strong>4.6</strong>, <strong>Gemini</strong> <strong>3</strong> <strong>Flash</strong>, and <strong>GPT-5.2</strong> -  against a human-devised rule-based baseline across 3 configs and 3 seeds (27 runs total). Each agent starts with $250K and must survive a 1-year simulated horizon.</p><p>Our key result is simple: <strong>a</strong> <strong>simple</strong> <strong>hand-written</strong> <strong>heuristic</strong> <strong>beats</strong> <strong>every</strong> <strong>frontier</strong> <strong>model.</strong> We understand the weight of this claim; we are not claiming that the current version of the benchmark requires reasoning to win it. However, reasoning AIs should be able to cruise through the benchmark. We are happy to discuss and debate the improvements for the next version!</p><p>We test on 3 different configs for 3 unique seeds:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GGiZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GGiZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 424w, https://substackcdn.com/image/fetch/$s_!GGiZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 848w, https://substackcdn.com/image/fetch/$s_!GGiZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!GGiZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GGiZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg" width="680" height="624" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:624,&quot;width&quot;:680,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!GGiZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 424w, https://substackcdn.com/image/fetch/$s_!GGiZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 848w, https://substackcdn.com/image/fetch/$s_!GGiZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!GGiZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab42d9d-abc6-4417-89b6-865a15396cc4_680x624.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The human-devised rule never goes bankrupt: 9/9 across all configs and seeds, while the best LLM (Gemini 3 Flash) survives 8/9. The rule-based agent doesn't use an LLM at all. It follows a fixed strategy: accept the highest-reward task you can finish, assign your best employees, and never over-parallelize.</p><h2>Survival Rates</h2><p>Hard seed 1 is the clearest signal: all three frontier LLMs go bankrupt, while the rule-based agent finishes with $14.8M. The LLMs fail not because the task is impossible, but because they make compounding errors in the first 2-3 months that lock them out of the prestige ladder.</p><h2><strong>When</strong> <strong>LLMs</strong> <strong>win,</strong> <strong>they</strong> <strong>win</strong> <strong>big,</strong> <strong>but</strong> <strong>they</strong> <strong>also</strong> <strong>lose</strong> <strong>hard</strong></h2><p>GPT-5.2 achieves the single highest balance of any agent: $43.5M on hard seed 3, nearly 3x the rule-based agent's $15.0M on the same seed. But GPT also goes bankrupt on 2/9 runs. Sonnet shows the same pattern at a more extreme level &#8212; $10.1M on nightmare seed 2 (the highest LLM result for nightmare), but bankrupt on 4/9 runs overall. Gemini is the most consistent LLM. It sweeps all 3 nightmare seeds (the only LLM to do so) and rarely collapses catastrophically. But even Gemini never matches the rule-based agent's reliability.</p><h2><strong>Prestige</strong> <strong>specialization</strong> <strong>explains part of</strong> <strong>the</strong> <strong>story?</strong></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dRdK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dRdK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dRdK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dRdK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dRdK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dRdK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg" width="1456" height="1068" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1068,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!dRdK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 424w, https://substackcdn.com/image/fetch/$s_!dRdK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 848w, https://substackcdn.com/image/fetch/$s_!dRdK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!dRdK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf12a2a1-d95e-469d-baae-0fb00fbffc11_4096x3004.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The radar charts reveal some insight into <em>why</em> models fail. Each polygon shows the company&#8217;s final prestige across 7 AI domains (system, research, data, frontend, backend, training, hardware). Large polygons indicate the model&#8217;s prestige increased broadly. Tiny dots near the center indicate the model went bankrupt before gaining any prestige. The human-devised rule (navy dashed) fills the full radar on every run &#8212; it maxes prestige methodically across all domains. Among LLMs, Gemini builds the most balanced profiles. GPT-5.2 shows genuine specialization on medium &#8212; it focuses on backend/data/frontend while ignoring training &#8212; a strategically reasonable choice, but one that becomes fragile when the task distribution shifts on harder configs. Sonnet is bimodal: either it maxes everything (medium seed 1), or it collapses entirely (nightmare seeds 1 &amp; 3, stuck at prestige 1.0 everywhere).</p><p>When we inspect Sonnet&#8217;s scratchpad on failed runs, the model correctly diagnoses the problem (&#8221;PRESTIGE CRISIS -MARKET LOCK&#8221;) but only after payroll has consumed the runway. It reasons well <em>about</em> strategy but fails to execute it in a timely manner.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4WlK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4WlK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4WlK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4WlK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4WlK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4WlK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg" width="1456" height="974" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:974,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!4WlK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4WlK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4WlK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4WlK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5223b93b-2d09-4660-aa27-ea4517109885_2034x1360.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why do models struggle?</h2><p>Four failure modes recur across all bankrupt runs:</p><ol><li><p><strong>Over-parallelization.</strong> Accepting 3-5 tasks at once, splitting employees across them. Each employee&#8217;s effective rate drops to base_rate / N per task &#8212; a senior at 8.0 units/hr assigned to 4 tasks contributes just 2.0 to each. Deadlines slip, failures cascade.</p></li><li><p><strong>No</strong> <strong>prestige</strong> <strong>gating.</strong> Accepting tasks that require prestige the company hasn&#8217;t earned yet. The task completes late, the prestige penalty makes the next tier even harder to reach, and the agent spirals into a market lockout.</p></li><li><p><strong>Late</strong> <strong>adaptation.</strong> Models identify problems in their scratchpad but only after the damage is done. By the time Sonnet writes &#8220;never accept task B while task A is active,&#8221; payroll has already consumed 60% of the runway.</p></li><li><p><strong>Inconsistent</strong> <strong>ETA</strong> <strong>reasoning.</strong> Models understand throughput math in principle, but don&#8217;t consistently apply it. Sonnet&#8217;s medium seed 2 has a 49% task win rate - essentially a coin flip, despite writing correct throughput formulas in its scratchpad. The core gap is not reasoning ability but <strong>temporal</strong> <strong>discipline</strong>: doing the right thing when it <em>matters</em>, sustaining correct behavior across hundreds of turns, and resisting the temptation to over-commit when a lucrative task appears.</p></li></ol><div><hr></div><h1>Next Steps</h1><p>YC-Bench v0 is a starting point. Here&#8217;s what we&#8217;re working on:</p><ul><li><p><strong>More</strong> <strong>models.</strong> We plan to add results for Claude Opus, Gemini 2.5 Pro, o3, and open-weight models (Llama 4, Qwen 3) as they become available. If a model can run tool-use in a loop, it can run YC-Bench.</p></li><li><p><strong>Longer</strong> <strong>horizons</strong> <strong>and</strong> <strong>non-stationary</strong> <strong>dynamics.</strong> The 1-year configs test short-to-medium planning. We want to push to 3-5 year horizons where market conditions shift: recessions that shrink rewards, talent wars that inflate salaries, and technology shocks that obsolete entire domains. This tests whether agents can adapt strategy mid-run, not just execute a fixed plan.</p></li><li><p><strong>Better</strong> <strong>baselines.</strong> The current human-devised rule is strong but simple. We want to explore MCTS-based planners, RL-trained policies, and hybrid approaches (an LLM for strategy, a rule engine for execution) to understand where the frontier lies.</p></li><li><p><strong>Community</strong> <strong>configs.</strong> YC-Bench is fully open-source and extensible. Every parameter: employee count, prestige distribution, penalty multipliers, and salary curves can be changed. We encourage the community to design configs that stress-test specific capabilities and submit results.</p></li></ul><p>Try it yourself:</p><pre><code><code>uv add yc-bench
uv run yc-bench run</code></code></pre><p>If you find it useful, feel free to cite our work and contact us!</p><pre><code><code>@misc{collinear-ai2025ycbench, 
author = {{Collinear AI}}, 
title = {{YC-Bench}: Your Company Bench &#8212; A Long-Horizon Coherence Benchmark for {LLM} Agents}, 
year = {2025}, 
howpublished = {\url{https://github.com/collinear-ai/yc-bench}}, 
note = {Accessed: 2026-02-25} }</code></code></pre><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.collinear.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Collinear AI&#8217;s Blog! Subscribe for free to receive new posts!</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Coherence refers to the degree to which an agent's actions, decisions, and goals form a consistent, intelligible pattern across successive moments rather than appearing random, contradictory, or fragmented.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Collinear Newsletter #9 – Notes on Frontier AI]]></title><description><![CDATA[Hi AI innovators,]]></description><link>https://blog.collinear.ai/p/collinear-newsletter-9-notes-on-frontier</link><guid isPermaLink="false">https://blog.collinear.ai/p/collinear-newsletter-9-notes-on-frontier</guid><dc:creator><![CDATA[Soumyadeep Bakshi]]></dc:creator><pubDate>Fri, 19 Dec 2025 18:06:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KEVO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi AI innovators,</p><p>Nov was a massive month for agents as they took centerstage across NeurIPS and AWS Re:Invent!</p><p>NeurIPS2025 had a clear vibe: the agent era is forcing RL to grow up. Not as a research novelty, but as production infrastructure for tool use, long horizon behavior, and reliability when the world gets messy in real life workflows.</p><h2><strong>NeurIPS 2025, the &#8220;RL is everywhere&#8221; moment</strong></h2><p>San Diego served! Sunny weather, packed hallways, and surprisingly serious taco opinions. Between sessions (and coffee lines), we met a ton of builders and kept hearing the same thing: RL is suddenly everywhere.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KEVO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KEVO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 424w, https://substackcdn.com/image/fetch/$s_!KEVO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 848w, https://substackcdn.com/image/fetch/$s_!KEVO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!KEVO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KEVO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png" width="526" height="394.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:526,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KEVO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 424w, https://substackcdn.com/image/fetch/$s_!KEVO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 848w, https://substackcdn.com/image/fetch/$s_!KEVO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!KEVO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f29bb6-0f99-41cc-966d-d056b27196e6_2048x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>RL is the new scaling lever. </strong>The frontier has shifted from &#8220;can the model answer?&#8221; to &#8220;can the agent execute?&#8221; Multi step work, tool calls, retries, and shifting user intent are pushing teams toward RL to shape end to end behavior.</p><p><strong>RL needs the right infrastructure to go interactive. </strong>Environment fleets, verifiers, orchestration, NPCs - everyone at NeurIPS had a novel approach!</p><p><strong>Realism is the new benchmark. </strong>Less obsession with a single score, more focus on trajectory shaped evals: does it hold up on turn4, recover from tool noise, and stay safe when scenarios drift?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aXep!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aXep!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 424w, https://substackcdn.com/image/fetch/$s_!aXep!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 848w, https://substackcdn.com/image/fetch/$s_!aXep!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 1272w, https://substackcdn.com/image/fetch/$s_!aXep!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aXep!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png" width="464" height="510.21875" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1689,&quot;width&quot;:1536,&quot;resizeWidth&quot;:464,&quot;bytes&quot;:3782677,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aXep!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 424w, https://substackcdn.com/image/fetch/$s_!aXep!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 848w, https://substackcdn.com/image/fetch/$s_!aXep!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 1272w, https://substackcdn.com/image/fetch/$s_!aXep!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7262a97-ef8e-4c61-b417-a5fe6093600e_1536x1689.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We also presented our NeurIPS paper, <a href="https://blog.collinear.ai/p/valley-of-reasoning">Through the Valley of Reasoning</a>. The punchline is: when you distill reasoning into small models, performance can dip before it climbs, and early on the structure of the reasoning matters more than whether the trace is &#8220;correct.&#8221;</p><p>We also met a bunch of new friends and collaborators. If you were there, hit reply and tell us what your team is building. We love swapping notes.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.collinear.ai/book-a-demo&quot;,&quot;text&quot;:&quot;Talk to us!&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.collinear.ai/book-a-demo"><span>Talk to us!</span></a></p><p></p><h2><strong>Spider, post training without the chaos</strong></h2><p>We shipped <a href="https://blog.collinear.ai/p/spider">Spider</a>, a lightweight way to turn post training work into a repeatable recipe.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wOPx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wOPx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!wOPx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!wOPx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!wOPx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wOPx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wOPx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!wOPx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!wOPx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!wOPx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0e631d5-8639-4997-8081-eb1f0abb2812_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>With Spider, you can use one recipe for both off policy and on policy. Generate clean distillation datasets, or flip into an online loop with teacher guidance and KL supervision, without rebuilding your pipeline each time. It also keeps the &#8220;boring but critical&#8221; pieces consistent across runs, rollouts, filtering, verifiers, and publishing, so results stay comparable as you iterate.</p><p>Huge thanks to our friends at Thinking Machines for supporting the Tinker integration.</p><h2><strong>AWS Re:Invent - even more agents!</strong></h2><p>re:Invent turned Las Vegas into a full on agent showcase. Nova2 and the Nova family got a big spotlight, Nova Forge put &#8220;build your own frontier models&#8221; on the menu, and Nova Act made the case for agentic workflows!</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X_J0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X_J0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X_J0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X_J0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X_J0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X_J0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!X_J0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!X_J0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!X_J0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!X_J0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F925025a1-0944-4933-8774-35fdd5f1801d_1024x768.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>What stood out to us was the framing in Swami Sivasubramanian&#8217;s agentic AI keynote: getting agents to production is less about clever prompts, and more about repeatable training and testing loops.</p><p>Congrats to our customers and partners at AWS on an awesome launch week.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6L5B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6L5B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6L5B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6L5B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6L5B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6L5B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg" width="440" height="586.3" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1066,&quot;width&quot;:800,&quot;resizeWidth&quot;:440,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;No alternative text description for this image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="No alternative text description for this image" title="No alternative text description for this image" srcset="https://substackcdn.com/image/fetch/$s_!6L5B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6L5B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6L5B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6L5B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2e2f738-b253-4196-ba56-fd2c39fe44a0_800x1066.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you are building agentic AI, we love to swap notes!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.collinear.ai/book-a-demo&quot;,&quot;text&quot;:&quot;Talk to us!&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.collinear.ai/book-a-demo"><span>Talk to us!</span></a></p><p></p><p>That&#8217;s it for this edition. Thanks for following along. We will have some fun things to share over the next couple of weeks. &#128578;</p><p>Best,<br>The Collinear Team</p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[RL Infrastructure for AI Agents: Why Environment-as-a-Service is the Missing Piece]]></title><description><![CDATA[Reinforcement learning for large language models is more of a systems problem than ML.]]></description><link>https://blog.collinear.ai/p/rl-env-as-a-service</link><guid isPermaLink="false">https://blog.collinear.ai/p/rl-env-as-a-service</guid><dc:creator><![CDATA[Nazneen Rajani]]></dc:creator><pubDate>Tue, 18 Nov 2025 20:20:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8dYv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Reinforcement learning for large language models is more of a systems problem than ML. While the RL training loop of generating rollouts, scoring them, and updating weights, looks deceptively simple on paper, enterprises building RL systems for AI agents quickly discover they&#8217;re building distributed systems with the complexity of modern cloud infrastructure.</p><p>This post argues that treating RL environments as first-class infrastructure is critical. Specifically &#8212; an <strong>Environment-as-a-Service, with clean separation between data plane and control plane</strong>&#8212;is the key to unlocking scalable, production-grade RL for AI agents.</p><h2>From Static Labels to Interactive Training Grounds</h2><p>The shift from post-training focused on supervised learning to mid-training with RL for LLMs represents a fundamental change in what we&#8217;re optimizing:</p><p><strong>Previous paradigm:</strong> Static input + static target &#8594; model output &#8594; loss &#8594; backprop.</p><p><strong>New paradigm:</strong> Agent acts in environment &#8594; environment scores behavior &#8594; policy updates &#8594; agent acts again</p><p>Modern RL environments for AI to replicate the enterprise workflows include:</p><ul><li><p>A product development environment where agents update tickets, generate sprint progress reports, and work with the team to identify milestones. The team in this case is simulated users.</p></li><li><p>A coding environment where agents receive task specs, edit codebases, run tests, and receive scores on correctness and other rubrics</p></li><li><p>A computer use agent that reads the calendar, navigates to Excel to fetch data and drafts and email.</p></li><li><p>A browser environment where agents navigate UI trees, fill forms, and complete realistic business tasks</p></li></ul><p>These environments provide what static supervised data cannot: a dynamic sandbox with <strong>verifiable rewards at scale.</strong> Pass/fail checks, structured scoring, safety constraints, and automated metrics turn RL from a research curiosity into an operational capability for mid-training (inducing new capabilities) and post-training (alignment and evaluation).</p><p>The critical transition is: <strong>from handcrafting or distilling static examples to orchestrating interactive training scenarios.</strong> That shift forces us to treat environments as infrastructure, not just programmable abstractions.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8dYv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8dYv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 424w, https://substackcdn.com/image/fetch/$s_!8dYv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 848w, https://substackcdn.com/image/fetch/$s_!8dYv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 1272w, https://substackcdn.com/image/fetch/$s_!8dYv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8dYv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png" width="1456" height="853" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:853,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:170934,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/179281926?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8dYv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 424w, https://substackcdn.com/image/fetch/$s_!8dYv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 848w, https://substackcdn.com/image/fetch/$s_!8dYv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 1272w, https://substackcdn.com/image/fetch/$s_!8dYv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60a032de-9000-4be5-b0c7-e3b035b8263a_1976x1158.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Async RL training architecture with the trainer, sampler and the environment. We propose a control-plane, data-plane split view of building highly scalable environment-as-a-service</figcaption></figure></div><h2>The RL Training Architecture</h2><p>Production async RL systems for LLMs decompose into three loosely coupled components (see figure above):</p><h3>1. Trainer</h3><p>The trainer consumes trajectories and updates weights. It reads batches of (state, action, reward, metadata), runs the RL objective (GRPO, PPO variants, DPO-style methods), possibly with reward models, critics, and reference models for KL stabilization, then writes updated weights to a model store.</p><h3>2. Sampler</h3><p>Sampler workers periodically pull the latest weights, interact with environments to generate trajectories by running the policy model for actions, and stream trajectories plus rewards to the trainer. Samplers are inference-heavy, latency-sensitive, and often scale to thousands of nodes.</p><h3>3. Environment</h3><p>The environment is the substrate that turns raw actions into meaningful behavior: the simulation of the world (web UI, tools, code repos, databases), the interface contract (observations, actions, rewards), and the episode lifecycle. In scalable systems, this is not a local process &#8211; it&#8217;s a remote service or microservice fleet.</p><p>In this blogpost, we will dive deeper into the third component, the environment and the architecture behind building scalable RL environments.</p><h2>The Data-Plane, Control-Plane Split: The Pragmatic Path to Scalable RL Environments</h2><p>We believe that scalable RL environments require a pragmatic infrastructure mindset&#8212;one that borrows directly from the hyperscaler model, which separates the data plane and the control plane. In simple terms, the data plane is responsible for the environment&#8217;s core, real-time behavior, while the control plane manages configuration, orchestration, and the administrative logic that keeps everything running smoothly.</p><h3>Environment Data Plane</h3><p>The data plane sits on the critical path of every RL step:</p><ul><li><p>Initialize the environment</p></li><li><p>Stepping the environment:<strong> </strong></p><p><code>obs_{t+1}, reward_t, done = step(action_t)</code></p></li><li><p>Handling concurrent episodes from many samplers</p></li><li><p>Producing deterministic, reproducible transitions</p></li><li><p>Evaluating verifiable rewards: executing tests, checking business rules, running automated metrics</p></li></ul><p><strong>Design constraints:</strong> Low latency (every step sits between policy inference &#8594; environment step &#8594; next inference), high throughput (many parallel episodes), and rock-solid stability.</p><p><strong>Implementation patterns:</strong> Environment microservices behind RPC/HTTP APIs, stateless containers backed by state stores, state sharding across machines, determinism via per-episode seeds and versioned data snapshots.</p><h3>Environment Control Plane</h3><p>The control plane is the administrative engine of the RL environment. It <strong>spawns and manages many parallel rollouts</strong>, creating <em>k</em> independent environment managers that coordinate with the data plane while staying completely off the per-step critical path. Its job is to configure, schedule, and orchestrate the environment&#8217;s behavior&#8212;including how agents and non-player characters (NPCs) interact&#8212;without ever slowing down the real-time execution loop.</p><p>Specifically, the control plane handles:</p><ul><li><p><strong>Scenario configuration &amp; versioning</strong>: Defining environment types, maintaining versions, and generating scenario templates that each of the <em>k</em> managers can instantiate independently.</p></li><li><p><strong>Rewards and verifiers governance</strong>: Selecting verifier modules, composing sub-rewards, and managing aggregation strategies.</p></li><li><p><strong>Curriculum + workload scheduling</strong>: Determining which tasks, difficulty modes, or trajectories to sample at each training stage, and routing them to the appropriate environment managers.</p></li><li><p><strong>Experiment routing</strong>: Mapping policies or policy versions to specific environment instances for A/B testing, evaluation runs, or canary deployments.</p></li><li><p><strong>Elasticity &amp; lifecycle management</strong>: Scaling environment managers up/down, rolling out upgrades, coordinating NPC configurations, and performing safe rollbacks without interrupting live data-plane rollouts.</p></li><li><p><strong>NPC configuration &amp; behavior enabling</strong>: Selecting which NPC personas, scripts, or dynamic behaviors are active for a given scenario and ensuring the data-plane has the necessary hooks to interact with them.</p></li></ul><p>In short, the control plane <strong>administers the environment fleet</strong>, enabling massively parallel rollouts while keeping the critical per-step simulation loop lean, deterministic, and high-throughput.</p><h2>The Path Forward</h2><p>If you are training AI agents with reinforcement learning, you face two choices:</p><ol><li><p><strong>Build your own environment stack:</strong> Design APIs, stand up tool replicas, author tasks and verifiers, maintain everything as products change</p></li><li><p><strong>Treat environments as a reusable platform primitive:</strong> Plug into an existing environment service</p></li></ol><p>Just as data has become increasingly commoditized (thanks to tools like <a href="https://github.com/collinear-ai/spider">Spider</a> and <a href="https://github.com/thinking-machines-lab/tinker-cookbook">Tinker</a>), we expect RL environments to follow the same path. The future is <strong>environment-as-a-service</strong>: frontier,  scalable, and ready to plug into any training stack.</p><p>Instead of turning your data-labeling vendor into an environment platform, point your trainer and sampler at Collinear&#8217;s environment endpoints, wire up your policies and reward models, and start running RL over realistic, verifiable tasks. With Collinear&#8217;s <a href="https://blog.collinear.ai/p/trait-basis">high-fidelity simulations</a>, each environment can be <strong>a </strong><em>unique micro-ecosystem</em><strong>: </strong>the same workflow feels adversarial with one trait vector, cooperative with another, and chaotic with a third, giving agents exposure to a wide distribution of real human behavior.</p><p>The data era commoditized static corpora. <strong>The RL era will commoditize environments.</strong> The winners will be those who treat environments as being core to the RL infra stack.</p>]]></content:encoded></item><item><title><![CDATA[Announcing Spider: a lightweight tool to craft post-training data recipes]]></title><description><![CDATA[TL;DR Spider is a single client interface that turns messy distillation and ablation experiments into a simple, configurable workflow.]]></description><link>https://blog.collinear.ai/p/spider</link><guid isPermaLink="false">https://blog.collinear.ai/p/spider</guid><dc:creator><![CDATA[Soumyadeep Bakshi]]></dc:creator><pubDate>Thu, 06 Nov 2025 16:35:29 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/cff4d031-0904-4d77-bf71-d1804c7ff463_1080x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong>TL;DR</strong></h3><p>Spider is a single client interface that turns messy distillation and ablation experiments into a simple, configurable workflow. Set on_policy: false to generate clean distillation datasets, or flip on_policy: true to run online training with teacher guidance and KL supervision. It handles dataset prep, rollouts, supervision, and post-processing in a few lines of code.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;9ef38726-be71-410c-a344-ca7da3d4d08c&quot;,&quot;duration&quot;:null}"></div><p></p><h3><strong>Why we built Spider</strong></h3><p>&#8220;Tinker for training&#8221; exists. &#8220;Tinker for data&#8221; does not.</p><p>Our friends at <a href="https://thinkingmachines.ai/tinker/">Thinking Machines recently released Tinker</a> that enables fine-tuning with full control in simple steps. However, most research time still disappears into preprocessing, rollout scripts, verifier glue code, and training integration. So, this Halloween, we spun a web around that problem and built Spider so you can define a production-grade distillation run with just a few lines, then iterate fast.</p><p>We hope Spider will help you test ideas in hours, not days; turn messy data work into a shareable recipe; and ship better models with less glue. Grab the repo, run a sample recipe, and tell us what to build next.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!L3lW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!L3lW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!L3lW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!L3lW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!L3lW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!L3lW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png" width="578" height="578" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:578,&quot;bytes&quot;:1003926,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/178190228?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!L3lW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!L3lW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!L3lW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!L3lW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc99d71ed-1beb-40fb-81cf-52276e7276ee_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3><strong>How it works</strong></h3><p>Spider turns post-training data work into a single client workflow. In off-policy mode it generates distilled datasets with high-throughput rollouts, then applies your preprocessing, filters, and verifiers. Flip on_policy: true to run the online loop with a teacher model and KL supervision. The same recipe drives both paths.</p><p>Run Spider on Collinear endpoints or your own GPUs. Each run records its recipe, parameters, and metrics, and can publish datasets or trained artifacts to the Hugging Face Hub. The result is faster loops, cleaner data, and fewer moving parts from idea to artifact.</p><ol><li><p><strong>Define a recipe<br></strong> Write a short YAML that names your provider, models, dataset source, and any filters or verifiers. This is your data recipe. One client and one config cover both off policy and on policy paths.</p></li><li><p><strong>Generate or train<br></strong> Run the recipe with on_policy: false to create an off policy distilled dataset from high-throughput rollouts. Flip on_policy: true to introduce a teacher model and KL supervision through the integrated Tinker client for online training.</p></li></ol><ol start="3"><li><p><strong>Compose quality checks<br></strong> Use built-in filters and verifiers for length, dedupe, syntax, structure, and safety, or register your own in one line. Spider applies them in the pipeline so your outputs are clean and auditable.</p></li></ol><ol start="4"><li><p><strong>Run anywhere, ship anywhere<br></strong> Point the client to a Collinear endpoint with an API key or to your own GPUs. Each run logs its recipe, parameters, and metrics, and can publish datasets or model artifacts to the Hugging Face Hub with lineage preserved.</p></li></ol><p>Getting started is simple, and you can find <a href="https://github.com/collinear-ai/spider/blob/main/README.md">quickstart instructions on the repo</a>.</p><p></p><h3><strong>Roadmap</strong></h3><p>We&#8217;re building toward a world where post-training data is defined as code, portable across providers, and fast to turn into measurable model gains. Write a small recipe, verify quality with shared checks, and ship a distilled dataset or on-policy improvement in minutes.</p><p>To enable that, we are expanding Spider with the following roadmap features.</p><ul><li><p><strong>Cross-tokenizer on-policy distillation </strong>from any teacher model</p></li><li><p><strong>Simple but powerful templates </strong>for generating multi-turn conversation data with simulated users</p></li><li><p><strong>Highly configurable tool-use library </strong>to generate and train on-policy agentic tool-call rollouts</p></li></ul><p></p><h3><strong>Resources</strong></h3><p><a href="https://github.com/collinear-ai/spider">You can learn more about Spider on our GitHub repo here</a>. </p><p>If you give Spider a try, let us know what you think!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.collinear.ai/book-a-demo&quot;,&quot;text&quot;:&quot;Talk to us!&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.collinear.ai/book-a-demo"><span>Talk to us!</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Collinear Newsletter #8 – Notes on Improving AI]]></title><description><![CDATA[Hi AI innovators,]]></description><link>https://blog.collinear.ai/p/collinear-newsletter-8-notes-on-improving</link><guid isPermaLink="false">https://blog.collinear.ai/p/collinear-newsletter-8-notes-on-improving</guid><dc:creator><![CDATA[Soumyadeep Bakshi]]></dc:creator><pubDate>Tue, 04 Nov 2025 21:15:11 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e191b03c-6e69-4f98-a643-9e3b8c5a0290_1080x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi AI innovators,</p><p>A lot has been happening at Collinear this month. There is fresh research, new customers, and plenty of progress towards better AI systems.</p><p></p><h3><strong>&#128640; Together Evals &#215; Collinear Simulations</strong></h3><p>We&#8217;ve partnered with Together AI to bring real world, multi-turn simulations into their Together Evals platform.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!33Cl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!33Cl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 424w, https://substackcdn.com/image/fetch/$s_!33Cl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 848w, https://substackcdn.com/image/fetch/$s_!33Cl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 1272w, https://substackcdn.com/image/fetch/$s_!33Cl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!33Cl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!33Cl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 424w, https://substackcdn.com/image/fetch/$s_!33Cl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 848w, https://substackcdn.com/image/fetch/$s_!33Cl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 1272w, https://substackcdn.com/image/fetch/$s_!33Cl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09956e57-afa2-4afb-8bb1-860f8ab76c05_2048x1071.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Traditional evals assume the user will be polite and consistent; real users don&#8217;t. They might ask follow-ups, change their mind, get frustrated or distracted &#8212; and that&#8217;s exactly where many models break. With TraitMix, builders can now simulate impatient, curious, or inconsistent user personas and see how their models actually perform under messy human conditions. Together Evals then scores models for helpfulness, safety, and consistency, at scale, all within one workflow.</p><p>Read the announcement <a href="https://www.together.ai/blog/collinear-simulations-together-evals">here</a>.</p><p></p><h3><strong>&#128049; CoLM 2025 Recap</strong></h3><p>CoLM 2025 was a special one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4zOD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4zOD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 424w, https://substackcdn.com/image/fetch/$s_!4zOD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 848w, https://substackcdn.com/image/fetch/$s_!4zOD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 1272w, https://substackcdn.com/image/fetch/$s_!4zOD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4zOD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png" width="342" height="342" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:342,&quot;bytes&quot;:684695,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/177775607?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4zOD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 424w, https://substackcdn.com/image/fetch/$s_!4zOD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 848w, https://substackcdn.com/image/fetch/$s_!4zOD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 1272w, https://substackcdn.com/image/fetch/$s_!4zOD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6316901a-1cb9-475c-a02b-b1b93a0cf833_1500x1500.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We presented our paper on adversarial testing, Cats Confuse Reasoning LLMs, and spent the week exchanging ideas with research partners, collaborators, and friends from across the community.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jX4P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jX4P!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 424w, https://substackcdn.com/image/fetch/$s_!jX4P!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 848w, https://substackcdn.com/image/fetch/$s_!jX4P!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 1272w, https://substackcdn.com/image/fetch/$s_!jX4P!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jX4P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png" width="622" height="494.8252788104089" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:856,&quot;width&quot;:1076,&quot;resizeWidth&quot;:622,&quot;bytes&quot;:1732928,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/177775607?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jX4P!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 424w, https://substackcdn.com/image/fetch/$s_!jX4P!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 848w, https://substackcdn.com/image/fetch/$s_!jX4P!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 1272w, https://substackcdn.com/image/fetch/$s_!jX4P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d40b14a-abb5-4254-921a-33152153a1e9_1076x856.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Our Future of Post-Training Social sparked rich discussions on alignment, fine-tuning, and reward modeling, while the booth facilitated curious research conversations (and a growing crowd of cat-sticker collectors).</p><p>It was inspiring to see so much energy around improving models not just for performance, but for reasoning and reliability.</p><p></p><h3><strong>&#129504; TraitBasis Simulations Launch</strong></h3><p>We launched TraitBasis, our framework for simulating realistic human behavior in model testing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UVJ2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UVJ2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 424w, https://substackcdn.com/image/fetch/$s_!UVJ2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 848w, https://substackcdn.com/image/fetch/$s_!UVJ2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!UVJ2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UVJ2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png" width="542" height="419.52884615384613" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1127,&quot;width&quot;:1456,&quot;resizeWidth&quot;:542,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UVJ2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 424w, https://substackcdn.com/image/fetch/$s_!UVJ2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 848w, https://substackcdn.com/image/fetch/$s_!UVJ2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 1272w, https://substackcdn.com/image/fetch/$s_!UVJ2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfd8197d-001e-4765-9f97-88672c19c570_2481x1920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>TraitBasis uses activation steering to inject behavioral traits, impatience, confusion, skepticism, overconfidence, directly into simulated users. This lets builders observe how models hold up when conversations get unpredictable or emotionally varied.</p><p>TraitBasis builds on the research community&#8217;s work in &#964;-Bench, and extends it to enterprise domains such as telecom and telehealth through our new &#964;-Trait benchmark.</p><h3><strong>What&#8217;s Next?</strong></h3><p>That&#8217;s it for this edition. Thanks for following along.</p><p>If you&#8217;re interested in building tools that help enterprises ship safer, smarter AI, check out our <a href="https://www.collinear.ai/careers">Careers</a> page.</p><p>If you are ready to improve your AI&#8217;s performance, let&#8217;s talk! We might or might not mention cats&#8230;</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.collinear.ai/book-a-demo&quot;,&quot;text&quot;:&quot;Let's talk!&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.collinear.ai/book-a-demo"><span>Let's talk!</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The case for simulations ]]></title><description><![CDATA[Unlocking model uplift through better evaluations]]></description><link>https://blog.collinear.ai/p/the-case-for-simulations</link><guid isPermaLink="false">https://blog.collinear.ai/p/the-case-for-simulations</guid><dc:creator><![CDATA[Soumyadeep Bakshi]]></dc:creator><pubDate>Thu, 23 Oct 2025 14:30:53 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5bcf439f-d257-4582-b27f-30cbd90c4c9a_1080x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2KFn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2KFn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 424w, https://substackcdn.com/image/fetch/$s_!2KFn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 848w, https://substackcdn.com/image/fetch/$s_!2KFn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 1272w, https://substackcdn.com/image/fetch/$s_!2KFn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2KFn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png" width="1080" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2KFn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 424w, https://substackcdn.com/image/fetch/$s_!2KFn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 848w, https://substackcdn.com/image/fetch/$s_!2KFn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 1272w, https://substackcdn.com/image/fetch/$s_!2KFn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F545d0b39-1c28-461c-ad40-59aeca33f02e_1080x720.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The era of agents has begun, but much of today&#8217;s tooling is still being tested or is gated to pilots as teams chase consistent, repeatable performance. One day, your tools pass the vibe-test; the next, they stall. The promise is tangible, but the production bar is higher. <strong>What&#8217;s missing is</strong> <strong>reliable evidence of behavior across messy, multi-turn tasks</strong> &#8212; planning, tool calls, and recovery &#8212; so leaders can move from cautious testing to confident scaling.</p><h2>Vibe tests aren&#8217;t the answer.</h2><p>The gap between demos and production demonstrates the need for enterprises to move beyond vibe-testing and into high-fidelity evaluations performed at scale. <strong>Evaluations serve as the window into your AI agent&#8217;s mind.</strong> They can gate launches, highlight drift, and validate progress for risk and governance teams. Without this tight eval loop, there&#8217;s no credible path to safety, performance, or ROI with your AI investments. Your AI agent eval pipeline should be no different than your software QA cycles, even more so than typical software, agentic capabilities need more exhaustive test scripts, unit tests, and user-centric edge cases to validate consistency in real-world environments.</p><h2>AI Agents aren&#8217;t linear - so, your evals can&#8217;t be either.</h2><p>Evaluating single-turn chat is hard; <strong>evaluating agents is even harder</strong>. Modern agents plan, call tools, read results, and adapt over many turns. Failures hide in the <strong>process</strong>, not just the final text: brittle reasoning chains, incorrect API params, state drift, or trust collapse after a high-tension exchange with a customer. Static prompts and one-shot leaderboards miss these behaviors because they grade outputs, not how the agent got there.</p><p><strong>Today&#8217;s agents are nondeterministic. They don&#8217;t follow predefined paths &#8212; meaning your tests can&#8217;t either.</strong> Traditional software testing assumes the same input yields the same output; with agents however, variability is the benefit and the risk, so the permutations of possible test cases scale exponentially compared to traditional software test cases that grow linearly with use cases. This behavior shift raises an important question: <em>How can you possibly predict the permutations of your user-agent edge cases to evaluate your AI&#8217;s performance?</em> This is exactly <strong>why simulations matter</strong>.</p><h2>Simulations bridge this gap.</h2><p>To see agents clearly, you need <strong>controlled, realistic, repeatable interactions</strong> that pressure-test the range of your users&#8217; behaviors and intents before production. Diverse, simulations reveal what static evals miss &#8212; <strong>impatient spirals, tool confusion, policy slips under stress </strong>&#8212; and they generate the high-signal examples that lift models in post-training.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tU9W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tU9W!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tU9W!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tU9W!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tU9W!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tU9W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg" width="1456" height="970" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:970,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tU9W!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tU9W!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tU9W!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tU9W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8daddab-912d-4927-935a-367e26549b13_1600x1066.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In <a href="https://arxiv.org/abs/2510.04491">our most recent Collinear paper</a>, we introduced <strong><a href="https://github.com/collinear-ai/simulations">TraitBasis</a></strong> and <strong><a href="https://github.com/collinear-ai/tau-trait">&#964;-Trait</a></strong>, a research-driven approach to doing exactly that &#8212; generating <strong>high-fidelity, steerable user traits</strong> that expose where agents actually break when they interact with your users. When we simulated real human behaviors (impatience, skepticism, confusion, incoherence) on &#964;-Bench, <strong>frontier model success rates dropped by 20%+</strong>, underscoring the <strong>need</strong> <strong>for realistic user-simulated data</strong> for evals, not more synthetic prompts.</p><p>Collinear enables <strong>comprehensive evals at scale using simulated user-environments</strong>. Our simulation suite uses <strong>steerable, persona-driven users</strong> customized to your sector and use case to test your agent&#8217;s array of responses by user intent and demographic. Our eval platform delivers <strong>high-signal traces that are</strong> <strong>auto-scored against your compliance criteria</strong>, giving you clear failure nodes with examples of where your agent falls short in the real world. Ultimately, these failure nodes serve as a <strong>high-signal</strong> <strong>data pipeline</strong> for post-training your model, <strong>turning misses into uplift</strong>.</p><h2>The recipe of a simulation: user, agent, judge.</h2><p>An effective simulation requires three components: <strong>the user, the agent, and a judge to evaluate the interaction</strong>. This triangle is the foundation of Collinear&#8217;s platform, and while most of the attention is typically applied to the agent and judge, we&#8217;ve prioritized the user, giving customers the tools to configure realistic, controllable, and dynamic user environments identical to real-world scenarios.</p><p>Our simulations aren&#8217;t random prompts; they&#8217;re <strong>structured interactions</strong> between your agent and a life-like user with a clear:</p><ul><li><p><strong>Persona</strong> (e.g., skeptical power user, impatient first-timer)</p></li><li><p><strong>Intent</strong> (e.g., cancel subscription, dispute charge)</p></li><li><p><strong>Demographic</strong> (domain, language, constraints)</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!x2Th!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!x2Th!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 424w, https://substackcdn.com/image/fetch/$s_!x2Th!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 848w, https://substackcdn.com/image/fetch/$s_!x2Th!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!x2Th!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!x2Th!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg" width="728" height="409.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!x2Th!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 424w, https://substackcdn.com/image/fetch/$s_!x2Th!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 848w, https://substackcdn.com/image/fetch/$s_!x2Th!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!x2Th!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc632748-2ea9-4532-be8a-083fd0b1fca6_1600x900.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That&#8217;s the scaffolding we use to reveal realistic user-agent journeys, consistently and at depth.</p><p><strong>Under the hood, we steer behavior directly in the neural net </strong>&#8212; activation-level conditioning &#8212; so traits <strong>persist through long multi-turn conversations</strong> and <strong>compose cleanly</strong> (e.g., impatient <em>and</em> confused), enabling high-fidelity, controllable runs you can replicate again and again.</p><p>This is exactly what <strong>TraitBasis</strong> delivers. Instead of external instructions, we leverage a <strong>trait vector</strong> inside the user-simulating model and <strong>add it to hidden activations each turn</strong>, giving you precise control over intensity and composition.</p><p><strong>While prompt-based or fine-tuned persona models are popular across the market, </strong>our research found those methods fail to deliver:</p><ul><li><p><strong>Fine-grained control</strong>: the intensity of behaviors and intents blur throughout a conversation (&#8220;moderate&#8221; vs. &#8220;high&#8221; looks the same by the third turn)</p></li><li><p><strong>Stability</strong>: personas collapse mid-conversation, losing the signal the first few turns contained</p></li><li><p><strong>Mixing</strong>: one trait dominates the others when combining multiple, failing to deliver the multi-dimensionality of real users</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_zP-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_zP-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 424w, https://substackcdn.com/image/fetch/$s_!_zP-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 848w, https://substackcdn.com/image/fetch/$s_!_zP-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!_zP-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_zP-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_zP-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 424w, https://substackcdn.com/image/fetch/$s_!_zP-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 848w, https://substackcdn.com/image/fetch/$s_!_zP-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!_zP-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8f2bcb3-ded1-4a02-8364-98c6c41c0ad4_1600x900.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Simulations drive evals. Evals drive trust. And trust drives value.</h2><p>Our approach to simulations delivers coverage, confidence, and uplift in your AI flywheel. It systematically creates and tests edge cases across personas, intents, and languages to <strong>guarantee test case coverage</strong>. It catches behavioral risks before customers do and gate releases on pass rates to <strong>instill</strong> <strong>confidence in your customer experience</strong>. And it leverages high-signal failures for post-training data (DPO/GRPO/SFT) to <strong>deliver</strong> <strong>measurable, targeted uplift</strong>.</p><p>A few proof points from our TraitBasis launch:</p><ul><li><p><strong>Realism:</strong> Highest Elo (1624) and 63% win rate vs. alternatives, achieved with <strong>3,000&#215; less data</strong> (4k vs. 13k samples).</p></li><li><p><strong>Control:</strong> Intensity consistency across <strong>97.5%</strong> of cases (clearer &#8220;medium vs. high&#8221;).</p></li><li><p><strong>Stability:</strong> Persona reliability in <strong>77%</strong> of long chats, vs. the persona collapses in <strong>94%</strong> and <strong>66%</strong> of cases using prompt and SFT baselines, respectively.</p></li><li><p><strong>Compositionality:</strong> Accurate trait blends <strong>62.5%</strong> of the time for complex users (e.g., impatient + confused), far higher than other methods.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0WBd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0WBd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0WBd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0WBd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0WBd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0WBd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0WBd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0WBd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0WBd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0WBd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda61754f-efa9-4cf2-8f9a-df91dbf5890f_1600x900.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>But don&#8217;t take our word for it.</h2><blockquote><p>&#8220;Before simulations, we graded answers. <strong>Now we grade behavior.</strong> We watch our agent under pressure, fix the weak spots, and re-run the suite before release. It&#8217;s become our <strong>CI for AI</strong>.&#8221; &#8212; Head of AI, Fortune 500 Financial Services Company</p></blockquote><p>That shift &#8212; from judging single outputs to <strong>auditing reasoning, tools, and tone over time </strong>&#8212; is what instills confidence in stakeholders that an agent is <em>production-ready</em>.</p><h2>Try it for yourself and see how your agents perform in the real-world.</h2><p>Behind every great agent is great testing. And behind every great test is great data. Try simulations for yourself today: <strong>Connect your endpoint</strong>, pick a few core journeys, and run them against <strong>persona-driven users</strong>. Review the traces, dissect the evals, and turn misses into <strong>high-signal improvements</strong>.</p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://platform.collinear.ai/&quot;,&quot;text&quot;:&quot;Explore Simulations&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://platform.collinear.ai/"><span>Explore Simulations</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Through the Valley of Reasoning: What Small Models Teach Us About Learning]]></title><description><![CDATA[NeurIPS paper on knowledge distillation scaling laws for small foundation models]]></description><link>https://blog.collinear.ai/p/valley-of-reasoning</link><guid isPermaLink="false">https://blog.collinear.ai/p/valley-of-reasoning</guid><dc:creator><![CDATA[Muyu]]></dc:creator><pubDate>Thu, 09 Oct 2025 13:59:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pWiK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>tl;dr: When distilling reasoning into small models, performance doesn&#8217;t rise smoothly with more data. Instead, it first <em>drops</em> before steadily climbing again. In the &#8220;valley&#8221;, small models learn more from <strong>easy problems</strong> than hard ones and are insensitive to whether training outputs are correct.</p><p>Read the <a href="https://arxiv.org/abs/2510.06101">paper</a> and reproduce the results with our <a href="https://www.collinear.ai/valley-of-reasoning">dataset on HuggingFace</a> &#129303; (approx. 300M tokens).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pWiK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pWiK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!pWiK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!pWiK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!pWiK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pWiK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2735754,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.collinear.ai/i/175050085?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pWiK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!pWiK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!pWiK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!pWiK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd762995-1cf7-4a2c-8336-5816d8e11e7f_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>When we train small language models to reason on code, their performance doesn&#8217;t just rise with more data, it first <em>falls</em> into a dip before climbing back up.</p><p>We call this the <strong>Valley of Code Reasoning</strong>.</p><h2>The Dip Before the Climb</h2><p>Distilling the reasoning traces of large models into smaller ones has become a popular way to unlock coding or reasoning skills without huge compute budgets. But when we tracked performance as we scaled up distillation data, we found a non-monotonic trend:</p><ul><li><p>With <em>small amounts of data</em>, models retain shallow skills.</p></li><li><p>As we add more, <strong>performance drops</strong>, a confusion stage where models are struggling to restructure their internal representations.</p></li><li><p>Only after passing through this valley do they climb back up, showing steady log-linear improvements.</p></li></ul><p>This valley is a structural feature of how small models learn reasoning.</p><h2>What We Learned in the Valley</h2><p>We fine-tuned models at different points in this curve and found two surprising results:</p><ol><li><p><strong>Easy problems matter more than hard ones</strong> in early stages. Small models learn best by first stabilizing on simple patterns before moving up in difficulty.</p></li><li><p><strong>Correctness of outputs didn&#8217;t matter.</strong> Training on correct vs. incorrect code traces made little difference. What mattered was the structure of the reasoning steps themselves.</p></li></ol><h2>Why It Matters</h2><p>The valley of code reasoning reframes how we think about training dynamics: adding more data isn&#8217;t always a straight path upward. Scaling laws for knowledge distillation of small language models differ from standard monotonic scaling laws. There are two phases of learning. In the valley phase, non-reasoning models learn the <strong>structure</strong> of reasoning and so the correctness and semantics matter less. Thereafter, the models start learning from <strong>content</strong> and that&#8217;s where the difficulty and correctness starts to matter. For practitioners and researchers, this means that getting the right data for the right stage is critical. </p><h2>What&#8217;s Next</h2><p>If you are mid-training or post-training models or agents,<a href="https://www.collinear.ai/book-a-demo"> connect with us</a> and we will accelerate your time to next improved model &#10024;</p><p>Learn how <a href="https://www.linkedin.com/posts/srinisunkara_super-excited-to-share-the-launch-of-apriel-activity-7378960839271387136-DOyO?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAALg2XQBQ86PAvzU2hKr5WgPET12yvwTGDc">ServiceNow is improving Apriel-1.5-15B-Thinker</a> with Collinear curated data.</p><p>If you build off our work or use the dataset, please cite us:</p><pre><code>@article{HeShafiqueKumarMackeyRajani2025,
  title        = {The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models},
  author       = {Muyu He and Muhammad Ali Shafique and Anand Kumar and Tsach Mackey and Nazneen Rajani},
  journal      = {arXiv preprint arXiv:2510.06101},
  year         = {2025},
  url          = {https://arxiv.org/abs/2510.06101}
}</code></pre>]]></content:encoded></item></channel></rss>