Collinear AI’s Blog
Subscribe
Sign in
Home
Archive
About
Latest
Top
Discussions
The Measure of Hard but Fair
A benchmark task can defeat a capable system and still tell us little about how systems differ. Three tasks from our security evaluation show what pass…
Sep 11
•
Guga Gogia
16
CWE-bench: Measuring Coding Agent Capabilities on Every Class of Known Vulnerability
CWE-bench is a frontier benchmark for AI Cybersecurity
Sep 2
•
Gonzalo Gonzalez
,
Adit Jain
, and
Nazneen Rajani
18
August 2026
The Vulnerability in the Reward
When the score is also the reward
Aug 28
•
Guga Gogia
14
The Simulated World Is Being Born In Cybersecurity
AGI needs simulated worlds that are hard enough to expose failure and fair enough to make that failure useful. Cybersecurity is forcing these worlds to…
Aug 25
•
Nazneen Rajani
13
The Uninterpretable User
What Simulated Users Say When the Agent Isn’t Listening
Aug 6
•
Parker Seegmiller
,
Chinmaya Andukuri
, and
Guga Gogia
10
July 2026
The User Went for a Cigarette
Modeling Partial Observability for High-Fidelity User Simulation
Jul 1
•
Parker Seegmiller
,
Chinmaya Andukuri
, and
Guga Gogia
9
May 2026
Is your RL environment fair to your agent?
or ensuring that your hillclimbing budget is spent right :)
May 14
•
Adit Jain
9
Whose Taste?
More data won't fix the AI verification problem. Different taste might.
May 7
•
Sachin
15
April 2026
Collinear Newsletter #11 - Notes on Frontier AI
Hi AI innovators,
Apr 30
•
Soumyadeep Bakshi
13
AI's U-235 Problem
Nuclear physics solved for k_eff. What's the AGI equivalent?
Apr 23
•
Jed Gresham
11
SimLab: The self-serve staging playground for real-world agents
Agents fail on real tool calls, long workflows, and messy data. SimLab lets you find those failures in simulation, not in production.
Apr 2
•
Sachin
13
March 2026
Collinear Newsletter #10 - Notes on Frontier AI
Happy March from the Collinear team!
Mar 24
•
Soumyadeep Bakshi
and
Nazneen Rajani
11
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts