Discussion about this post

User's avatar
Karam Elabd's avatar

Enjoyed this, Jasper. The expiry angle on lab uplift studies, in your 'rocky terrain' section, feels obvious in hindsight but I hadn't thought about it before.

On expiry: assuming that models only get more capable, if model A1 showed uplift, A2 presumably clears the same bar, the capability threshold has been crossed. The question for newer models becomes whether safeguards are holding.

Separately, I wonder if there's room (a sub-tier) between MCQ-style evals and full wet lab RCTs. Benchmarks test model knowledge but not interactive uplift. Wet lab studies capture that, but they're expensive and expire with each model generation. What about structured thought exercises or lab simulations where experts work through realistic scenarios with model assistance, producing a lab protocol as the auditable artifact that a trusted practitioner could then validate in an actual lab? You'd still compare against a control group without model access to measure uplift, but without requiring participants to physically be in a lab with expensive precursors. Might be cheaper and more repeatable across model generations.

Adam Howes's avatar

Copied from LinkedIn in-case here is better for replying:

Nice post! This reminds me of our conversation from the weekend Peter Peneder r.e. different tiers of evidence. I think that the correlates of uplift idea is nice, though probably methodologically challenging to get right. For e.g. the Active Site RCT one could look at the LLM logs, analyze the extent to which failure might have been due to inaccurate or misleading information provided by the model. Then you could try to connect that to whether or not a newer model might have answered more helpfully, and then if that in turn might have changed the participants probability of success. This of course requires a lot of assumptions, and likely won't be a clean easy answer. Still, there could be qualitative case studies where you could say "yes but newer models wouldn't have made that mistake" (and that has some value). This is probably something we could do if it seems worth it!

2 more comments...

No posts

Ready for more?