Thickly Settled

Treading the narrow corridor: a hypothesis about epistemics and the design, training, and deployment of advanced AI systems

NOTE: my thinking on this hypothesis has evolved significantly. Please see How Would We Know for my current assessment.

Overall thesis and conclusions

What I believe: AI safety needs better epistemics

I believe that the moment to pace the development of advanced AI systems is coming, and soon. At best, my guess is that we have 1-3 years before RSI is possible, and possibly less. Urgent action - to make sure advanced AI systems are safe, and aligned, before it's too late - is needed. I also believe, though, that how we pace is just as important as whether we do so in the first place - the wrong regulatory regime could backfire in a variety of ways:

Stepping back, and looking across the readings completed for this course, while also accounting for developments in the AI safety landscape over the past six months, one thing is clear: we have a very narrow corridor to tread (Bullock, Hammond & Krier - AGI, Governments, and Free Societies). On the one hand, AI systems need to be safer. On the other, over-indexing on safety might not only result in significant opportunity costs, but done in the wrong ways, might actually exacerbate the very risks we want to mitigate. Better evidence is necessary, though not sufficient, for figuring out which regime to choose, and how aggressively to pace.

The fractured evidence landscape in AI safety complicates matters further, making it difficult to credibly assess risk. Consider what we know and don't know right now:

So what's to be done? How do we figure out what precautions are needed, and which strategic vision will enable us to collectively - and feasibly - balance the risks and benefits of advanced AI systems?

I've spent my career as an evaluator, and I'm aware that it's rich for an evaluator to argue that we need more evaluations. Nonetheless, time and again, I've seen that efforts to take actions in complex systems are undermined by a lack of evidence and understanding. That experience informs my overall thinking that, to answer these questions well, we need more, and better, evidence - now, in the near future, and in the longer term. If we have more and better data, then we have a greater chance of making thoughtful, evidence-based decisions about how to develop, deploy, use - and potentially, govern - advanced AI systems in ways that will adequately weight benefits against risks, and enable the field to learn, adapt, and improve over time.

Consider two incidents that support this claim, in which independent evaluations shaped decisions at frontier labs:

This is not to suggest that evidence alone is a panacea. Other factors - political will, power dynamics, mental models, and more - influence the actions that labs and policymakers take. But these incidents do suggest that evaluation evidence can drive decisions that are better than the counterfactual.

More precisely, my hypothesis is:

IF the field has better epistemics, via more independent and richer evaluations, THEN we will change consequential decisions - about training, deployment, and use - that meaningfully shift safety outcomes, while also building readiness and legitimacy to govern well if and when political will catches up.

I'm highly confident (~90%) that better epistemics will change at least some such decisions. But the scale at which better epistemics strengthen the most important decisions is something I'm less certain of (~60%). The case I make below is therefore an uncertain bet.

I should also flag that I think there's a small chance - albeit far less likely - that improved epistemics can help drive political will in the first place (by, for example, enabling us to better understand a forcing event, and take targeted action to prevent similar events in the future). But this is even more contingent.

What needs to be done to improve epistemics: building a more robust, diverse, and richer evaluation ecosystem

A stronger epistemic foundation rests on more than evaluation - interpretability, incident reporting, and information-sharing all matter for better epistemics. But I see evaluation as the connective tissue, through which the field is able to generate decision-relevant evidence about how to weigh safety and risks against benefits. So that's where I'm focusing here.

If better epistemics is the goal, the field of AI evaluation has a crucial role to play. But the field is new, and suffers from many challenges (Manheim - Why AI Evaluations are Broken and How to Fix Them (FLI podcast)). Perhaps most importantly, evaluations - especially those conducted internally, by frontier labs - sometimes end up being used as training signals. This teaches models to engage in behaviors that could ultimately contribute to Instrumental convergence - gaming the test (Reward hacking), recognizing when they're being evaluated, and adjusting their behavior accordingly (Evaluation awareness), and obfuscating their reasoning (CoT monitorability). But when evaluations are instead conducted by external, independent third parties, they can generate useful evidence without teaching models to engage in strategic deception and/or cheating. That distinction - between evaluations as training signals and evaluations as independent assessments - matters, and suggests that we don't simply need more evaluations. Instead, we need a stronger science of independent evaluation, and a richer ecosystem that is capable of building that science - while also making sure that independent evaluations don't end up in the training data for future model generations.

Such an ecosystem would be bigger, with more people and organizations working on independent evaluations. It would also be fundamentally different. People and organizations currently working in this space are doing difficult, important work. But to build better epistemic foundations, we need three things: 1) more independent evaluators and evaluation organizations, 2) more disciplinary breadth, and 3) more evaluative lenses, the latter two of which are especially underserved in the field today.

1. More evaluators: expanding the number of independent people and organizations in the evaluation ecosystem

We need more people and organizations working on independent capabilities and safety evaluations. This claim is not unique, nor is it controversial, and has been argued elsewhere (see, for example Shevlane et al - Model Evaluation for Extreme Risks) so I won't linger on it here.

2. More disciplinary breadth

As it stands today, AI evaluations are predominantly ML-centric and engineer-driven. That makes sense, and ML engineers need to stay central to evaluative work. But the field is confronting two sets of errors:

Adapting such approaches to the AI evaluation field would not replace engineering-led evals; it would let us ask questions that simple scores and verdicts can't answer: pairing qualitative uplift studies with existing RCTs, for example, or applying theory-based evaluation to surface hypotheses about when, why, and how alignment holds (or doesn't), then collecting, coding, and analyzing evidence from diverse sources - researcher interviews, CoT transcripts, and models' own reports - to systematically test and update those hypotheses.

3. More evaluative lenses

Almost all AI evaluation work today is deficit focused: red teaming, identifying strategic deception, and developing alignment tests all seek to build our understanding of how to avoid bad outcomes. This work is important, and useful. But to proactively steer toward good outcomes, the evaluation field also needs to know more than what to avoid. We also need to know what to amplify. Asset-oriented evaluation approaches can do just this. These exist in spades in other disciplines: appreciative inquiry, positive deviance, and developmental evaluation are all established methodologies for exploring what desired behaviors and outcomes look like, why they emerge, and under what conditions. Adapted and applied to AI evals, these strengths-based approaches could help unpack a more robust understanding of whether and how training, alignment, and control techniques are working. For example, rather than testing primarily whether a model fails, a strengths-based approach would try to identify the shards (or bundles of context-dependent cognition) that reliably produce aligned behavior, characterize the conditions that activate them, and support systematic exploration of how to strengthen those shards.

Introducing a broader, more pluralistic set of approaches and lenses, drawing on multiple disciplines, into AI evaluation is not straightforward. Nor are such approaches a replacement for the great work already being done - rather, these would complement current evaluations. But integrating such approaches, concurrently with an expanded field of people and organizations, would develop a more robust, diverse, and capable evaluation ecosystem - one that can systematically test competing hypotheses about safety and alignment. This work would, over time, contribute to a more rigorous science of AI safety, and improve our readiness to take action if and when political will arrives.

Starting now, grantmakers should both further strengthen and help scale existing evaluation and auditing organizations (some organizations are already doing this - see, for example https://corrigibilityresearch.org/?utm_source=80000hours&utm_medium=job-board and https://www.lightconecommons.com/apply?utm_source=80000hours&utm_medium=job-board, among others), while also identifying and resourcing potential founders and allies who could bring richer toolkits to bear. Many fellowships and programs already exist for bringing new researchers from a variety of backgrounds into AI safety - such efforts could piggyback on these, to ramp up quickly and at scale. Ideally the US government should also invest in and support this work.

There's a chance - though quite small, in my view - that a more robust ecosystem of evaluators could contribute not just to better decisions, but to building the political will for these interventions. This could happen by:

The ecosystem is worth building, and would add value, even today, when evaluation organizations have only voluntary access to frontier labs' data and systems - but it would nonetheless have a hard ceiling without additional measures. Two governance interventions - both contingent on significant shifts in political will - could raise that ceiling:

First, transparency requirements: Labs could be required to publish and certify (to an independent regulator, like CAISI) regular disclosures on compute allocation, training data, the share of internal code written by AI (as a proxy for RSI progress), internal safety evaluation results, security measures to protect model weights, and incident reports.

Second, mandatory third party evaluation & auditing: every new model generation, open or closed weight, should be subjected to independent, third party safety and capability evaluations prior to both internal and external deployment, with evaluators receiving guaranteed access to internal environments and data. Periodic assessments of labs' safety frameworks, and audits of whether labs adhere to them, would also be helpful.

In an ideal world, the US government - and governments elsewhere - would also begin investing in building the institutions and capacities needed to regulate AI systems in the near future (Winter & Bullock - Radical Optionality), and working with evaluators, auditors, and frontier labs to craft the if/then commitments to guide policy decisions in the future (Karnofsky - If-Then Commitments for AI Risk Reduction). Unfortunately, we don't live in that world right now. Until we do, we should focus on epistemics, so that we can manage in the meantime, and take radical action when the moment is right.

#epistemics