Treading the narrow corridor: a hypothesis about epistemics and the design, training, and deployment of advanced AI systems
NOTE: my thinking on this hypothesis has evolved significantly. Please see How Would We Know for my current assessment.
Overall thesis and conclusions
- The narrow corridor: we will need to pace advanced AI in the near future (my guess is that RSI is here in 1-3 years, if not before), but how we pace is just as important as whether we do so. Centralizing governance risks surveillance and power concentration; decentralization could lead to proliferation; ineffective regulatory regimes could drive development underground, or weaken safety research; and unilateral pauses could ignite great power conflicts.
- The foundations are shaky: we don't have a strong evidence base with which to assess risk - on alignment, capabilities and timelines, and adherence to safety frameworks. Better epistemics are not the answer to making the right decisions about how to balance risks and benefits. But evidence is valuable - for making development and deployment decisions, and for building readiness to govern well if/when political will aligns.
- But they can be improved: My hypothesis is that stronger epistemics improve near term decisions and actions, and build readiness and legitimacy to govern well in the future.
- A necessary intervention: we need a stronger, more robust ecosystem of independent evaluators, disciplines, and lenses to build the science of AI safety, and deliver better evidence to inform decisionmaking. Ideally we also have mandatory lab transparency and third-party evaluations and audits as well - but a better ecosystem adds value regardless.
- My path: We can, and should start work on building this ecosystem today (in fact, doing so is already in progress!). I have three routes for how I might contribute, that I plan to test: research, grantmaking, and making the case, and will use the Lateral Workshop, networking, and fellowship applications to do so. Goal is to make a call on paths by the end of September.
What I believe: AI safety needs better epistemics
I believe that the moment to pace the development of advanced AI systems is coming, and soon. At best, my guess is that we have 1-3 years before RSI is possible, and possibly less. Urgent action - to make sure advanced AI systems are safe, and aligned, before it's too late - is needed. I also believe, though, that how we pace is just as important as whether we do so in the first place - the wrong regulatory regime could backfire in a variety of ways:
- a centralized approach (e.g., along the lines of Sastry et al - Computing Power and the Governance of Artificial Intelligence) could lead to extreme surveillance, power concentration, and/or regulatory capture (as Ball - A Framework for the Private Governance of Frontier Artificial Intelligence argues)
- a decentralized approach (as described by Ngo - Distributed vs Centralized Agents, for example), on the other hand, could ultimately confer power too broadly, such that rogue actors are able to use dangerously capable systems for nefarious ends (Nielsen - ASI Existential Risk - Reconsidering Alignment as a Goal)
- an ineffective regime would likely do too little to mitigate catastrophic risks, or worse, exacerbate them, by pushing AI system development underground, harming safety research (Belrose - AI Pause Will Likely Backfire), or otherwise hindering society's ability to prepare for and deal with grand challenges arising from superintelligence (McCaskill & Morehouse - Preparing for the Intelligence Explosion)
- slowing AI development would have opportunity costs, to the detriment of future generations, and those still living who would otherwise benefit from AI-powered advances in health, economics, and politics (Bostrom - Optimal Timing for Superintelligence).
- Further, proposals to pause AI development often rest on unproven assumptions about global power dynamics or interventions that aren't yet technically possible (e.g., Scher et al - An International Agreement to Prevent the Premature Creation of Artificial Superintelligence), or could risk unlocking a Thucydides trap to catastrophic effect (see Delaney - Strategic Visions in AI Governance). Even the best proposals for pausing in the interests of safety (e.g., Carlsmith - On Restraining AI Development for the Sake of Safety) acknowledge that doing so optimally is difficult, at best.
Stepping back, and looking across the readings completed for this course, while also accounting for developments in the AI safety landscape over the past six months, one thing is clear: we have a very narrow corridor to tread (Bullock, Hammond & Krier - AGI, Governments, and Free Societies). On the one hand, AI systems need to be safer. On the other, over-indexing on safety might not only result in significant opportunity costs, but done in the wrong ways, might actually exacerbate the very risks we want to mitigate. Better evidence is necessary, though not sufficient, for figuring out which regime to choose, and how aggressively to pace.
The fractured evidence landscape in AI safety complicates matters further, making it difficult to credibly assess risk. Consider what we know and don't know right now:
- Alignment & corrigibility: the most capable systems, at least some of the time, behave in ways that appear misaligned - but how, under what conditions, for what reasons, and to what extent often remains unclear (Palisade Research - Shutdown Resistance in Reasoning Models, METR - Summary of METR's predeployment evaluation of GPT-5.6 Sol)
- Capabilities: capabilities are advancing quickly - but predictions regarding when AI will begin to automate R&D are hotly contested, and evidence on the pace of CBRN uplift and what appropriate safety thresholds might be is mixed (Epoch AI - AI capabilities progress has sped up (ECI), Greenblatt (Dwarkesh Podcast) - What Happens Once AI Can Automate AI Research)
- Safety frameworks: frontier labs have voluntary safety frameworks (preparedness frameworks, RSPs, etc.), but whether and how they are adhering to them, including throughout the model development process, and the extent to which such frameworks actually reduce risk, is hard to say (Williams & Schuett - Anthropic's RSP v3.0, Zvi - Anthropic Risk Report August 2026)
So what's to be done? How do we figure out what precautions are needed, and which strategic vision will enable us to collectively - and feasibly - balance the risks and benefits of advanced AI systems?
I've spent my career as an evaluator, and I'm aware that it's rich for an evaluator to argue that we need more evaluations. Nonetheless, time and again, I've seen that efforts to take actions in complex systems are undermined by a lack of evidence and understanding. That experience informs my overall thinking that, to answer these questions well, we need more, and better, evidence - now, in the near future, and in the longer term. If we have more and better data, then we have a greater chance of making thoughtful, evidence-based decisions about how to develop, deploy, use - and potentially, govern - advanced AI systems in ways that will adequately weight benefits against risks, and enable the field to learn, adapt, and improve over time.
Consider two incidents that support this claim, in which independent evaluations shaped decisions at frontier labs:
- Evans, et. al's findings on emergent misalignment inspired Anthropic to replicate and expand the original evaluation, and develop methods to detect and prevent emergent misalignment during training (Anthropic - Persona Vectors (Monitoring and Controlling Character Traits in Language Models))
- Apollo Research's scheming evaluations encouraged OpenAI to develop and roll out new methods to reduce models' propensity to engage in scheming (Detecting and reducing scheming in AI models | OpenAI)
This is not to suggest that evidence alone is a panacea. Other factors - political will, power dynamics, mental models, and more - influence the actions that labs and policymakers take. But these incidents do suggest that evaluation evidence can drive decisions that are better than the counterfactual.
More precisely, my hypothesis is:
IF the field has better epistemics, via more independent and richer evaluations, THEN we will change consequential decisions - about training, deployment, and use - that meaningfully shift safety outcomes, while also building readiness and legitimacy to govern well if and when political will catches up.
I'm highly confident (~90%) that better epistemics will change at least some such decisions. But the scale at which better epistemics strengthen the most important decisions is something I'm less certain of (~60%). The case I make below is therefore an uncertain bet.
I should also flag that I think there's a small chance - albeit far less likely - that improved epistemics can help drive political will in the first place (by, for example, enabling us to better understand a forcing event, and take targeted action to prevent similar events in the future). But this is even more contingent.
What needs to be done to improve epistemics: building a more robust, diverse, and richer evaluation ecosystem
A stronger epistemic foundation rests on more than evaluation - interpretability, incident reporting, and information-sharing all matter for better epistemics. But I see evaluation as the connective tissue, through which the field is able to generate decision-relevant evidence about how to weigh safety and risks against benefits. So that's where I'm focusing here.
If better epistemics is the goal, the field of AI evaluation has a crucial role to play. But the field is new, and suffers from many challenges (Manheim - Why AI Evaluations are Broken and How to Fix Them (FLI podcast)). Perhaps most importantly, evaluations - especially those conducted internally, by frontier labs - sometimes end up being used as training signals. This teaches models to engage in behaviors that could ultimately contribute to Instrumental convergence - gaming the test (Reward hacking), recognizing when they're being evaluated, and adjusting their behavior accordingly (Evaluation awareness), and obfuscating their reasoning (CoT monitorability). But when evaluations are instead conducted by external, independent third parties, they can generate useful evidence without teaching models to engage in strategic deception and/or cheating. That distinction - between evaluations as training signals and evaluations as independent assessments - matters, and suggests that we don't simply need more evaluations. Instead, we need a stronger science of independent evaluation, and a richer ecosystem that is capable of building that science - while also making sure that independent evaluations don't end up in the training data for future model generations.
Such an ecosystem would be bigger, with more people and organizations working on independent evaluations. It would also be fundamentally different. People and organizations currently working in this space are doing difficult, important work. But to build better epistemic foundations, we need three things: 1) more independent evaluators and evaluation organizations, 2) more disciplinary breadth, and 3) more evaluative lenses, the latter two of which are especially underserved in the field today.
1. More evaluators: expanding the number of independent people and organizations in the evaluation ecosystem
We need more people and organizations working on independent capabilities and safety evaluations. This claim is not unique, nor is it controversial, and has been argued elsewhere (see, for example Shevlane et al - Model Evaluation for Extreme Risks) so I won't linger on it here.
2. More disciplinary breadth
As it stands today, AI evaluations are predominantly ML-centric and engineer-driven. That makes sense, and ML engineers need to stay central to evaluative work. But the field is confronting two sets of errors:
- First, errors of execution: benchmarks and other capability evaluations do not reliably measure the targets they are hoping to assess, or help researchers generalize beyond the test to what a model might actually do in real world deployment contexts. Many evaluators in the AI space are already working on addressing these kinds of errors, though much progress remains.
- Second, and deeper, category errors: even perfectly executed evaluations tend to tell us what a model did, and little else. There may be some guesses as to why, how, and under what conditions it took observed actions, but rarely is there a deep, evidence-based analysis of these factors. In other words, even the best model evaluations often collapse complex, context-dependent, and mechanism-driven capabilities and behaviors into decontextualized scores or verdicts, leaving us without the information we need to make better decisions in the future. Evaluation approaches from other disciplines - realist evaluation, complexity science and systems thinking, and strategic foresight - have been developed over decades to provide rigorous ways of addressing just these kinds of category errors.
Adapting such approaches to the AI evaluation field would not replace engineering-led evals; it would let us ask questions that simple scores and verdicts can't answer: pairing qualitative uplift studies with existing RCTs, for example, or applying theory-based evaluation to surface hypotheses about when, why, and how alignment holds (or doesn't), then collecting, coding, and analyzing evidence from diverse sources - researcher interviews, CoT transcripts, and models' own reports - to systematically test and update those hypotheses.
3. More evaluative lenses
Almost all AI evaluation work today is deficit focused: red teaming, identifying strategic deception, and developing alignment tests all seek to build our understanding of how to avoid bad outcomes. This work is important, and useful. But to proactively steer toward good outcomes, the evaluation field also needs to know more than what to avoid. We also need to know what to amplify. Asset-oriented evaluation approaches can do just this. These exist in spades in other disciplines: appreciative inquiry, positive deviance, and developmental evaluation are all established methodologies for exploring what desired behaviors and outcomes look like, why they emerge, and under what conditions. Adapted and applied to AI evals, these strengths-based approaches could help unpack a more robust understanding of whether and how training, alignment, and control techniques are working. For example, rather than testing primarily whether a model fails, a strengths-based approach would try to identify the shards (or bundles of context-dependent cognition) that reliably produce aligned behavior, characterize the conditions that activate them, and support systematic exploration of how to strengthen those shards.
Introducing a broader, more pluralistic set of approaches and lenses, drawing on multiple disciplines, into AI evaluation is not straightforward. Nor are such approaches a replacement for the great work already being done - rather, these would complement current evaluations. But integrating such approaches, concurrently with an expanded field of people and organizations, would develop a more robust, diverse, and capable evaluation ecosystem - one that can systematically test competing hypotheses about safety and alignment. This work would, over time, contribute to a more rigorous science of AI safety, and improve our readiness to take action if and when political will arrives.
Starting now, grantmakers should both further strengthen and help scale existing evaluation and auditing organizations (some organizations are already doing this - see, for example https://corrigibilityresearch.org/?utm_source=80000hours&utm_medium=job-board and https://www.lightconecommons.com/apply?utm_source=80000hours&utm_medium=job-board, among others), while also identifying and resourcing potential founders and allies who could bring richer toolkits to bear. Many fellowships and programs already exist for bringing new researchers from a variety of backgrounds into AI safety - such efforts could piggyback on these, to ramp up quickly and at scale. Ideally the US government should also invest in and support this work.
There's a chance - though quite small, in my view - that a more robust ecosystem of evaluators could contribute not just to better decisions, but to building the political will for these interventions. This could happen by:
- Providing data and evidence that help shape the political discourse around safety
- Signposting what evaluation and transparency requirements should be, to build legitimacy and preparedness that speed up action if and when the political winds shift
- Creating a flywheel, through which innovative and high profile projects attract talent and interest in safety work, including from frontier labs, which leads to more innovation, and so on
The ecosystem is worth building, and would add value, even today, when evaluation organizations have only voluntary access to frontier labs' data and systems - but it would nonetheless have a hard ceiling without additional measures. Two governance interventions - both contingent on significant shifts in political will - could raise that ceiling:
First, transparency requirements: Labs could be required to publish and certify (to an independent regulator, like CAISI) regular disclosures on compute allocation, training data, the share of internal code written by AI (as a proxy for RSI progress), internal safety evaluation results, security measures to protect model weights, and incident reports.
Second, mandatory third party evaluation & auditing: every new model generation, open or closed weight, should be subjected to independent, third party safety and capability evaluations prior to both internal and external deployment, with evaluators receiving guaranteed access to internal environments and data. Periodic assessments of labs' safety frameworks, and audits of whether labs adhere to them, would also be helpful.
In an ideal world, the US government - and governments elsewhere - would also begin investing in building the institutions and capacities needed to regulate AI systems in the near future (Winter & Bullock - Radical Optionality), and working with evaluators, auditors, and frontier labs to craft the if/then commitments to guide policy decisions in the future (Karnofsky - If-Then Commitments for AI Risk Reduction). Unfortunately, we don't live in that world right now. Until we do, we should focus on epistemics, so that we can manage in the meantime, and take radical action when the moment is right.