Thickly Settled

Applying systems thinking and evaluation science to AI safety: four demonstrations

Over the course of my career I've worked at the intersection of systems thinking, evaluation, and strategy - guiding my partners as they seek to better understand, navigate, and shape complex systems. I have helped foundations, governments, researchers, and nonprofits - working across disciplines, in dozens of countries - surface tacit assumptions, test their thinking and practice against diverse, rich evidence, and systematically figure out the conditions and mechanisms under which their desired outcomes are likely to emerge, or not. Evaluation as a field has spent decades building a robust toolkit for just this work, from realist and developmental evaluation approaches to complexity and systems thinking, to strategic foresight. And much of this toolkit has not yet been adapted to AI safety.

Here are four concrete demonstrations of how that toolkit could complement existing ML-driven approaches to AI evaluation and alignment research, all of which could stand as independent proposals OR be applied within existing organizations. They share an underlying method: surface the implicit assumptions, make explicit what would confirm or disconfirm those assumptions, gather evidence across traditional and nontraditional sources, and synthesize to systematically assess not just what is happening, but how, why, and under what conditions.

  1. Weak signals about the pace of progress. Combine systems thinking and strategic foresight to identify the field's key assumptions - including those held by researchers in China and the Global South - about what is driving the pace of AI system development, then set up monitoring systems to regularly collect, surface, and use weak signals to test and update those assumptions over time. This would enable the field to systematically assess which key assumptions about the pace of progress are holding, and which aren't, thereby giving us a more grounded, continuously updated basis for reasoning about when, for example, AGI is likely to emerge, and in what ways.

  2. Richer uplift studies. Take a richer approach to user-focused uplift studies, systematically collecting and synthesizing qualitative evidence alongside existing RCT-based approaches, to better unpack whether and how user capabilities improve when working with AI systems, under what conditions, and via what mechanisms. This would give the field a more robust sense of whether uplift is generalizable across individuals and fields, whether gains are more likely to be concentrated among a small set of users and domains, and where the greatest risks and opportunities lie.

  3. Strengthening the alignment evidence base. Apply realist and developmental evaluation to alignment research itself: work with an alignment team to surface and structure their implicit assumptions about when, how, and under what conditions alignment holds or doesn't; test those assumptions by collecting confirming and disconfirming evidence across a range of traditional and nontraditional sources (existing benchmarks and metrics, interrogation of the say-versus-do gap in models' own self-reports, and the tacit insights of the research team); and synthesize findings across them. This would enable a team to see which of its alignment assumptions are holding, which aren't, and under what conditions alignment is most likely to break, helping to build an evidence base for a more robust science of alignment over time, and, in doing so, driving sharper, better-targeted alignment evaluations for both existing and future models.

  4. Realist-informed transcript analysis. Develop a codebook based on a set of hypotheses about the various factors that may have misaligned behavior, captured in an existing eval and/or incident report (such as the METR investigation into the Hugging Face attach). Rigorously code agent transcripts (full COT, tool calls, etc.) according to that codebook (rather than, for example, running more general regex classifications). Analyze all the coded excerpts to pull out crosscutting patterns and themes about agent behavior, then distill findings and conclusions that unpack the conditions that drove misaligned behavior, - such that we could then use those findings to develop and test hypotheses to inform decisions about how to build internal RL and eval environments to mitigate such behavior in the future.

Taken together, these are not separate ideas so much as four applications of the same discipline, and aim to showcase how systems-informed evaluation has real, concrete, and complementary value to offer AI safety work, alongside the technical methods already in use.

#evals #strategy