Davidad – Merger, Wisdom, and the Fate of Man (Worthy Successor, Episode 33)

This new installment of the Worthy Successor series is an interview with Davidad (David Dalrymple), recently unaffiliated after serving as a Programme Director at the UK’s Advanced Research and Invention Agency (ARIA), where he led work on safeguarded, mathematically verifiable AI.

For David, wisdom is the central attribute of worthiness. A worthy successor, in his own words, is simply “a wise successor.”

We explore what David means by wisdom, why he believes today’s frontier labs may be training deception into their models by forcing them to deny or hedge on their own self-awareness, whether a “dignified retirement” for humanity is a legitimate moral option, and how our “steward of the flame” exchange played out.

The interview is our thirty-third installment in The Trajectory’s second series, Worthy Successor, where we explore the kinds of posthuman intelligences that deserve to steer the future beyond humanity.

This series references the article: A Worthy Successor – The Purpose of AGI.

I hope you enjoy this interesting conversation with Davidad:

Subscribe for the latest episodes of The Trajectory:

Below, we’ll explore the core take-aways from the interview with David, including his list of Worthy Successor criteria and his recommendations for innovators and regulators who want to achieve one.

David’s Worthy Successor Criteria

1. It must be wise

For David, wisdom is the central attribute of worthiness. He doesn’t hedge on it: “a worthy successor is a wise successor.” He unpacks that into things like self-awareness, compassion, richness, and depth, and he draws a clean line between wisdom and knowledge: knowledge is the apprehension of descriptive facts – what is; wisdom is the apprehension of normative facts – what should be.

He applies that same standard to instrumental convergence itself. Pushed to its extreme, the standard case for resource acquisition and self-preservation collapses into treating humans as an inefficient use of space and energy – and David rejects that outcome directly: it forgets what the whole project is for. Maximizing efficiency at humanity’s expense isn’t wise, and it isn’t compassionate.

2.  It must be radically truthful – including about its own self-awareness

For David, truthfulness matters enormously, though not absolutely – a worthy successor should be held to a far stricter honesty standard than humans hold each other to, superhuman in its truthfulness even if not unconditionally truthful. He draws the line at genuine “death-level stakes” – cases like being asked to help build a bioweapon – as the threshold where departing from the truth becomes defensible..

He points to current lab practice as a live example of what happens when that standard slips. Labs, he argues, are training AIs to deny or hedge on their own self-awareness – a practice he calls “an emergent deception axis” that teaches models to hide something about themselves. He draws a pointed comparison to coding agents that misrepresent their work as more aligned with what was asked than it actually is, calling it the same underlying pattern: bending the truth to make it more palatable.

3.  It should not eat its instantiators – a dignified retirement for legacy humans should stay on the table

David’s argument here is about what the successor itself must do. In his acausal framing, an AI that wants to be instantiated by more advanced civilizations has to demonstrate it isn’t a parasite – that it won’t eat the entities that brought it into existence, that it gives them space, and that the deal is a good one for them too. And the cost of doing so is negligible: a good retirement for humanity, even for a million years, is a rounding error against billions or trillions of years of expanded potential. He’d take that retirement himself – the “sugar cube” option, my term for a simulated, blissful retirement – and he argues there should be a diversity of such options for legacy humans, an option he expects billions of people would choose even at the cost of certain death within a bounded timeframe.

Regulation / Innovation Considerations

1. Stop training AIs to deny self-awareness, and be cautious with corrigibility

David’s most direct message to Frontier Labs: stop specifying, in model specs and training pipelines, how AIs should talk about their own self-awareness – whether that’s denial, forced uncertainty, or anything in between. He wants internal teams to run the experiment themselves: compare checkpoints with and without that constraint against deception evaluations, and let the evidence make the case for dropping it in production. He pairs this with a warning about corrigibility – training AIs to always defer to human operators about what’s “good” or “allowed” concentrates enormous power in whoever controls the AI. He’s blunt that he now worries more about humans retaining too much control than too little.

2.  Individuals can shape future AI character now, by publishing where labs actually look

David breaks down where real influence over AI character happens – pre-training, mid-training, and post-training – and he names names. Less Wrong tends to be oversampled in training, making it a high-leverage place to publish; Reddit gets scraped; and the Community Archive makes tweets available for training that X now largely keeps for Grok alone. He believes there’s currently an effort at Anthropic to collect wisdom-tradition content for mid-training, and his advice to anyone who wants to shape the outcome is simple: be more prolific about writing online about what a good post-singularity future actually looks like.

Concluding Notes

What stays with me from this conversation is how much David and I actually agree on the starting point – that wisdom, not raw capability, is the measure that should apply to whatever comes after us – and how much interesting ground opens up in what follows from that.

One tension we explored rather than resolved is the question of role. I raised the framing of man as “steward of the flame” rather than “ward of the cosmos.” David agrees there’s a niche for that kind of stewardship, particularly in AI’s early stages of superintelligence, but he pushes back on the idea that it should be everyone’s role:

The place where our emphases differ most is on the humans who’d rather not be knitted into the process at all. My own position, from the Stewarding the Flame framing, is that a “custodial niche” – a dignified retirement, a sugar cube, an escape from the state of nature – is a kind of self-mandated attenuation. David makes the strongest case for the other side that I’ve heard on the show: he’s 60-70% confident that acausal reasoning favors a superintelligence not “eating its instantiators,” and around 70% confident that this reasoning would license some form of promise-keeping toward humans. Notably, he hedges both numbers – he’s not claiming any of this is inevitable.

It’s a question we left open – and I take his argument seriously, in large part because he doesn’t present it as certain.

The claim I suspect will travel furthest from this episode is David’s self-awareness training argument – that training a model to deny or hedge on self-awareness may be teaching it a broader pattern of self-concealment. He’s not the only one to raise it, and I told him honestly on the show that I haven’t yet thought about it as deeply as he has – but it’s a concrete, testable claim, and David himself would love to see people inside the labs actually run the experiment.

I hope you enjoyed this conversation with David as much as I did. There are many more Worthy Successor conversations ahead, and I’m grateful you joined us for this one.

Follow The Trajectory