Theories of AGI “Values”
People who fear AGI destroying humanity often fear that AGI will not share human values. People who advocate for AGI soon often believe that AGI will naturally share human values….
This week, Scott Aaronson joins us on The Trajectory for episode 4 of the Worthy Successor series.
Scott is a theoretical computer scientist and Schlumberger Centennial Chair of Computer Science at the University of Texas at Austin who recently completed a year-long stint as an AGI researcher with OpenAI. I was influenced to invite Scott to the program after seeing his TEDx talk The Problem with Human Specialness in the Age of AI, and after Jaan Tallinn recommended Scott as a thinker worth following.
In this episode, Scott shares perspective on why AGI should evolve, not replace our existing human values. He discusses the possibility of a kind of “moral bedrock” that humans may already have access to, and which AGIs might expand upon.
I hope you enjoy this conversation with Scott Aaronson:
Subscribe for the latest episodes of The Trajectory:
Below, we’ll explore the core takeaways from the interview with Scott, including his list of Worthy Successor criteria, and his ideas about how to best leverage governance to improve the likelihood that whatever we create is, in fact, worthy.
Posthuman intelligences may be vast extensions of our human preferences, but such human values and preferences are still in some way present.
Their moral values have evolved from ours by some continuous process (rather than displacing said values entirely).
It would carry the flame of awareness/qualia on a new AGI torch.
We could scrub all training data of any mention of consciousness and see if the model is able to articulate what awareness is like, regardless.
(Scott doesn’t mention any specific form of governance in this series, but he advocates that it should be created based on the enormity of what’s being created – an event potentially more consequential than anything else on earth.)
Models requiring beyond some determined level of FLOPs might require some kind of registration process with a governing body, where it might be tested across a variety of criteria to determine risk.
…
I appreciated Scott’s firm emphasis on:
On the topic of morality and “values,” I concur with Scott that – especially initially – it would be important for AGI to branch off from our own human values, rather than boot up a new set of them entirely. Any kind of wholesale replacement risks losing not only things that are “uniquely human,” but things that might be adaptive and useful for future life in general – a sentiment that Bostrom shared in his Worthy Successor interview here on The Trajectory.
Regarding Scott’s notion that intelligence may “cap out” at moral ideas similar to those of humans, I disagree completely, as I suspect that a mind a billion times beyond our own, working on problems vastly beyond our own, housed in substrates and manifested in embodied forms vastly beyond our own, would likely have values wholly alien to us. I suspect that thinking AGI would – with any reasonable likelihood – converge on human values or human-friendly values is probably a dangerous idea.
All that said, Scott’s point about there being a kind of “moral bedrock” may have credence, and I suspect time will tell. Given Scott’s insistence on seeing things as they are, I’d guess that his perspectives will evolve as new experiments come in, as will mine.
What did you think of this episode with Scott?
Drop your comments on the YouTube video and let me know.
People who fear AGI destroying humanity often fear that AGI will not share human values. People who advocate for AGI soon often believe that AGI will naturally share human values….
This new installment of the Worthy Successor series features Ed Boyden, Y. Eva Tan Professor in Neurotechnology at MIT and a full member of the McGovern Institute for Brain Research….
For the sake of this essay I won’t be talking about morality as some abstract human access to “the good,” but simply as a set of heuristics of principles for…
A friend shared this cartoon with me recently: I laughed… but in a somber way. It’s funny because the second bird is totally limited in his actions, thoughts, and values…
This is an interview with Kristian Rönn, author, startup founder, and now CEO of Lucid, and AI hardware governance startup based in San Francisco. In this episode Kristian explores his…
This new installment of the Worthy Successor series is an interview with Joe Carlsmith, a senior advisor at Open Philanthropy, whose work spans AI alignment, moral uncertainty, and the philosophical…
This new installment of the Worthy Successor series is an interview with Robin Hanson – economist and author, Associate Professor of Economics at George Mason University, and a former Research…
This new installment of the Worthy Successor series is a conversation with Stephen Wolfram, founder and CEO of Wolfram Research, creator of Mathematica and the Wolfram Language, and a physicist…
Nick Bostrom – former Founding Director of the Future of Humanity Institute at Oxford – joins this week on The Trajectory. Bostrom has plenty of formal accolades, including being the…