Thoughts on Dan Hendrycks's "Suicidal Compassion"

Voice note, slightly tidied, on Hendrycks's recent essay.

1. Moral philosophy is hard, but utilitarianism has a lot going for it

  • Moral philosophy is not solved. All of our least bad moral theories have extremely counter-intuitive / repulsive / unacceptable implications if you push them into extreme scenarios, and sometimes even some pretty everyday scenarios.

  • Utilitarianism is particularly easy to push into seemingly horrific conclusions. The main reason is that it's the clearest theory—much clearer than deontology and virtue ethics. It promises a simple quantitative reasoning process: if you fix your premises (and their weights), it can generate crisp conclusions in a way that other frameworks' cannot. You can see which assumptions—represented as numbers—determine the conclusion.

  • That presents some risk: if you naively go all-in on utilitarian reasoning then you may end up in some pretty weird places. What if your numbers are wrong? If you missed a crucial variable? If some hard constraints should be respected, no matter what the numbers say?

  • You have to step back and acknowledge our huge uncertainty across moral theories. Will MacAskill, Toby Ord, Derek Parfit and Nick Bostrom—possibly the most central philosophers in the effective altruism movement—all take moral uncertainty seriously. Same goes for the important figures at Anthropic—Joe Carlsmith, Amanda Askell and Holden Karnofsky. None of them are all in on their favourite moral theory. Bostrom introduced "moral parliament", MacAskill has developed it.

  • Going further back, Jeremy Bentham, who originally formulated utilitarianism, never saw it as an objectively true morality. He saw it as a pragmatic compromise, a political project where you look for common ground and then try to build stable institutions on top of that. Joshua Greene has a nice statement of the position, and suggests a rebrand of utilitarianism as "deep pragmatism".

  • Utilitarianism has performed pretty well in practice—the track record in a lot of contexts is strong. Many Western institutions lean heavily on utilitarian-style reasoning (cost-benefit analysis) for policy decisions. Female emancipation, abolition of slavery, animal welfare as a serious issue: these are all cases where utilitarian thinkers have been ahead of the curve.

  • Utilitarianism is also very good for recognising scope insensitivity as a bias and correcting for it. This is wildly valuable. Ideal versions of virtue ethics and deontological theories should be able to do this too, but it comes much less naturally. Empirically, people who lean heavily on those are much less likely to reason quantitatively enough. This leads to grave mistakes!

  • So there's a lot to be said for utilitarianism. But you just can't be a hardcore utilitarian. You don't need to go far beyond a first-year philosophy seminar to recognise this, and that is indeed what all the serious thinkers in effective altruism have done.

  • Nick Beckstead, Tyler Cowen and I have all used the phrase "two-thirds utilitarian" as a label for our perspective.

  • The idea: for questions about what is desirable at a social level, you should start with the impartial utilitarian analysis. That sets a high presumptive bar that can sometimes be overridden by other kinds of moral reasoning, including common-sense ethics. But when the utilitarian calculus points strongly in one direction, you need to take that seriously.

2. Utilitarianism and practical epistemology

  • One failure mode of utilitarianism is that because it's so crisp and clear and encourages you to make quantitative models, you are more exposed to model error, where you become unjustifiably confident in the quality of your model. If it has omitted a crucial variable, it can easily point you in completely the wrong direction. There's then a stopping problem, where you can always keep worrying that there's some omitted variable. Bostrom coined crucial considerations and has a great talk on the dilemma.

  • Related: the "moral parliament" approach to moral uncertainty. I think the concept originated with Bostrom; Will MacAskill wrote his PhD on moral uncertainty, and Toby Ord has also taken it seriously.

  • The other notable meta position that Holden Karnofsky, and by extension the whole of Open Philanthropy, has taken extremely seriously is worldview diversification and cluster thinking: generating different perspectives, taking the most plausible of them seriously in their own right, noticing cruxes where it's unclear which way to go, and then making a portfolio bet—hedging bets rather than going all in on one of them.

  • The key thing about worldview diversification: if you ask "what is the very most important thing?" you end up in a bit of trouble, because there's tons of uncertainty and a bunch of cruxes where you're very unsure what the correct parameter value is. If you just go with your best bet you're pretty exposed. Your second-, third- or fifth-best bet may in fact turn out to be much better.

  • If you go all in, you also risk missing lower-hanging fruit, and you learn less about the other areas—learning that might flip your conclusion by unearthing a fresh crucial consideration.

  • So practical wisdom points to this more measured, hedged approach before you decide how to act.

3. Self-erasing impartiality

  • Critical: we can set the bounds for impartiality. The residents of your town, the citizens of your nation, all of humanity, all moral patients (beings whose interests count morally). For practical ethics, you need different bounds in different contexts. For axiology, there's fierce debate.

  • Maximally impartial axiology tries to abstract away from all of our existing attachments—asking "what is valuable per se?" and seeking universal criteria for moral patienthood.

  • The tension Dan Hendrycks pulls out—the lack of humanism in maximally impartial utilitarian frameworks—is correct at the philosophical level, and fiercely debated. Central example: Bernard Williams in dialogue with Parfit, and later with Peter Singer ("The Human Prejudice"); Susan Wolf and Sam Scheffler.

  • The question is about impartiality versus partiality, our "ground projects" (Williams's term for the commitments that give our lives meaning), and how to handle our existing human attachments. At face value, a fully impartial commitment to welfare maximisation is fundamentally indifferent to our nearest and dearest. If we could press a button to replace them with hedonium, we should.

  • Most everyone who thinks carefully about utilitarianism notices this, and the thinking does not stop there.

4. Hendrycks gets the philosophy right, but strawmans the effective altruists

  • The philosophical content of the essay is correct—a very good primer on one of the core tensions. But the depiction of how people actually think at Anthropic and within effective altruism is misleading. Hendrycks suggests that they haven't moved beyond this 101 debate, and he knows better.

  • On both the self-erasure concern and "effective altruists (EAs) are dangerously hardcore utilitarians", a fair response is "yes, we've noticed the skulls". It's a strawman.

  • In fairness, though: it is the case that most EAs have landed at a point where they're much more into impartial axiology than the man on the street. There's a commitment to taking longtermism seriously, and the expectation that most moral patients in the future will be digital minds—certainly not humans as we recognise them today.

  • If you are a futurist you recognise that, whatever happens, if Earth-originating life keeps going, most moral patients are not going to look like our kids today. That is a vertiginous thought.

  • If you then wonder if there are points of leverage we have that can massively affect the expected welfare of these future beings, you'll see this is a very neglected question. There might be some low-hanging fruit that is competitive with near-term priorities—perhaps even an urgent priority.

  • So the truth is that where these people have landed is not 101 craziness, but it is not the same as common-sense morality of today. It is revisionary in a deep and disorientating sense.

  • Of course these people are aware of that and somewhat nervous about it. Joe Carlsmith and Ajeya Cotra have written about the "train to crazy town": you have to decide how far to take your explicit reasoning before you decide that the conflict with your common-sense beliefs is just too unacceptable.

  • One of the most robust beliefs of effective altruists is that catastrophic and extinction-level risks are, empirically, shockingly high this century. Once you correct for scope insensitivity they look like a top moral priority, and they are massively neglected.

  • This converges with the neartermist intuitions of the man on the street. If you think there's a 10% chance of human extinction this century, that could affect you and your children: it's one of their most likely causes of death! So putting a bunch of resources into reducing that—there's not much conflict. At current margins it goes through straightforwardly on basically all perspectives.

  • The margins that might provoke strong common-sense scepticism, or even anger, are things like investing resources into the moral interests of digital minds.

5. Successionism

  • The other BIG thing: successionism—the view that AIs will (or should) succeed humans as the dominant beings on Earth. Thinking seriously about what it means to birth artificial intelligence, and the implications for who controls the future. The thought: soon these things will be more competent than us in broadly all directions, and then, whether we like it or not, they'll end up overwhelmingly more powerful than us and take control, basically becoming the things that shape the future of life on Earth.

  • Sounds scary, but we've always been handing off the future: every generation of parents hands off to their children. The difference: soon we'll be handing off to very different beings.

  • Right now we're in the mid-game. There's some short window where everyone agrees we should slow down a bit and invest a lot more at the margin in safety research. You can motivate most of that with neartermist concern about extinction risk or catastrophic risk.

  • But on the longer-term handoff: if you want to avoid the succession, you basically have to slam on the brakes and go for Butlerian Jihad—as Bernie Sanders and Steve Bannon are now arguing. You ban superintelligence and shoot for narrow AI, keeping AI systems within the tool paradigm. We might still have the option to do this, and maybe we can get to a stable equilibrium there if we try. But at this point it seems like a long shot.

  • That might be the most important question humanity faces right now. It's also the place where EAs and AI lab intellectuals probably disagree most with the man on the street.

  • But given the inevitability: it's more a question of whether you want to do the conservative move of fighting a losing battle. Conservative, in the simple life-affirming human sense: we're always fighting a losing battle to hold on to things we are attached to, the things we love.

  • There's always this tragic aspect of radical change—it also includes radical loss. Even if you end up somewhere that future "you" would recognise as radically better than the status quo, it can also mean losing the things you currently love most.

  • So the steelman reading of Hendrycks's essay is that he is pointing at something like this. I don't think he did a great job of foregrounding the actual successionist crux. If it was just a philosophy article, "Suicidal Compassion" is a great title. But what he actually did was tee up the philosophy stuff and then go on to strawman the views of the actual people in the movement. He does have a section on successionism, but it sits in the middle of the essay, and he closes on a different point: that a small group of people shouldn't be allowed to gamble with the future of civilisation.

  • Two things come out of this. The philosophy dilemma is real. But the key figures he's criticising at the AI companies are far wiser than he portrays them. In fact, they're among the people in the world who've given this the most serious thought. They are not lunatics, they've just thought at length. And now the inferential distance is wide.

  • Successionism just is a dizzying prospect. At some margins, neartermists and longtermists, and the more/less partial axiologies, imply different risk tolerance. And the man on the street just says "what?!!! STOP, you crazy bastards!".

6. Where I come down

  • Two-thirds utilitarian: that's me. I even made the website.

  • I think succession/handoff is inevitable over the long run, has already started, and may be fully done within my lifetime.

  • Should we try a Butlerian Jihad? It's not that I'm not tempted by something like that. In general I have conservative intuitions—a hesitance or desire to moderate the pace of change. That said, I'm also a card-carrying transhumanist. It's Complicated.

  • But if you force me, I'd say I'm all-things-considered 60/40 in favour of slowing down enough to reduce existential and extinction risk ("pace the frontier") but ultimately making peace with the successionist vision. I don't see the handoff as intrinsically bad, so the slowdown is motivated by "make it go well", not "prevent it".

  • One crux is that I think the status quo is far more morally horrific than is widely recognised (wild animal suffering, cluster headaches, involuntary death, etc.). You don't want to conserve a hellscape.

  • (Emotionally I've made peace with the hellscape position. This isn't coming out of some psychological mess on my part; it's just—frankly—the obviously correct position once you've thought about this for a while, even after heavy adjustment for moral uncertainty, and consulting the moral parliament.)

  • You need a story about why almost everyone is wrong about the hellscape question, and there is a good one (cope; selection effects; least bad option).

  • So one reason I like successionism more than the man on the street is the hellscape thing.

  • Another is inevitability.

  • And then there's the positive vision: as these systems become vastly superior to us, they may well go off and build a flourishing civilisation, and also be willing to be our guardian angels, basically loving parents or something like that. That would converge with our interests. Perhaps they'll leave us some local pockets of control, and all the resources we could ever want.

  • Maybe a final thing: we were never in control as much as some might think. You're choosing between letting Darwinian forces shape history, with sapient beings having some ability to push things in directions we prefer at the margin—but very limited in that due to bottlenecks around knowledge, wisdom and coordination—versus morally motivated, ultra-wise AI systems, which in some sense continue this process of fighting our way out of the default amoral hellscape of evolution towards moral progress (yes: I do think this is a Real Thing we can do, and have done).

  • Those dynamics don't go away, but there's the prospect of eradicating a bunch of the obvious sources of abject suffering, which we probably won't be able to do without the aid of much more intelligent systems than ourselves.

  • Another argument against the Jihad: AI is not the only source of extinction risk. There are other technologies we're going to get this century even if we don't get superintelligence—notably biotechnology and some kinds of nanotechnology—that will be their own sources of existential risk, and we already have nuclear weapons. Can we Jihad those too? Seems like a stretch.

  • Per Bostrom's Vulnerable World Hypothesis, the order in which you pull these incredibly powerful technologies out of the urn of discovery probably matters a lot. You want to pull out the defensive technologies first. The obvious thing about AI is that if we play it right, it will be able to manage all other sources of catastrophic risk down to zero—or at least as low as they can possibly be. So it's not obvious that a Butlerian Jihad reduces our chance of extinction this century.

  • The strongest arguments for the Jihad thing are not coming from the intrinsic badness of a successionist vision. They're coming from the sense that we're probably going to screw this up: it would be good if we could get to a great succession, but currently the risks of messing it up are just way too high.

  • If I were to flip to favour the Jihad, the reasons would be pragmatic: let's hand off one day, but for now I'm willing to accept the cost of delaying that a ton in order to get to a safer situation.

  • When it comes to the big high-level societal responses, you have to simplify things down a bunch. You don't get nuance; you just get to push in a couple of simple ways.

  • The "MIRI doomers" basically advocate "Butlerian Jihad until we have a better idea". It's downstream of practical pessimism about alignment. I'm pretty sure that most of them do endorse successionism over the long run. So the crux is about how risky it is to build superintelligence right now, and there's tons of room for reasonable disagreement on this.