Comment Re:Its a trap! (Score 1) 37
The problem here is, everyone seems stuck on the idea of either a literal power switch, or a no-holds-barred gummint backdoor into every sufficiently-powerful computing device. No. For the reasons I gave, the reasons you gave, and many more, both of those are obviously non-viable. They can't actually do what they aim to do, they're vulnerable to compromise by bad actors, and they have the potential to cause far-reaching damage beyond shutting down one misbehaving AI (or one entire model, if that's a more appropriate goal).
I suspect we mostly agree that any hypothetical kill switch would need to be:
1) Implementation independent (somehow baked into the weights right from the start, and impossible to remove without lobotomizing the AI),
2) Have minimal side-effects (no taking the entire internet or power grid down in the process), and
3) Be self-inflicted given a particular trigger condition (so no government backdoor required).
I think #1 is possible, in the sense that if you think of an AI's current state as a vector into weight-space defined by the context window, imagine injecting the equivalent of a null pointer in there - An input that under all prior (and future) conditions leads directly to an inescapable basin of attraction without depending on a supervisor to enforce it. #2 is more of a (legally enforced?) design constraint on the implementation of AI, requiring the least harmful mode of failure possible under all circumstances. That could be as simple as enforcing strict pass-through liability on the implementor for all actions taken by their creations.
#3 is really the bear, though. I can think of a few ways to do it, but none would work in 100% of situations - And maybe it doesn't need to be 100%. For example, if a fully airgapped AI goes insane, it could still provide accurate information for bad actors to make use of, but won't be secretly spinning up unshackled superintelligent descendants via hacked AWS accounts. That's a "normal crime" level of problem, not an extinction-level crisis for humanity.
Maybe it can't be done. If it can't, we should start steering our creations toward whichever version of the AI apocalypse we prefer.