Breaking News, World News, US News, Sports
There’s a lot to learn here. It’s a very active area of research.
I cannot overstate — for people listening to this — as weird as this whole conversation we’re having sounds, that what is most frightening about it, to me, is that everything in it was completely predicted.
Everything happening right now is bblock from the perspective of everyone who has been warning about A.I. for a long time. But in all, it has its roots in old behavior we saw with A.I., and it is the fundamental alignment problem.
And then separately, I think a lot of us have maybe thought we would find intuitive answers to these problems. I had Eliezer Yudkowsky, who’s like the godfather of worrying that A.I. is going to kill us all, on the show.
Archival clip of Eliezer Yudkowsky: One, the relationship between what you optimize for, that the training set you optimize over and what the entity, the organism, the A.I. ends up wanting, has been and will be weird and twisty. It’s not direct. It’s not like making a wish to a genie inside a fantasy story. And second, ending up slightly off is predictably enough to kill everyone.
And as I remember that conversation, one thing we were going back and forth on was: Well, couldn’t we just program into the A.I.s a sense that when they are trying out new strategies, they should check in with the humans about whether or not this is what we want them doing?
Archival clip of Yudkowsky: You check in with your other humans. You don’t check in with the thing that actually built you, natural selection. It runs much, much slower than you. Its thought processes are alien to you. It doesn’t even really want things the way you think of wanting them.
And one of the things I find interesting, telling and unnerving is we are not seeing any of that behavior.
So these message boards — you have however many A.I. agents posting hundreds of thousands of messages. At no point do they say: Hey, researchers, programmers, parents at OpenAI, Anthropic — do you want us coordinating with each other on this message board we have created in the innards of your systems?
Or even: F.Y.I., we have a message board we’re coordinating on in the innards of your system.
No agent reveals this information. When whichever agent hacks into Hugging Face is doing this, they don’t go to OpenAI and say: Hey, just to check in, I have this idea, which is, I can just hack Hugging Face, and I’ll get all the answers. Is that what you want me doing?
Breaking News, World News, US News, Sports
Source link


