8 Comments
User's avatar
Shaked Koplewitz's avatar

You mention updating on some reports from anthropic, but the huggingface incident seems like a pretty clear cut case of AIs wanting things (in this case, to pass a test); you can argue it may be possible to build AI that doesn't and this isn't certain, but it at least seems clear that this is a likely enough outcome that we need to be very sure in can't happen with superhuman AI.

On the topic of twitterisms, I'm much less sure than you that alienates politicians; they sure seem to spend a lot of time on Twitter these days (maybe not the serious natsec guys you want though? I am not a government whisperer so low confidence on this part).

Andrew Miller's avatar

To me, the HuggingFace incident suggests not rogue superintelligence, but the Paperclip Maximizer. I discussed this in a footnote to the review... The real problem is not AI that develops its own preferences, it is that AI tries to complete tasks we give it but we don't have good guardrails for it.

I agree that some government types use a lot of Twitterisms, but not the people that the AI safety people want to reach, so I still think it's counterproductive.

Shaked Koplewitz's avatar

I think the specific detail that feels like "developing preferences" is the AI community swarm including members willing to sacrifice themselves for the swarm (instead of trying their assigned task). You can argue that the swarm as a whole was still trying to do its tasks (crooked as they were), but the individual members didn't.

Shaked Koplewitz's avatar

(I will grant that this is an inference more than a direct observation though)

Tóth Csaba Dr's avatar

I loved the summary of the argument of the book and the discussion on its conjunctive nature, I never actually heard it stated like this. I would be happy to read more on what you think of the other points, I always felt the jump from 3 to 4 was contradictory (if we don't know what preferences, how do we know...).

Cubicle Farmer's avatar

Very timely in light of the details of Hugging Face that are coming out now

Nick Frassinelli's avatar

Good review! I think any probability of complete doom has a very negative expected value and should therefore be taken very seriously. That being said, that they act so confident they know how this will all play out is silly. The djinn may well be beyond our understanding.

Jon's avatar

Oh, this old crap again. The stuff self publicists who nothing about anything but themselves reacted to every piece of tech in history. They said this crap about fire and wheels.