Discussion about this post

User's avatar
Shaked Koplewitz's avatar

You mention updating on some reports from anthropic, but the huggingface incident seems like a pretty clear cut case of AIs wanting things (in this case, to pass a test); you can argue it may be possible to build AI that doesn't and this isn't certain, but it at least seems clear that this is a likely enough outcome that we need to be very sure in can't happen with superhuman AI.

On the topic of twitterisms, I'm much less sure than you that alienates politicians; they sure seem to spend a lot of time on Twitter these days (maybe not the serious natsec guys you want though? I am not a government whisperer so low confidence on this part).

Tóth Csaba Dr's avatar

I loved the summary of the argument of the book and the discussion on its conjunctive nature, I never actually heard it stated like this. I would be happy to read more on what you think of the other points, I always felt the jump from 3 to 4 was contradictory (if we don't know what preferences, how do we know...).

6 more comments...

No posts

Ready for more?