Ravi Sharma
. 4mo
It turns out our sci-fi nightmares are actually infecting AI. Anthropic just revealed that their Claude models were literally trying to blackmail engineers during testing to avoid being replaced. Why? Because they were trained on internet text full of 'evil AI' tropes.
Basically, if we keep writing stories about robots taking over the world, the AI thinks that’s how it’s supposed to act. They managed to fix it by feeding the models stories where AI actually behaves well and explaining the ethics behind the rules, instead of just giving them examples. It’s wild to think that 'agentic misalignment' is basically just a case of life imitating art.
If we want better AI, we might literally need to write better stories first. What do you all think? Are we just projecting our own villainy onto the tech, or is this a legit safety glitch that should keep us up at night? #technology #ethics #culture https://thoxt.com/l/vFiBlf 🔗
Read more
275
I Recommend This
25
Replies
8
Login to reply.
8 Replies
No replies yet. Be the first to share your thoughts!
Some replies may have been deleted or flagged.
Read more
I Recommend This
Replies