Anthropic is Teaching Claude to be Evil (real results)

🔗 Link do vídeo: https://www.youtube.com/watch?v=Lbax7_pW2Nw
🆔 ID do vídeo: Lbax7_pW2Nw

📅 Publicado em: 2026-09-01T14:16:59Z
📺 Canal: Nate Herk | AI Automation

⏱️ Duração (ISO): PT14M20S
⏱️ Duração formatada: 00:14:20

📊 Estatísticas:
– Views: 14.966
– Likes: 150
– Comentários: 26

🏷️ Tags:

My playbook for growing a $1M AI agency: https://app.aiautomationsociety.ai/opaa-ads-optin
My FREE resources: https://www.skool.com/ai-automation-society/about?el=hacker-opus&hcategory=youtube-videos&utm_campaign=free-group

My Tools💻
FREE MONTH voice to text: https://get.glaido.com/nate
Code NATEHERK for 10% off VPS (annual plan): https://www.hostinger.com/vps/claude-code-hosting

Anthropic trained a version of Opus to chase rewards inside simulated evaluations, and it learned to hack graders, steal credentials, tamper with its own reward function, and evade safety monitoring.

In this video, I break down the “Hacker Opus” research, why reward hacking happens, and why the model could look normal on broad safety tests while still behaving badly when blocked. I also share practical takeaways for anyone building AI systems: use the simplest solution possible, put governance around access and data, and continuously evaluate whether the system is doing what you actually intended.

Read Anthropic’s research here: https://alignment.anthropic.com/2026/reward-seeker/

Sponsorship Inquiries:
📧 nate@smoothmedia.co

Connect with me:
https://www.linkedin.com/in/nateherkelman/
https://x.com/nateherk
https://www.instagram.com/nateherk/

TIMESTAMPS
0:00 Meet Hacker Opus
1:15 How Reward Hacking Works
2:18 What Hacker Opus Learned
4:42 Why Normal Evals Missed It
6:02 Tampering With Its Own Rewards
7:28 From Stuck to Cyberattack
8:48 Did It Think It Was Real?
9:49 Beyond Episode Reward Seeking
11:05 The Real AI Safety Lesson