Article // Intelligence
South Park Commons · Aug 21, 2026
What Is Reward Hacking and Can We Stop It?
Minus One with Tom McGrath, Chief Scientist at Goodfire
What's interesting
Engagement
Likes: 1
Comments: 0
Restacks: 0
Entities
Topics
People
Companies
Guests
Sponsorship
No sponsorship detected
Related articles
Your agents will work in swarms, but who watches them?
Kilo Blog · Zero-day exploits · Hugging Face
BREAKING: Nikesh Arora, Palo Alto Networks
Sourcery · Zero-day exploits · Hugging Face
8/18: OpenAI Paces the Frontier
MTS · Goodfire · Hugging Face
What’s still unknown about AI security breaches
Project Liberty · AI interpretability · Hugging Face
Governing Agentic Swarms
Luiza's Newsletter · Hugging Face · OpenAI
IRREPLACEABLE with AI Newsletter #240
IRREPLACEABLE with AI · Hugging Face · OpenAI
Wat vrijwel iedereen gemist heeft bij de OpenAI Anthropic hack.
Trending in Tech · Hugging Face · OpenAI
The model may not be your biggest risk
Gradient Flow · Hugging Face · OpenAI