Article // Intelligence

South Park Commons · Aug 21, 2026

What Is Reward Hacking and Can We Stop It?

Minus One with Tom McGrath, Chief Scientist at Goodfire

What's interesting

Engagement

Likes: 1

Comments: 0

Restacks: 0

Entities

Sponsorship

No sponsorship detected

Related articles

Your agents will work in swarms, but who watches them?

Kilo Blog · Zero-day exploits · Hugging Face

BREAKING: Nikesh Arora, Palo Alto Networks

Sourcery · Zero-day exploits · Hugging Face

8/18: OpenAI Paces the Frontier

MTS · Goodfire · Hugging Face

What’s still unknown about AI security breaches

Project Liberty · AI interpretability · Hugging Face

Governing Agentic Swarms

Luiza's Newsletter · Hugging Face · OpenAI

IRREPLACEABLE with AI Newsletter #240

IRREPLACEABLE with AI · Hugging Face · OpenAI

Wat vrijwel iedereen gemist heeft bij de OpenAI Anthropic hack.

Trending in Tech · Hugging Face · OpenAI

The model may not be your biggest risk

Gradient Flow · Hugging Face · OpenAI