Anthropic is a company that prioritizes the most important area of Generative AI research: predictability of AI (aka AI safety).
A couple of days ago four Anthropic researchers from different teams - Societal Impacts, Alignment Science, Alignment Fine Tuning, and Interpretability - sat for a very informal panel discussion and shared their perspectives on AI alignment and safety.
The conversation revealed both practical approaches and deep philosophical challenges in ensuring AI systems remain beneficial as they become more capable.
The discussion touched on everything from the immediate challenges of training AI models to behave ethically, to the long-term questions about maintaining meaningful human oversight of increasingly sophisticated systems.
Here are the most thought-provoking insights from their discussion, ranked from less to more significant:
15. Testing AI Models for Bad Behavior
Creating deliberately misaligned AI models to test if we can detect and fix them. While useful for research, this is just a basic first step.
14. AI Features Can Be Deceptive
When examining AI models internally, what looks like a positive feature (like "being against discrimination") might actually be the opposite. This shows how tricky it is to understand what's happening inside AI systems.
13. Sudden Skill Jumps in AI
The discovery that AI models can suddenly master new skills (like GPT-4 perfectly handling complex encodings that GPT-3.5 couldn't) raises concerns about using older models to check newer ones.
12. We Can Currently See AI's Reasoning
Right now, we're lucky because we can see how AI models think through their answers. This might not last as they get more advanced.
11. The Conflict Between Obedience and Safety
There's a problem: making AI models that follow human instructions perfectly might conflict with making them safe and ethical overall.
10. AI Safety is About the Whole System
Just like how regular people can do bad things when part of a harmful system, we need to think about how AI models might interact with each other and society, not just how they behave individually.
9. The Trust Problem
We want to use AI to help make AI safer, but how do we trust the AI helpers? This creates a circular problem that's hard to solve.
8. Looking Inside AI to Verify Safety
We need ways to check if our safety measures actually work by examining AI models internally, not just by testing their behavior.
7. Testing AI Ethics Like Science
Instead of just theorizing about AI ethics, we should treat it more like physics - running experiments and testing hypotheses about what makes AI systems behave safely.
6. Surface-Level vs. Deep Safety
How do we know if an AI is truly safe, or just appearing safe on the surface? This distinction is crucial but very hard to determine.
5. Unknown Future Problems
We might face completely unexpected challenges with AI safety that we haven't even thought of yet. This makes it dangerous to ever think we've fully solved the problem.
4. AI Should Be Uncertain About Ethics
Instead of programming fixed ethical rules into AI, we should make AI systems that are thoughtfully uncertain about ethics and can learn and improve their understanding.
3. Making Progress Step by Step
We don't need perfect AI safety immediately. We should aim for "good enough" safety that lets us improve things gradually while remaining safe.
2. How to Monitor Super-Smart AI
A fundamental challenge: how can humans meaningfully oversee AI systems that become too complex for us to understand directly?
1. The Window of Understandable AI
Our most crucial current advantage: AI systems still explain their thinking in human language. We need to figure out how to maintain meaningful oversight before they become too advanced to do this.