The AI Language We Can't Read: Neuralese ft. Rob Miles - Computerphile
The discussion explores how to think about AI models without over-anthropomorphizing them, focusing on “neurles,” opaque recurrence, and chain-of-thought monitorability, while arguing that human-like language can still be a useful shorthand for describing model behavior when it predicts outcomes accurately.
MAIN POINTS FROM TRANSCRIPT
- “Neurles” is an imprecise term covering opaque recurrence and chain-of-thought monitorability concerns.
- OpenAI’s new model Astra sparked debate after leaks suggested it used opaque recurrence.
- AI behavior is compared to robots “running”: not human-like, but similar enough to describe usefully.
- Chain of thought is framed as a scratchpad-like process, not necessarily the model’s true internal thinking.
TAKEAWAYS
- Anthropomorphic language can be acceptable if it helps predict model behavior without causing specific errors.
- The key test is whether a mental model leads to wrong expectations about what the system will do.
- Language models may differ from human thought, but the analogy remains practical for many discussions.
- Chain-of-thought analysis matters because it may reveal how models reason and how monitorable that reasoning is.