Even some of the best AI can’t beat this new benchmark
The Center for AI Safety and Scale AI have introduced "Humanity's Last Exam," a benchmark featuring crowdsourced questions to challenge frontier AI systems in various subjects.
MAIN POINTS
- "Humanity's Last Exam" is a new benchmark for testing AI systems.
- It includes thousands of crowdsourced questions.
- Subjects covered are mathematics, humanities, and natural sciences.
- Developed by the Center for AI Safety and Scale AI.
TAKEAWAYS
- The benchmark aims to evaluate the capabilities of advanced AI systems.
- Crowdsourcing ensures a diverse range of questions and perspectives.
- Collaboration between nonprofit and commercial entities highlights the importance of AI safety.
- This initiative could guide future AI development and safety standards.