P
Pac-Bench
Open Source๐ Alt to SWE-benchTest LLM spatial reasoning by one-shot building a Pac-Man game
๐ณ Self-Hostable๐ No Sign-upโก Traction Score: 71/100โ
17 Stars
git clone https://github.com/jonclegg/pacman-bakeoff.gitPac-Bench is an interactive evaluation framework designed to test frontier models on complete game development and spatial reasoning tasks. It challenges LLMs to one-shot build a fully working Pac-Man game, providing a standardized benchmark for complex code generation.
Challenges models to generate an entire playable Pac-Man game in a single prompt response.
Evaluates how well models understand coordinate geometry, collision detection, and grid mechanics.
Provides instant feedback loops to verify if the generated code results in an executable, playable game.
Benchmarking the spatial reasoning and code synthesis limits of newly released frontier LLMs
Comparing coding agent capabilities across complex, multi-component application development
Researching failure modes of LLMs when handling game loop architecture and state management
Unlike SWE-bench which focuses on real-world GitHub issues, Pac-Bench specifically targets spatial reasoning and end-to-end creative code generation in a controlled sandbox environment.
AI researchers, ML engineers, and developer tool creators evaluating frontier code generation models.
Compare other trending developer tools and open-source projects in this space.