Core Navigation
โšก All Radar Feed๐Ÿค– AI Agents & Workflows๐Ÿง  AI & Machine Learning๐Ÿ’ป DevTools & CLI๐Ÿ”„ Open Source Alternatives๐Ÿ“ฆ Frameworks & Libraries๐Ÿ—„๏ธ Database & Storageโ˜๏ธ DevOps & Cloud๐Ÿ›ก๏ธ Security & Pentestingโšก Productivity & Workflow๐ŸŽจ Design & Frontend๐Ÿงช Testing & Benchmarks๐ŸŒ APIs & Web Scraping
Directory & Community
โ„น๏ธ About ToolsRadar+ Submit a Tool๐Ÿ“œ Privacy Policy๐Ÿ™ GitHub Source Code โ†—
Pac-Bench logo

Pac-Bench

Open Source๐Ÿ”„ Alt to SWE-bench

Test LLM spatial reasoning by one-shot building a Pac-Man game

๐Ÿณ Self-Hostableโšก Traction Score: 71/100โ˜…17 Stars
๐Ÿ’กAnalyst Verdict & Strategic Take
AI Editorial Assessment
"An exceptionally creative and rigorous benchmark that pushes LLMs beyond basic code completion into multi-file spatial coordination. Essential for AI researchers evaluating true frontier reasoning capabilities."
๐Ÿ”’https://jonclegg.github.io
Open Site โ†—
Live Web Application

Pac-Bench

Test LLM spatial reasoning by one-shot building a Pac-Man game

โšก

Quick Installation / Run

git clone https://github.com/jonclegg/pacman-bakeoff.git

๐Ÿ’ก What Problem Does Pac-Bench Solve?

Pac-Bench is an interactive evaluation framework designed to test frontier models on complete game development and spatial reasoning tasks. It challenges LLMs to one-shot build a fully working Pac-Man game, providing a standardized benchmark for complex code generation.

Commercial AlternativeSWE-bench
Self-HostableYes (Docker/Bare-metal)
Sign-up BarrierNo (Instant Access)
License ModelOpen Source
Discovery Sourcehackernews

โš–๏ธ Pros & Cons Analysis

๐ŸŸข Key Advantages
  • โœ“Tests holistic application design rather than isolated algorithmic snippets
  • โœ“Visually engaging and immediately intuitive evaluation metric (playable game)
  • โœ“Bridges the gap between raw text generation and tangible software execution
๐ŸŸก Things to Consider
  • !Narrow domain focus limited specifically to arcade game clones
  • !Relies on deterministic interpretation of code correctness which can be tricky for generative outputs

โšก Core Architecture & Key Capabilities

01One-Shot Game Generation

Challenges models to generate an entire playable Pac-Man game in a single prompt response.

02Spatial Reasoning Assessment

Evaluates how well models understand coordinate geometry, collision detection, and grid mechanics.

03Interactive Evaluation

Provides instant feedback loops to verify if the generated code results in an executable, playable game.

๐ŸŽฏ Practical Applications & High-Value Use Cases

Scenario 01

Benchmarking the spatial reasoning and code synthesis limits of newly released frontier LLMs

Scenario 02

Comparing coding agent capabilities across complex, multi-component application development

Scenario 03

Researching failure modes of LLMs when handling game loop architecture and state management

๐Ÿ”„ Why Choose Pac-Bench Over SWE-bench?

Unlike SWE-bench which focuses on real-world GitHub issues, Pac-Bench specifically targets spatial reasoning and end-to-end creative code generation in a controlled sandbox environment.

๐ŸŽฏ Target Audience & Who is this for?

AI researchers, ML engineers, and developer tool creators evaluating frontier code generation models.

Top Related Alternatives in AI & Machine Learning

Compare other trending developer tools and open-source projects in this space.