Opper Releases Jevman, A Pac-Man Benchmark For Decision Models
By scoring decisions per junction under a hard time limit rather than grading generated text, Opper's benchmark gives probabilistic decision models a shared, measurable task where response speed is recorded as its own result alongside accuracy.
Reporting from 1 source: GIGAZINE.
Opper published jevman, an open-source benchmark that has AI models play Pac-Man to compare decision-making. Six models, including Jev, Kev, Laya, Clef, and GPT-6 Luna Decisions, each play 100 games against standard-rule ghosts. Models receive maze, dot, and ghost data as JSON and return up, down, left, or right as probabilities within a 2-second limit per junction. Ranking uses the 100-game average score, with a 95 percent confidence margin, so models inside that margin are treated as tied.
Jev is a probability model from TypeSafe AI that outputs a likelihood for each preset option instead of generating prose, which makes it suited to decisions that need to be fast. Jevman turns that property into a game: at every junction, the model gets the maze, dot placement, and ghost positions as JSON and answers with probabilities for up, down, left, or right. The game moves according to those probabilities, and any junction where the model fails to answer within 2 seconds is handled by a simple alternative algorithm. That count is recorded too, so speed is scored, not assumed.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.