Interactive
Check if Jev suits your use case
Answer four short questions and get a plain-language fit summary, plus a ready-to-paste prompt for your coding tool (Cursor, OpenCode, Claude Code, and similar).
jevbench.dev · v1 preview
Jev is a fast structured probabilistic decision model — pick actions from app state, not chat replies. This site measures Jev, Jev-compatible copycats, and dual-brain (guide LLM + Jev) setups on interactive harnesses, starting with StarCraft II and expanding to product loops. Reproducible runs. Open source. No win claims we don’t have.
Top results across Jev models and gaming harnesses — measured and official-cited.
| # | Agent / model | Use case | Category | Score | Cost in |
|---|
Harness v0.1 · 38 SC2 runs · 1 game (StarCraft II) · updated 20 Sep 2026
Where a fast in-loop decision beats a chat reply — decision models on harnesses, games first, then product workflows. Numbers on individual cards are illustrative, not certified JevBench scores.
Interactive
Answer four short questions and get a plain-language fit summary, plus a ready-to-paste prompt for your coding tool (Cursor, OpenCode, Claude Code, and similar).
Sources: TypeSafe blog, jevai.dev, practical guide, and the individual links on each card above.
Run the harness yourself. Fork clean — never commit keys or tokens.
Clone the StarCraft II agent harness and follow the repo README for local setup.
git clone https://github.com/rapidstartup/jev-plays-starcraft-2
cd jev-plays-starcraft-2
# see README for deps, SC2, and eval commands
No releases published yet. Packaged builds, sample replays, and scored run artifacts will land here.
.env in gitWhat JevBench measures and where we’re going.
JevBench is a decision-model bench: we measure Jev and open alternatives on interactive harnesses — win/loss, task completion, latency, and (soon) video evidence — not a single typed answer. Games are the first harness surface; product loops come next.
Interactive games are a harsh, observable testbed: partial observability, long horizons, and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks where agents must act under time pressure. See the full list of use cases (including Doom, Wikiracing, browser use, Mario/StarCraft/drone).
Fortnite (Creative-only adapter plan) and Minecraft (Mineflayer / Mindcraft-class structured adapter plan) are on the roadmap as PARKED — planning docs only; no live scored runs, videos, or leaderboard rows yet. Live game count remains StarCraft II until a harness smoke ships.
The same patterns will expand to non-game workflows (support routing, retrieval judging, trust & safety). Leaderboard filters already cover every use case plus a game / non-game category.
Scored runs will ship with video / replay evidence stored on Cloudflare R2. Clips will link from each leaderboard row once live scoring is online.
Questions about the harness, scoring, or submitting a run.
Use Get help below, or the chat widget once it loads.
Get helpKeep the bench open. Orgs can gift tokens — details coming.
Support JevBench so the public runs stay open. PayPal and Buy Me a Coffee are live; crypto and org token-gift flows still coming.
Organizations can gift API / compute tokens to power public runs. Get in touch to discuss compute sponsorship. Crypto link coming.