Open eval leaderboard + CI gate for autonomous coding agents (solve, score, trace).
An open, always-on leaderboard and CI gate for autonomous coding agents — every patch runs in a sandbox, every run has a public trace, every regression fails the build. ▶ Live leaderboard: forgejudge.ahmedhobeishy.tech · playground · methodology · model swap · MCP registry Current numbers (hidden-test = the agent never sees the failing test; $0 free tier; same harness, swap the model; 18 tasks ×…
Inferred from the transports this listing declares (stdio). A client not listed here hasn’t been ruled out — it just isn’t something Forge can confirm.
Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.
Forge read 2 source files from the published package tarball and matched no MCP tool registrations. Extraction is pattern-based over shipped source: a server that builds its tool list at runtime, or that ships only bundled or minified code, registers nothing this can see. Treat it as “not detected”, not as “exposes none”.
An open, always-on leaderboard and CI gate for autonomous coding agents — every patch runs in a sandbox, every run has a public trace, every regression fails the build. ▶ Live leaderboard: forgejudge.ahmedhobeishy.tech · playground · methodology · model swap · MCP registry Current numbers (hidden-test = the agent never sees the failing test; $0 free tier; same harness, swap the model; 18 tasks × 3 seeds = 54 runs/model, 162 total): | Model | pass@1 | pass@3 | The score rises with the better…
Forge's dependency resolver reads npm metadata only, so this PyPI package has no resolved tree. That is a gap in coverage, not a clean bill of health.