io.github.mauricekleine/nonobench

MCPcommunitylive
v1.1.0io.github.mauricekleineUnknownUpdated 5d agoGitHub

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

Endpoint healthlive
checked 2 days ago · 675ms
100% of the last 1 check reached this endpoint
Works in
ClaudeCursorCopilotChatGPTGemini

Inferred from the transports this listing declares (streamable-http). A client not listed here hasn’t been ruled out — it just isn’t something Forge can confirm.

Automatically indexed from public sources. Not yet verified by the developer on Forge.Claim this listing →
6GitHub stars
1Forks
5d agoLast update
Package
Authorio.github.mauricekleine
LicenseUnknown
Version1.1.0
Sourcemcp-registry
Trust Status
B
60/100Good
✓Listed in Forge index+10/10
—Publisher identity verified+0/30
→ Publisher: run `forge publish` from the repo to claim ownership
—Domain verification+0/10
→ Not currently available for this listing type — the domain-verification check only runs for npm-backed packages today, so this row cannot be earned here yet regardless of what's hosted at the domain.
✓Prompt-injection scan · clean+30/30
✓Obfuscation / exfil scan · clean+20/20
Paste into Claude Code, Cursor, or any AI assistant to fix all gaps
StatusCommunity-indexed
PublisherUnverified
SignatureUnsigned
Domain—
Provenance—
DependenciesNot audited
Tool surface14 tools · none privileged
Security scan✓ CleanvHEAD · 2d agoHow well does this scan work?
EvalsNone
IndexedOct 2, 2026

Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.

Tools

14 tools · none privileged
Statically extracted from the published packagevHEAD · 2d ago

Read out of the source npm actually ships, at scan time. The package was never executed. Tools registered dynamically at runtime, or hidden inside bundled or minified code, can be missed — so this is a floor on the tool surface, not a complete census of it.

get_leaderboardNonobench models ranked by accuracy, overall or for one grid size. Effort defaults to all levels; use effort=best to match the homepage.

Nonobench models ranked by accuracy, overall or for one grid size. Effort defaults to all levels; use effort=best to match the homepage.

No input schema was published for this tool.

list_providersProvider ids, names, families and variant counts.

Provider ids, names, families and variant counts.

No input schema was published for this tool.

list_familiesModel families, efforts and best variants.

Model families, efforts and best variants.

No input schema was published for this tool.

compare_modelsCompare two or more model ids or family names side by side.

Compare two or more model ids or family names side by side.

No input schema was published for this tool.

set_filtersUpdate the visible leaderboard filters and URL.

Update the visible leaderboard filters and URL.

No input schema was published for this tool.

get_model_resultsAccuracy, cost and latency for one model, per grid size.

Accuracy, cost and latency for one model, per grid size.

No input schema was published for this tool.

list_puzzlesThe benchmark puzzles with ids and row/column clues.

The benchmark puzzles with ids and row/column clues.

No input schema was published for this tool.

get_puzzle_resultsPer-model outcomes for a puzzle, optionally including parsed grids.

Per-model outcomes for a puzzle, optionally including parsed grids.

No input schema was published for this tool.

get_model_puzzlesWhich puzzles a model solved, missed, or did not run.

Which puzzles a model solved, missed, or did not run.

No input schema was published for this tool.

check_solutionCheck a nonogram grid (row-major 0/1 string) against a puzzle's clues.

Check a nonogram grid (row-major 0/1 string) against a puzzle's clues.

No input schema was published for this tool.

open_puzzleShow a puzzle in the puzzle explorer on this page.

Show a puzzle in the puzzle explorer on this page.

No input schema was published for this tool.

nonobenchNo description published

This tool published no description. Forge does not invent one.

get_puzzleOne puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.

One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.

No input schema was published for this tool.

list_runsIndividual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.

Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.

No input schema was published for this tool.

13 of 14 tools published a description.

Tool names and descriptions are written by the publisher and shown verbatim as inert text. They are the strings an MCP client passes to a model, so Forge scans them for prompt-injection patterns — any finding appears with the security scan above. “Privileged” is a keyword match on the tool name, not an audit of what the tool does: a benign-sounding name can still do anything.

About

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

Keywords
mcp
Alternatives
Comparing tool surfaces…

No dependency coverage

This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.