An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
Inferred from the transports this listing declares (streamable-http). A client not listed here hasn’t been ruled out — it just isn’t something Forge can confirm.
Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.
Read out of the source npm actually ships, at scan time. The package was never executed. Tools registered dynamically at runtime, or hidden inside bundled or minified code, can be missed — so this is a floor on the tool surface, not a complete census of it.
get_leaderboardNonobench models ranked by accuracy, overall or for one grid size. Effort defaults to all levels; use effort=best to match the homepage.Nonobench models ranked by accuracy, overall or for one grid size. Effort defaults to all levels; use effort=best to match the homepage.
No input schema was published for this tool.
list_providersProvider ids, names, families and variant counts.Provider ids, names, families and variant counts.
No input schema was published for this tool.
list_familiesModel families, efforts and best variants.Model families, efforts and best variants.
No input schema was published for this tool.
compare_modelsCompare two or more model ids or family names side by side.Compare two or more model ids or family names side by side.
No input schema was published for this tool.
set_filtersUpdate the visible leaderboard filters and URL.Update the visible leaderboard filters and URL.
No input schema was published for this tool.
get_model_resultsAccuracy, cost and latency for one model, per grid size.Accuracy, cost and latency for one model, per grid size.
No input schema was published for this tool.
list_puzzlesThe benchmark puzzles with ids and row/column clues.The benchmark puzzles with ids and row/column clues.
No input schema was published for this tool.
get_puzzle_resultsPer-model outcomes for a puzzle, optionally including parsed grids.Per-model outcomes for a puzzle, optionally including parsed grids.
No input schema was published for this tool.
get_model_puzzlesWhich puzzles a model solved, missed, or did not run.Which puzzles a model solved, missed, or did not run.
No input schema was published for this tool.
check_solutionCheck a nonogram grid (row-major 0/1 string) against a puzzle's clues.Check a nonogram grid (row-major 0/1 string) against a puzzle's clues.
No input schema was published for this tool.
open_puzzleShow a puzzle in the puzzle explorer on this page.Show a puzzle in the puzzle explorer on this page.
No input schema was published for this tool.
nonobenchNo description publishedThis tool published no description. Forge does not invent one.
get_puzzleOne puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.
No input schema was published for this tool.
list_runsIndividual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.
No input schema was published for this tool.
13 of 14 tools published a description.
Tool names and descriptions are written by the publisher and shown verbatim as inert text. They are the strings an MCP client passes to a model, so Forge scans them for prompt-injection patterns — any finding appears with the security scan above. “Privileged” is a keyword match on the tool name, not an audit of what the tool does: a benign-sounding name can still do anything.
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
Linked names open Forge’s index of every entry observed exposing that tool. Browse all indexed tools.
This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.