Benchmark AI models on real prompts. Find cheaper, faster alternatives across 340+ models.
MCP server that benchmarks AI models on your actual prompts and finds cheaper, faster alternatives. Works with Claude Code, Cursor, Windsurf, and any MCP-compatible tool. Sign up at llmtest.io and grab your API key from the dashboard. Cursor / Windsurf / Other MCP clients: Add to your MCP config file: Just ask in natural language: "Check my LLMTest status" "Find cheaper models for my AI calls"…
Inferred from the transports this listing declares (stdio). A client not listed here hasn’t been ruled out — it just isn’t something Forge can confirm.
Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.
Read out of the source npm actually ships, at scan time. The package was never executed. Tools registered dynamically at runtime, or hidden inside bundled or minified code, can be missed — so this is a floor on the tool surface, not a complete census of it.
statusNo description publishedThis tool published no description. Forge does not invent one.
list_flowsNo description publishedThis tool published no description. Forge does not invent one.
get_suggestionsNo description publishedThis tool published no description. Forge does not invent one.
update_suggestionNo description publishedThis tool published no description. Forge does not invent one.
run_benchmarkNo description publishedThis tool published no description. Forge does not invent one.
list_new_modelsNo description publishedThis tool published no description. Forge does not invent one.
optimize_promptNo description publishedThis tool published no description. Forge does not invent one.
get_autopilot_statusNo description publishedThis tool published no description. Forge does not invent one.
enable_autopilotNo description publishedThis tool published no description. Forge does not invent one.
disable_autopilotNo description publishedThis tool published no description. Forge does not invent one.
list_active_optimizationsNo description publishedThis tool published no description. Forge does not invent one.
revert_optimizationNo description publishedThis tool published no description. Forge does not invent one.
get_accountNo description publishedThis tool published no description. Forge does not invent one.
seed_samplesNo description publishedThis tool published no description. Forge does not invent one.
list_samplesNo description publishedThis tool published no description. Forge does not invent one.
0 of 15 tools published a description.
Tool names and descriptions are written by the publisher and shown verbatim as inert text. They are the strings an MCP client passes to a model, so Forge scans them for prompt-injection patterns — any finding appears with the security scan above. “Privileged” is a keyword match on the tool name, not an audit of what the tool does: a benign-sounding name can still do anything.
MCP server that benchmarks AI models on your actual prompts and finds cheaper, faster alternatives. Works with Claude Code, Cursor, Windsurf, and any MCP-compatible tool. Sign up at llmtest.io and grab your API key from the dashboard. Cursor / Windsurf / Other MCP clients: Add to your MCP config file: Just ask in natural language: "Check my LLMTest status" "Find cheaper models for my AI calls" "Run a benchmark on my blog-writer flow" "What models are trending?" LLMTest is a proxy that sits…
Linked names open Forge’s index of every entry observed exposing that tool. Browse all indexed tools.
This package was last scanned before Forge began storing the resolved tree. The next scan will record it.