# @octocrawl/mcp

Web scraper for agents: blocked, empty and wrong pages reported as such, with an Evidence Record.

- **Type:** MCP server
- **Trust:** 65/100 (B), scored on the package rubric
- **Verification:** verified (build provenance)
- **Version:** 0.4.1
- **Author:** io.github.77777R7
- **License:** AGPL-3.0-only
- **npm:** @octocrawl/mcp
- **Endpoints:** streamable-http https://mcp.octocrawl.dev/mcp
- **Source:** https://github.com/77777R7/Octocrawl
- **Endpoint health:** reachable (last checked 2026-10-10T00:47:51.125Z, 1 sample) — uptime is not a security property and is not part of the trust score
- **Compatible clients:** claude-code, cursor, copilot, chatgpt, gemini (basis: transport)

## Trust

65/100 (B), scored on the package rubric
- Publisher verified: no
- Build provenance: verified attestation
- npm trusted publishing (OIDC): yes
- Install scripts: suspicious script found
- Prompt-injection scan: not run
- Obfuscation scan: not run
- Evidence age: 1 day

## Security scan

- **Status:** warnings
- **Scanned:** 2026-10-10T13:36:32.405Z
- **Version scanned:** 0.4.1
- **CVEs:** none found by OSV at scan time
**Findings**
- injection-shaped content (warning) in the `create_delivery_destination` tool: Exfiltration-shaped instruction

## Tools

32 declared. Statically extracted from the shipped source — a floor on the surface, not a census.
- `preview_monitor` — Capture a nonpersistent sample and assess identity, fields, evidence, and missing reasons. Start with preset firecrawl-introduction.
- `create_monitor` — Create a public-document Monitor. Defaults to paused so a delivery destination can be configured first. Use preset firecrawl-introduction for first use.
- `list_monitors` — List current Monitor state and freshness.
- `get_monitor` — Check a Monitor baseline, latest run, latest event, and next schedule.
- `run_monitor` — Queue a durable manual run. Returns runId immediately; disconnection does not cancel execution.
- `get_monitor_run` — Inspect a run and field assessment with evidence and failure reasons.
- `pause_monitor` — Pause scheduling and cancel active Monitor execution.
- `resume_monitor` — Resume Monitor scheduling; first run becomes due immediately.
- `cancel_monitor_run` — Explicitly cancel a queued or running Monitor run.
- `create_delivery_destination` — Register an HTTPS webhook for a Monitor. The secretEnv names an operator environment variable; never send the secret value.
- `list_delivery_destinations` — List webhook destinations, optionally for one Monitor (monitorId) or one crawl or batch (jobId, the taskId); custom header names are listed, never their values.
- `list_deliveries` — Page through delivery state and failures, for a Monitor (monitorId) or a crawl or batch (jobId, the taskId). Defaults to 20 compact results.
- `get_delivery` — Inspect one delivery and its retry attempts.
- `retry_dead_letter` — Explicitly retry a dead-letter delivery with the same eventId.
- `scrape_product` — Get evidence-backed JSON for one anonymous Amazon.sg /dp/{ASIN} product. No schema or model setup needed.
- `batch_products` — Queue 1-1000 distinct Amazon.sg product URLs with the reviewed JSON schema. Returns taskId; page results with get_batch_items.
- `scrape` — Fetch one URL through the Octocrawl coverage ladder. Compact by default; set debug=true for the full audit. The result's warnings name what its content cannot v
- `get_scrape` — Read the record of one scrape call by the scrapeId its response carried (metadata.scrapeId): the request (header values replaced by their names), who made it (o
- `map`
- `crawl` — Start a multi-page crawl. Returns { taskId } (HTTP 202 equivalent). By default it follows links in the start URL's path subtree on its host and www twin, folds 
- `get_crawl` — Read a crawl by task id. Returns a CrawlReport.
- `get_crawl_pages` — Read a paginated list of crawl page results by task id (the latest attempt's unless attemptId is given). Pages omit the routing audit and trace unless debug is 
- `get_crawl_errors` — Read a paginated list of crawl errors by task id.
- `cancel_crawl` — Cancel a crawl task. Completed pages remain queryable.
- `resume_crawl` — Restart a paused or failed crawl with the options it was started with. Returns { taskId }; poll get_crawl.
- `list_active_crawls` — List the crawls the API process is running (those it started and those it resumed at startup; never a batch): each with its id, start URL, status, pages so far 
- `batch_scrape` — Persist and run 1-1000 explicit URLs. Returns a taskId (with ignoreInvalidURLs also invalidURLs, the entries skipped); use get_batch_items for paginated results
- `get_batch_errors` — The items of a batch that did not succeed, across every attempt (a resumed batch keeps its earlier failures): errors [{ id, timestamp, url, status, code, error,
- `hand_off_batch`
- `import_login`
- `list_logins` — The person's saved logins (import_login, octocrawl login import): { logins: [{ domain, savedAt, cookieCount, localStorage, sessionSha256 }] }, never a cookie or
- `remove_login` — Forget the person's saved login to a site (a domain or a page URL on it).

## Install

**Verdict: do-not-install** — Do not install: 1 injection-shaped pattern found in this entry's own text — it may try to steer the model that loads it.
**Blocking**
- 1 injection-shaped pattern found in this entry's own text — it may try to steer the model that loads it. — tool:create_delivery_destination: Exfiltration-shaped instruction
**Client configuration withheld.** Client configs are withheld because this entry has a blocking finding. Show the warnings below to the person installing it.
If they have seen the findings and still want to proceed, request the plan again with acknowledge_warnings=true.

## Blast radius

Extensive blast radius — deletes data; runs locally and hosted.
- Floor 56, ceiling 56 (tier: extensive)
- This is impact, not likelihood. A high radius is not a defect: a filesystem server is supposed to write files. It is never part of the trust score.

## Machine-readable views of this entry

- Signed JSON: https://forgeregistry.com/api/v1/packages/%40octocrawl%2Fmcp
- Install plan: https://forgeregistry.com/api/v1/packages/%40octocrawl%2Fmcp/install-plan
- Alternatives: https://forgeregistry.com/api/v1/alternatives/%40octocrawl%2Fmcp
- HTML page: https://forgeregistry.com/registry/%40octocrawl%2Fmcp
- MCP: POST https://forgeregistry.com/api/mcp → `forge_get_package` / `forge_install_plan`

## About this document

Generated by Forge (https://forgeregistry.com) — a compact rendering of the same record served, signed, at the JSON URL above. Trust and scan facts are the registry's own measurements; anything Forge did not measure is named as unmeasured rather than omitted.
