Universal web content extraction — any URL to LLM-ready markdown. HTML, YouTube, PDF, DOCX.
Universal web content extraction — any URL to LLM-ready markdown. HTML — BeautifulSoup + content density filtering (removes nav, sidebar, ads) YouTube — transcript extraction with timestamps PDF — text extraction with page structure DOCX — paragraph and heading extraction Auto-fallback — tries lightweight httpx first, falls back to Playwright for JS-heavy pages Async-first — built on httpx and…
Inferred from the transports this listing declares (stdio). A client not listed here hasn’t been ruled out — it just isn’t something Forge can confirm.
Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.
Forge read 0 source files from the published package tarball and matched no MCP tool registrations. Extraction is pattern-based over shipped source: a server that builds its tool list at runtime, or that ships only bundled or minified code, registers nothing this can see. Treat it as “not detected”, not as “exposes none”.
Universal web content extraction — any URL to LLM-ready markdown. HTML — BeautifulSoup + content density filtering (removes nav, sidebar, ads) YouTube — transcript extraction with timestamps PDF — text extraction with page structure DOCX — paragraph and heading extraction Auto-fallback — tries lightweight httpx first, falls back to Playwright for JS-heavy pages Async-first — built on httpx and Playwright async APIs Optional extras for specific content types: For HTML pages, if the initial httpx…
Forge's dependency resolver reads npm metadata only, so this PyPI package has no resolved tree. That is a gap in coverage, not a clean bill of health.