Schema-driven document extraction with local OCR + LLM. Document in, Structured JSON out.
Document in, Structured JSON out. Locally. With your schema. docpick is a lightweight, schema-driven document extraction pipeline that combines local OCR engines with local LLMs to extract structured JSON from any document — invoices, receipts, bills of lading, tax forms, and more. Zero cloud dependency — runs entirely on your machine (CPU or GPU) Custom schemas — define your own Pydantic models…
Inferred from the transports this listing declares (stdio). A client not listed here hasn’t been ruled out — it just isn’t something Forge can confirm.
Verification confirms publisher identity (repo ownership), not code safety. The security scan covers known CVEs and suspicious install scripts.
Forge read 0 source files from the published package tarball and matched no MCP tool registrations. Extraction is pattern-based over shipped source: a server that builds its tool list at runtime, or that ships only bundled or minified code, registers nothing this can see. Treat it as “not detected”, not as “exposes none”.
Document in, Structured JSON out. Locally. With your schema. docpick is a lightweight, schema-driven document extraction pipeline that combines local OCR engines with local LLMs to extract structured JSON from any document — invoices, receipts, bills of lading, tax forms, and more. Zero cloud dependency — runs entirely on your machine (CPU or GPU) Custom schemas — define your own Pydantic models or use 8 built-in document schemas Validation built-in — checkdigit verification, cross-field rules,…
Forge's dependency resolver reads npm metadata only, so this PyPI package has no resolved tree. That is a gap in coverage, not a clean bill of health.