Skip to content
LalitSatvikPublic

About

Offline-first document intelligence — multi-format ingest, unified schema proposal, extraction with provenance

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

doc-unify

Feed it messy, heterogeneous documents (investor reports, PDFs, images, docx/pptx — video coming later) and it figures out what structured data can be pulled out of them, proposes one unified schema across the whole corpus (catching when "Revenue" and "Net Sales" are the same metric, and flagging when they only look the same), lets you approve/adjust that schema conversationally, then extracts every document against it with full provenance (source doc/page/snippet per cell) and confidence scoring.

Built entirely with free, local tools — Ollama for the LLM and embeddings, Postgres+pgvector for storage/retrieval — so it runs fully offline via Docker Compose. CloudProvider (an OpenAI-compatible client, e.g. for Groq's free tier) implements the same LLMProvider interface as OllamaProvider, so a hosted deployment is a config change (LLM_PROVIDER=cloud + CLOUD_* env vars), not a code change.

Why

Manually unifying data across hundreds of investor reports — each with a different layout, different field names for the same metric, different units, sometimes different accounting methodology — is slow and error-prone. This project automates the two hard parts: figuring out what's extractable across an unfamiliar corpus, and reconciling differently-labeled fields into one consistent table without silently merging things that only look alike.

Screenshots

Real output from a local run — two synthetic investor updates for the same fictional company, one labeling revenue "Total Revenue" and the next "Net Sales", with a GAAP vs. non-GAAP net income pair thrown in.

Intake — drop reports in, watch them get ingested, chunked, and embedded:

Intake

Schema — candidate fields clustered across both documents into one proposal each; note "Revenue" lists both Total Revenue and Net Sales as its as-written members, and "Net Income" lists both the GAAP and non-GAAP figures it clustered together for review:

Schema

Ledger — the unified table, one row per document, every cell stamped with a confidence score that opens its exact source citation on click, plus a review queue for anything that couldn't be mechanically normalized (here, a unit: 'employees' on the headcount figure):

Ledger

Chat — the same pipeline, driven conversationally through the tool-calling agent (shown mid-request; a 3B local model takes a couple of minutes to work through a multi-round tool call on CPU/Metal). The real reply to this exact message, from the same local run:

You have two documents ingested: "Acme_Q3_2025_Investor_Update.pdf" and "Acme_Q4_2025_Investor_Update.pdf". The proposed schema includes the following fields: 'Revenue', 'Gross Margin', 'Total Employee Count', 'Net Income', and 'Adjusted EBITDA'. We can proceed to extract the 'Revenue' field from these documents. Would you like to proceed with extracting the 'Revenue' field from the documents? If so, I can run the extraction process.

Chat

(Run locally against qwen2.5:3b for speed rather than the default qwen2.5:7b. Clustering is embedding-driven and model-independent; the LLM conflict-flagging step and the chat latency both improve with the larger default model or a GPU.)

Status

The full pipeline is implemented end to end: ingest (PDF/image/docx/pptx, with OCR fallback for scans) → chunk + embed into pgvector → propose a unified schema (candidate extraction → clustering → LLM cluster review, with explicit conflict flags) → approve/rename fields → extract + normalize against the approved schema (unit/scale conversion, confidence scoring, a review queue for anything that can't be mechanically normalized) → query it all conversationally through a tool-calling chat agent, or through the ledger-themed frontend. Video ingestion remains an explicitly deferred stub (app/ingestion/video.py) against the same Extractor interface.

Backend: 74 tests, ruff clean. Frontend: next build (TypeScript + ESLint) clean.

Architecture

backend/app/
  ingestion/   per-format extractors -> common ContentBlock interface
  embedding/   chunking + pgvector retrieval
  schema/      candidate-field extraction, clustering, LLM cluster review
  extraction/  structured extraction against an approved schema + normalization
  agent/       tool-calling chat agent over the above
  llm/         LLMProvider interface (OllamaProvider / CloudProvider)
  db/          models + migrations
  api/         FastAPI routes
frontend/      Next.js app (upload, chat, schema review, unified data grid)

Running locally

docker compose up --build

# first run only -- pull the models Ollama needs:
docker compose exec ollama ollama pull qwen2.5:7b
docker compose exec ollama ollama pull nomic-embed-text

Running the backend tests

cd backend
pip install -e ".[dev]"
pytest
ruff check .

License

MIT

About

Offline-first document intelligence — multi-format ingest, unified schema proposal, extraction with provenance

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages