Skip to content

Repository files navigation

DocMind

Document intelligence API powered by Gemini and Vertex AI RAG Engine.

Upload documents, index them into a Vertex RAG corpus, retrieve relevant chunks, and answer questions with grounded citations.

Features

  • Vertex AI RAG corpus management (auto-create or reuse existing corpus)
  • File ingestion to Vertex RAG (upload_file) with chunking controls
  • Semantic retrieval via rag.retrieval_query
  • Vertex-native LLM reranking (RagRetrievalConfig.ranking.llm_ranker)
  • Grounded answer generation with Gemini Flash and optional Pro fallback
  • Auto-ingestion from docs/ with file watcher
  • Optional API key auth and local/GCS storage

Prerequisites

  • Python 3.12–3.13
  • GCP project with Vertex AI API enabled
  • Application Default Credentials:
gcloud auth application-default login

Quick Start

git clone https://github.com/EhsanulHaqueSiam/docmind.git && cd docmind
uv sync
cp .env.example .env

Set at least:

GCP_PROJECT_ID=your-project-id
GCP_LOCATION=us-central1

Run:

uv run uvicorn src.main:app --reload --host 0.0.0.0 --port 8000

Open:

  • http://localhost:8000 (documentation page)
  • http://localhost:8000/docs (Swagger UI)

API Endpoints

Method Endpoint Description
POST /api/v1/query Ask a question with RAG grounding
POST /api/v1/search Retrieve relevant chunks (with optional rerank)
POST /api/v1/documents/upload Upload + ingest a document
GET /api/v1/documents List ingested documents
DELETE /api/v1/documents/{doc_id} Delete an ingested document
POST /api/v1/documents/ingest Re-ingest all files from docs/
GET /api/v1/health Health check (Vertex RAG connectivity)
GET /api/v1/stats Vertex RAG corpus/file stats

Configuration

Variable Default Description
GCP_PROJECT_ID — Google Cloud project ID
GCP_LOCATION us-central1 Vertex AI region
GEMINI_FLASH_MODEL gemini-2.5-flash Default answer model
GEMINI_PRO_MODEL gemini-2.5-pro Fallback answer model
VERTEX_RAG_CORPUS_NAME — Existing corpus resource name/ID to reuse
VERTEX_RAG_CORPUS_DISPLAY_NAME docmind-corpus Corpus display name when auto-creating
VERTEX_RAG_CORPUS_DESCRIPTION DocMind managed corpus Corpus description
VERTEX_RAG_EMBEDDING_MODEL publishers/google/models/text-embedding-005 Embedding model for corpus backend
VERTEX_RAG_CHUNK_SIZE 512 Chunk size for upload transformations
VERTEX_RAG_CHUNK_OVERLAP 100 Chunk overlap
VERTEX_RAG_VECTOR_DISTANCE_THRESHOLD — Optional retrieval filter
VERTEX_RAG_ENABLE_LLM_RERANK true Enable Vertex LLM reranking
VERTEX_RAG_RERANKER_MODEL gemini-2.5-flash LLM reranker model
TOP_K 10 Retrieval candidates
RERANK_TOP_K 5 Final chunks used
STORAGE_MODE local local or gcs
DOCS_DIRECTORY docs Watched docs directory
GCS_BUCKET — Required when STORAGE_MODE=gcs
API_KEY — If set, require X-API-Key
MAX_UPLOAD_SIZE_MB 100 Upload size limit
WATCHER_DEBOUNCE_SECONDS 2.0 File watcher debounce

Architecture

Upload file
  -> save to storage
  -> vertexai.rag.upload_file(corpus, chunking config)

Search
  -> vertexai.rag.retrieval_query(top_k, optional llm ranker)

Query
  -> search_with_rerank()
  -> build context with retrieved chunks
  -> Gemini generate content (Flash -> optional Pro fallback)

Deploy to Cloud Run

  • Build and push image.
  • Set env/secrets (project ID, region, bucket, optional API key).
  • Deploy deploy/cloudrun.yaml.

Notes

Vertex RAG supported regions and availability can change. Check official Vertex AI RAG docs before production rollout.

About

RAG system with Vertex AI Gemini, Docling, and Qdrant

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages