Document intelligence API powered by Gemini and Vertex AI RAG Engine.
Upload documents, index them into a Vertex RAG corpus, retrieve relevant chunks, and answer questions with grounded citations.
- Vertex AI RAG corpus management (auto-create or reuse existing corpus)
- File ingestion to Vertex RAG (
upload_file) with chunking controls - Semantic retrieval via
rag.retrieval_query - Vertex-native LLM reranking (
RagRetrievalConfig.ranking.llm_ranker) - Grounded answer generation with Gemini Flash and optional Pro fallback
- Auto-ingestion from
docs/with file watcher - Optional API key auth and local/GCS storage
- Python 3.12–3.13
- GCP project with Vertex AI API enabled
- Application Default Credentials:
gcloud auth application-default logingit clone https://github.com/EhsanulHaqueSiam/docmind.git && cd docmind
uv sync
cp .env.example .envSet at least:
GCP_PROJECT_ID=your-project-id
GCP_LOCATION=us-central1Run:
uv run uvicorn src.main:app --reload --host 0.0.0.0 --port 8000Open:
http://localhost:8000(documentation page)http://localhost:8000/docs(Swagger UI)
| Method | Endpoint | Description |
|---|---|---|
POST |
/api/v1/query |
Ask a question with RAG grounding |
POST |
/api/v1/search |
Retrieve relevant chunks (with optional rerank) |
POST |
/api/v1/documents/upload |
Upload + ingest a document |
GET |
/api/v1/documents |
List ingested documents |
DELETE |
/api/v1/documents/{doc_id} |
Delete an ingested document |
POST |
/api/v1/documents/ingest |
Re-ingest all files from docs/ |
GET |
/api/v1/health |
Health check (Vertex RAG connectivity) |
GET |
/api/v1/stats |
Vertex RAG corpus/file stats |
| Variable | Default | Description |
|---|---|---|
GCP_PROJECT_ID |
— | Google Cloud project ID |
GCP_LOCATION |
us-central1 |
Vertex AI region |
GEMINI_FLASH_MODEL |
gemini-2.5-flash |
Default answer model |
GEMINI_PRO_MODEL |
gemini-2.5-pro |
Fallback answer model |
VERTEX_RAG_CORPUS_NAME |
— | Existing corpus resource name/ID to reuse |
VERTEX_RAG_CORPUS_DISPLAY_NAME |
docmind-corpus |
Corpus display name when auto-creating |
VERTEX_RAG_CORPUS_DESCRIPTION |
DocMind managed corpus |
Corpus description |
VERTEX_RAG_EMBEDDING_MODEL |
publishers/google/models/text-embedding-005 |
Embedding model for corpus backend |
VERTEX_RAG_CHUNK_SIZE |
512 |
Chunk size for upload transformations |
VERTEX_RAG_CHUNK_OVERLAP |
100 |
Chunk overlap |
VERTEX_RAG_VECTOR_DISTANCE_THRESHOLD |
— | Optional retrieval filter |
VERTEX_RAG_ENABLE_LLM_RERANK |
true |
Enable Vertex LLM reranking |
VERTEX_RAG_RERANKER_MODEL |
gemini-2.5-flash |
LLM reranker model |
TOP_K |
10 |
Retrieval candidates |
RERANK_TOP_K |
5 |
Final chunks used |
STORAGE_MODE |
local |
local or gcs |
DOCS_DIRECTORY |
docs |
Watched docs directory |
GCS_BUCKET |
— | Required when STORAGE_MODE=gcs |
API_KEY |
— | If set, require X-API-Key |
MAX_UPLOAD_SIZE_MB |
100 |
Upload size limit |
WATCHER_DEBOUNCE_SECONDS |
2.0 |
File watcher debounce |
Upload file
-> save to storage
-> vertexai.rag.upload_file(corpus, chunking config)
Search
-> vertexai.rag.retrieval_query(top_k, optional llm ranker)
Query
-> search_with_rerank()
-> build context with retrieved chunks
-> Gemini generate content (Flash -> optional Pro fallback)
- Build and push image.
- Set env/secrets (project ID, region, bucket, optional API key).
- Deploy
deploy/cloudrun.yaml.
Vertex RAG supported regions and availability can change. Check official Vertex AI RAG docs before production rollout.