Feature/adaptive services - #23
Merged
Merged
Conversation
…ord fallback Ladder step 2 of never-fail: when embedding fails with a non-transient class (monthly quota exhausted / bad credentials) OR zero usable providers remain, pending canonical units complete gen-less and the session finalizes as mode='bm25' — answers keep working at reduced quality instead of 'Indexing failed'. Keying off error CLASS (not breaker state) because a single-provider registry throws after one batch and never accumulates threshold failures. Query-time: vector sessions whose provider died between questions are answered via retrieveBm25Only with degraded:true/qualityTier:'keyword' + plain-language notice; /ai/status exposes qualityTier. Case L reproduces the exact scenario (always-429 provider, vector-size corpus): ready/bm25 → retrieval serves keyword tier → healthy key arrival upgrades to vector on next reindex.
Health monitor service (60s, unref'd): snapshots breaker states + registered generations into a Mongo ProviderState collection so cooldown context survives Render restarts. Tracks consecutiveUnhealthyChecks — after RAG_DISCONTINUE_AFTER_CHECKS (3) the provider is flagged 'discontinued' for routing hints. Real inference requests remain the only availability probe (doc §21); this loop only RECORDS what they taught us. Wired to boot alongside recoverPendingIndexes.
New rag/embedding/selector.ts centralizes the full eligibility predicate (registered ∧ permitted ∧ breaker-closed ∧ monitor-streaks) and owns the ModelRegistry — model-id → providers table so future same-model provider pairs fail over WITHOUT a generation rebuild (doc §34 Strategy A; ready even though today's models are unique). Orchestrator.embed() now orders candidates via the selector each call: healthiest/fastest first instead of static env order. Env RAG_PROVIDER_ORDER remains as the base registry construction.
File Analyzer (rag/analyzer/file-analyzer.ts): classifies extracted content — kind (pdf/docx/xlsx/text/image), scanned-PDF detection (pages-without-text ratio), conservative token estimate, workload tier (low/medium/high/very_high) and an actionable notice when content needs OCR but the host can't run it. File size is never used alone; scoring runs on EXTRACTED reality (review §16). Extraction router: image files (png/jpg/jpeg/webp/bmp) are supported ONLY when RAG_OCR_ENABLED=true AND the memory profile grants ≥300MB workload ceiling — tesseract.js loads lazily per file, input capped by RAG_OCR_MAX_BYTES, RSS projection checked before recognition starts, worker terminated after. TINY hosts keep OCR off: images surface an honest 'enable RAG_OCR_ENABLED' notice instead of silent skips or OOM. Selector integration: pipeline passes each batch's conservative token estimate into embedWithFailover → orchestrator orders providers with real workload context (selector ctx). New dependency: tesseract.js (lazy, opt-in via env). Tests: unit 61/61 (analyzer matrix incl. boundary fix) · integration 12/12.
/health rag block now includes per-provider durable state (id/generationId/state/discontinued) from the health monitor, so operators can see WHY routing degraded straight from the dashboard. .env.example documents all new adaptive-services knobs (OCR tier-gate, monitor interval, discontinue threshold).
The OCR commit (fa3d18b) imported tesseract.js but the package.json / lockfile changes never got committed — local node_modules masked it while CI's npm ci correctly had nothing to install. Dependency is lazy-imported and tier-gated; no runtime cost unless RAG_OCR_ENABLED.
Deploying quick-share with
|
| Latest commit: |
0918789
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://80617208.quick-share-1a4.pages.dev |
| Branch Preview URL: | https://feature-adaptive-services.quick-share-1a4.pages.dev |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.