YachayBot is a source-grounded AI search application for exploring public Peruvian cultural and educational knowledge. It turns a question into an inspectable path: source cards, cited answers, evidence strength, and refusal behavior when the indexed corpus cannot support a confident response.
Built with Next.js, TypeScript, Zod, and optional Mistral answer polishing, the project demonstrates production-minded AI engineering through typed API contracts, deterministic retrieval, citation validation, reproducible tests, and transparent product scope.
Start with the reviewer system proof, then inspect the demo evidence, architecture, migration plan, methodology, limitations, eval fixtures, and API boundary docs.
| Review need | Where to look |
|---|---|
| Fast technical path through the system | docs/system-proof.md |
| Demo routes, static artifacts, and smoke coverage | docs/demo-evidence.md and docs/demo-script.md |
| Architecture and service boundaries | docs/architecture.md and docs/backend-migration.md |
| Source-grounding and evaluation approach | docs/methodology.md, docs/limitations.md, docs/eval-v2-metric-deltas.md, and evals/README.md |
| FastAPI sidecar boundary | api/README.md |
Validation commands:
npm run evidence:demo
npm run validate
cd api
C:\Users\abiga\AppData\Local\Programs\Python\Python313\python.exe -m pytestI built YachayBot around one principle: AI education tools are only useful when people can inspect where an answer came from. The project connects cultural access, public educational resources, and responsible AI by making evidence visible instead of hiding it behind a chatbot response.
YachayBot showcases end-to-end product and engineering scope: retrieval design, API validation, model-output constraints, evaluation metrics, browser-tested flows, and user-facing polish.
-
Source-Grounded AI Search
Search public Peruvian cultural and educational resources through inspectable source cards, snippets, metadata, rights notes, and URLs. -
Citation-Aware Answering
Generate cited answers from retrieved evidence, with refusal behavior when indexed sources are too weak or absent. -
Model Output Safety
Optional Mistral polishing is accepted only when the model preserves known citation markers instead of dropping or inventing citations. -
Evaluation Dashboard
Review top-3 hit rate, top-5 hit rate, refusal pass rate, and latency for the deterministic retrieval baseline. -
Typed Full-Stack Implementation
Next.js route handlers, TypeScript contracts, Zod request validation, unit tests, and browser-smoked public flows.
| Route | Purpose |
|---|---|
/es |
Spanish search experience |
/qu |
Quechua-locale search shell with source-bound behavior |
/ay |
Aymara-locale search shell with source-bound behavior |
/es/sources |
Source browser with language, topic, and source-type filters |
/es/evals |
Retrieval and refusal evaluation dashboard |
/es/ai-bot |
Chat UI over the same retrieval, citation, and refusal rules |
Account and educator workspace routes are intentionally scoped out of the public build. /sign-in and /dashboard show scope pages, while /api/auth/* returns 503 FEATURE_UNAVAILABLE. This keeps the project focused on the AI search system instead of presenting unfinished authentication as a feature.
YachayBot is designed around a simple rule: the answer is not useful unless the evidence is visible.
- A user submits a query from search or chat.
- API routes validate input with Zod.
searchCorpusranks local chunks with deterministic token-overlap scoring.buildAnswerclassifies evidence as strong, moderate, or weak.- Weak evidence returns a refusal instead of a confident answer.
- Usable evidence returns a cited grounded answer.
- If
MISTRAL_API_KEYis configured, Mistral can polish the grounded answer. - Polished output is accepted only if it preserves known citation markers.
User question
|
v
Next.js localized UI
|--------------------------|
v v
/api/v1/search /api/chat
| |
v v
Shared corpus JSON -> Deterministic retrieval
|
v
Evidence strength check
| |
v v
Refusal for weak Cited grounded answer
evidence |
v
Citation marker validation
| |
v v
UI Optional Mistral polishing
Important implementation files:
src/app/[locale]/page.tsx: localized search-first homepagesrc/app/[locale]/sources/page.tsx: source browsersrc/app/[locale]/evals/page.tsx: evaluation dashboardsrc/app/[locale]/ai-bot/page.tsx: chat UIsrc/app/api/v1/*: public API routessrc/app/api/chat/route.ts: evidence-first chat endpointsrc/lib/v2-data.ts: corpus, retrieval, answer generation, and eval logicsrc/lib/v2-schemas.ts: Zod request validation schemassrc/lib/v2-types.ts: shared TypeScript contractsmigrations/001_v2_core.sql: Postgres/pgvector schema design
| Area | Implementation |
|---|---|
| Intended use | Discover and inspect public Peruvian cultural and educational resources |
| Retrieval source | Shared JSON corpus in data/corpus.json |
| Retrieval method | Deterministic token-overlap baseline |
| Generation | Template-based grounded answer, with optional Mistral polishing |
| Guardrails | Zod validation, evidence-strength refusal, citation marker validation, rate limiting |
| Citations | Answers cite retrieved chunks with markers such as [1] |
| Unsupported requests | Refused when indexed evidence is weak or absent |
| Reproducibility | Works locally without provider keys or hosted databases |
| Database design | Postgres/pgvector schema included in migrations/001_v2_core.sql |
- 15 public or official source records
- 5 curated methodology notes
- 1 inspectable chunk per source record
- Metadata for institution, language, region, topic tags, rights notes, and source URL
The Eval v2 set contains 50 questions: factual retrieval checks with expected source IDs, paraphrases, typo/noisy queries, mixed-language prompts, source-confusion cases, refusal checks for unsupported or unsafe requests, ambiguous prompts, multilingual boundary checks, off-topic prompts, and citation behavior checks.
| Metric | Value |
|---|---|
| Top-1 hit rate | 0.86 |
| Top-3 hit rate | 1.00 |
| Top-5 hit rate | 1.00 |
| Refusal pass rate | 1.00 |
| Citation pass rate | 1.00 |
| Citation coverage | 1.00 |
These metrics verify the local corpus, deterministic retrieval baseline, refusal behavior, citation behavior, and latency. Before/after deltas are documented in docs/eval-v2-metric-deltas.md. They are not presented as broad production accuracy claims.
- Frontend: Next.js, React, TypeScript, Tailwind CSS, next-intl
- API: Next.js route handlers
- Validation: Zod request schemas
- Retrieval: Shared JSON corpus with deterministic token-overlap scoring
- AI: Optional Mistral polishing after retrieval and citation checks
- Data Design: Supabase Postgres and pgvector migration schema
- Testing: TypeScript unit tests, Next.js build checks, browser smoke testing
npm install
npm run prisma-generate
npm run dev
Open http://localhost:3000/es.
If port 3000 is busy:
npm run dev -- -p 3001copy .env.example .env.local| Variable | Current role |
|---|---|
MISTRAL_API_KEY |
Enables optional answer polishing after source-grounded retrieval |
YACHAYBOT_SEARCH_SERVICE_URL |
Optional FastAPI sidecar base URL for /api/v1/search; local deterministic retrieval is used as fallback |
YACHAYBOT_PGVECTOR_ENABLED |
Opt-in flag for experimental FastAPI pgvector retrieval comparison |
YACHAYBOT_PGVECTOR_DATABASE_URL |
Optional Postgres connection string for experimental pgvector retrieval comparison |
PINECONE_API_KEY |
Reserved for vector-search migration experiments |
PINECONE_HOST |
Reserved for vector-search migration experiments |
The deterministic search and evaluation flows work without Mistral, Pinecone, or pgvector credentials.
docker build -t yachaybot .
docker run --rm -p 3000:3000 yachaybotThe Docker runtime serves the deterministic MVP without optional provider credentials. Mistral, Pinecone, and hosted database variables are reserved for optional or future paths.
GET /api/v1/documents: returns source records and inspectable chunks.POST /api/v1/search: returns ranked chunks, source metadata, latency, detected language, evidence strength, and an answer preview.POST /api/v1/answers: generates an answer from reviewed chunks only, or refuses weak evidence.GET /api/v1/evals/runs: returns the deterministic eval run and metrics.POST /api/chat: validates chat messages, retrieves from the corpus, builds a cited answer or refusal, and optionally calls Mistral.POST /v1/searchin the FastAPI sidecar: returns deterministic search and answer-preview parity over the shared corpus.GET /v1/evals/runsin the FastAPI sidecar: returns deterministic eval-run parity over the shared eval set.POST /v1/retrieval/comparein the FastAPI sidecar: returns deterministic baseline retrieval plus optional experimental pgvector comparison status/results.
. |-- src/ | |-- app/ # Localized pages and API routes | |-- components/ # UI components and search experience | |-- i18n/ # next-intl routing and request config | `-- lib/ # Corpus, retrieval, schemas, tests, types |-- docs/ # Architecture, methodology, limitations, ADRs |-- public/demo/ # Demo SVGs for review |-- migrations/ # Postgres/pgvector schema |-- data/ # Shared corpus, curated data notes, and source metadata |-- evals/ # Evaluation notes |-- api/ # FastAPI service boundary and sidecar search parity |-- web/ # Frontend boundary notes `-- prisma/ # Prisma setup retained for data/client generation
npm run lint
npm run typecheck
npm test
npm run build
npm run evidence:demo
npm run validate- TypeScript route and data tests
- Invalid API payload tests
- Search and answer behavior tests
- Mistral citation acceptance and rejection tests
- Auth scope route tests
- Eval run API tests
- Chat rate-limiting tests
- Production Next.js build
- GitHub Actions CI on pull requests and pushes to
main
To smoke-test public demo routes against a running app:
npm run build
npm run start
npm run smoke:public
Set SMOKE_BASE_URL to test a non-default host.
YachayBot placed third at INFORTELGRAF Peru Hackathon 2025. The project is aligned with educational access, responsible AI, and public-interest technology.