When I first read the requirements, it became clear that scalability was paramount. Accordingly, I implemented Sentinel-AI as a proof of concept designed to run on Kubernetes and scale seamlessly to millions of users. With the right production-level enhancements—such as optimized provisioning, autoscaling policies, and resilient networking—this prototype can be deployed in a very short timeframe and handle heavy loads at production scale.
Although this initial version is not fully agentic, it already leverages embeddings and a large language model (LLM) and can be easily connected to a private inference server (currently supporting SaaS providers like OpenAI and Anthropic). For an overview of how to integrate a high-throughput, private inference cluster at scale,
see the CogniX architecture.
At the moment, the app is not agentic, but converting the AI-powered components into fully agentic systems is straightforward—see the pseudocode in src/agentic for details.
The base idea is that, if needed, every microservice can be easily ocnverted into an agentic service and the pseudo code in src/agentic is a good starting point.
Sentinel-AI is an event-driven, microservice platform for real-time feed ingestion, filtering, ranking, and anomaly detection—designed to run on Kubernetes and scale to millions of users. 🚀🐳
This project implements a scalable, real-time newsfeed platform that aggregates, filters, stores, and ranks IT-related event data from multiple sources. It is built with an asynchronous, microservice-based architecture, where each service has a distinct responsibility and communicates via a message bus.
💡 Tip: If you are interested in a full agentic AI solution, please have a look at Reforge-AI
Key Features:
- 🔗 **Dynamic Ingestion:** Subscribe to any data feed (RSS, APIs, webhooks, etc.) and ingest events in real time.
- 🧹 **Smart Filtering:** Apply custom relevance rules or plug in ML models to filter events.
- ⚖️ **Deterministic Ranking:** Balance importance & recency with a configurable scoring algorithm; support on-the-fly reordering via APIs.
- 🔍 **Searchable Storage:** Persist full event metadata, embeddings, and scores in a vector database for fast semantic search.
- 🚨 Anomaly Detection: Automatically detect and flag unusual or malformed events.
- 📈 **High Scalability:** Built on NATS JetStream and Kubernetes auto-scaling to serve millions of users with minimal latency.
- 🖥️ **Interactive Dashboard:** List, filter, rerank, delete events and sources, and visualize feeds in real time through a web UI.
- Overview: A high-level summary of the platform's requirements and how they map to the different services.
- Architecture: A detailed look at the microservices architecture, data flows, and technologies used.
- API Service: Describes the main entry point for the system, responsible for ingestion, source management, and data retrieval.
- Scheduler Service: Describes how the platform manages and schedules data collection from sources.
- Connector Service: Explains how the platform fetches and normalizes data from external sources.
- Filter Service: Details the intelligent filtering and enrichment process using LLMs.
- Ranker Service: Explains the configurable ranking algorithm that scores events based on importance and recency.
- Inspector Service: Describes the service responsible for detecting and flagging anomalous or fake news events.
- Guardian Service: Outlines the role of the system's monitoring and health-checking component.
- Web Service: Describes the interactive web UI for managing and visualizing platform data.
This repository ships Docker-Compose manifests (plus helper scripts) that spin up the full micro-service stack in seconds. You can also run an individual service on your laptop for development.
- Docker Desktop or
docker≥ 20.10 anddocker-composeV2. - Linux / macOS (the scripts use
bash). - An API key for your preferred LLM provider (e.g. OpenAI, Anthropic).
deployment/.env.example # single file used by docker-compose
src/*/.env.example # one per micro-service (only needed if you run them separately)
• Docker-Compose path (recommended):
Rename deployment/.env.example → deployment/.env and add your LLM provider key(s).
All other variables already have sensible defaults.
• Local-service path:
When hacking on a service outside Docker, copy its .env.example to .env and tweak as needed.
From one directory above the repo root (so the docker build context remains small):
chmod +x deployment/*.sh # first time only
sudo deployment/start.sh # builds & launches the stackThe very first run downloads base images and builds all containers, so it can take several minutes. Subsequent starts are faster.
sudo deployment/stop.shHaving issues? Check container logs in Portainer or run
docker compose logs -f <service>.
After the stack is up, the script prints a Portainer URL (e.g. http://localhost:9000). On first visit you must create an admin user & password. From the dashboard you can:
- Inspect container logs.
- Restart a service if it failed to start (this PoC occasionally needs manual restarts).
Dashboards exposed by the compose file:
| Service | URL |
|---|---|
| Portainer | http://localhost:9000 |
| Qdrant | http://localhost:6333 |
| NATS | http://localhost:8222 |
| Postgres | http://localhost:5432¹ |
| ¹ psql/GUI only—no web UI included. |
💡 Tip: After adding a few sources and ingesting news, open the Qdrant dashboard to explore the stored vectors.
cd src/ranker
cp .env.example .env # edit variables if needed
pip install -r requirements.txt
python main.pyMake sure Docker Compose is already running NATS, Postgres, and Qdrant (or point the env vars to your own instances).
-
Qdrant Update Race Condition:
- Both the
rankerandinspectorservices use aretrieve-then-updatepattern to modify event records in Qdrant. This can create a race condition where concurrent updates might overwrite each other, leading to data loss. - Recommendation: Modify the
QdrantLogicclass to support theset_payloadoperation, which allows for atomic, partial updates to a record without overwriting the entire object.
- Both the
-
Externalize NATS Retry Policy:
- The message redelivery attempt count (
max_deliver) is currently hardcoded to3in the subscriber services (filter,ranker,inspector). - Recommendation: This should be moved to a
.envvariable (e.g.,NATS_MAX_DELIVER_COUNT) to allow for easier configuration without code changes.
- The message redelivery attempt count (
-
Service Startup Dependencies:
- When the cluster starts, the Web UI may become available before all backend services are ready, leading to initial errors.
- Recommendation: Implement health checks or dependencies in the Docker Compose configuration to ensure a graceful startup sequence.
-
Readiness Probes:
- The readiness probes for the
inspectorandwebservices are not fully functional and need to be corrected.
- The readiness probes for the
- Scheduler Scalability: Replace the current
APSchedulerimplementation with a more distributed and scalable solution suitable for a multi-node environment. - Authentication: Implement
Authentikto add user access control and integrate with existing organizational credentials. - Helm Chart: Create a Helm chart for streamlined deployment to a Kubernetes cluster.
- Improved Web UI: Enhance the user interface with more advanced features and a more polished design.
Powered by ❤️ for intelligent, AI-driven insights!