A multi-threaded C++ read-only cache server for database workloads
Anemo DB is a learning-focused systems project that places a read-only TCP cache layer in front of a database. It is designed to reduce repeated read-query load, absorb traffic bursts, and expose real-time telemetry for monitoring. The current implementation uses PostgreSQL connectors, but the same architecture can be adapted to other databases with minor connector-level changes.
- Multi-threaded TCP server for SQL-over-socket requests
- Read-through cache with
O(1)lookup + LRU ordering - TTL-based lazy expiration (no cleanup background thread)
- Request coalescing (single DB fetch for concurrent misses on same key)
- Bounded request queue with fast rejection when overloaded
- Fixed-size PostgreSQL connection pool (
libpqxx) - Telemetry endpoints via in-band commands (
STATS,STATS_JSON) - Flask + Chart.js dashboard for live status and load simulation
- Benchmark and dataset tooling for cache-vs-DB comparisons
/AnemoDB/Cache Components- Core C++ engine:
main.cpp,CacheEngine.hpp,Cache.hpp,ConnectionPool.hpp,ThreadSafeQueue.hpp
- Core C++ engine:
/AnemoDB/Cache Benchmark- SQL schema/data scripts + Python benchmark clients
/AnemoDB/Cache Monitor- Terminal monitor that polls server stats
/AnemoDB/Web Dashboard- Flask backend, traffic generator, HTML/CSS/JS dashboard UI
/AnemoDB/Bash Control- Helper shell scripts to start/stop/check PostgreSQL and run server
- Client sends a SQL query over TCP, terminated with
<EOQ>. - Listener thread accepts and enqueues requests.
- Worker thread processes request:
STATS/STATS_JSON=> telemetry response- SQL query => cache lookup/reservation
- On miss, leader thread fetches from the database using pooled connection.
- Cache line is fulfilled and waiting followers are notified.
- Response is returned with trailing
<EOQ>delimiter.
For deeper internals, see /AnemoDB/DOCUMENTATION.md.
- Request format:
<SQL_QUERY>\n<EOQ>\n
- Response format:
<PAYLOAD>\n<EOQ>\n
- Special commands:
STATSSTATS_JSON
- Linux environment (scripts use
systemctl+sudo) - PostgreSQL running locally or reachable over network
- C++17 compiler (
g++) libpqxxandlibpq- Python 3
- Python packages:
flask,psycopg2
psql -U postgres -d college_db -f "Cache Benchmark/create_db/01_schema.sql"
psql -U postgres -d college_db -f "Cache Benchmark/create_db/02_generate_data.sql"
psql -U postgres -d college_db -f "Cache Benchmark/create_db/03_indexes.sql"g++ -std=c++17 "Cache Components/main.cpp" -o anemo_db -lpqxx -lpq -pthreadOr run helper script:
bash "Bash Control/anemo_db.sh"./anemo_dbThe admin console prompts for DB and server configuration, then accepts commands:
showstatsclearhelpstop
cd "Web Dashboard"
python web_dashboard.pyOpen http://127.0.0.1:5000.
python "Cache Monitor/monitor_cache.py"
python "Cache Benchmark/client_test.py"
python "Cache Benchmark/benchmark_script.py"
python "Cache Benchmark/benchmark_script2.py"
python "Cache Benchmark/benchmark_script3.py"- Live server connectivity status
- Hit rate, throughput, average latency, queue depth, memory, threads
- Throughput and hit/miss charts
- Live cache vs direct-DB latency comparison
- Interactive traffic controls:
- target mode (
cache,db,both) - dynamic thread count
- start/stop load generation
- target mode (
- SQL scripts create a college-style schema with:
- 6 departments
- 120 courses
- 300 faculty rows
- 1,000,000 students
- 5,000,000 enrollments
- 5,000,000 marks
- Benchmarks include mixed query workloads (point lookups, joins, aggregations).
- Cache server query path is read-focused and intended for benchmark-style read workloads.
- Query result serialization currently returns the first row’s columns as a pipe-separated string.
- Security hardening (auth/TLS/input policy) is not implemented.
- Shell helper scripts are Linux/systemd specific.
No explicit license file is currently present in this repository.

