Skip to content

AI Cost Firewall

Rust License GitHub Release Docker Status

OpenAI-compatible control layer for AI cost, with optional privacy, security, audit, and compliance integrations

AI Cost Firewall is a lightweight OpenAI-compatible gateway that reduces unnecessary LLM API calls through exact and semantic cache reuse.

AI Cost Firewall can be deployed independently as a complete caching and cost-control gateway.

It can also integrate with separately licensed VCAL modules for privacy, security, audit, and compliance.

AI Cost Firewall and VCAL platform overview

AI Cost Firewall controls the request path between AI applications and model providers.

AI Cost Firewall is developed and maintained by VCAL Labs, Inc.


Why AI Cost Firewall?

LLM applications frequently generate repeated or semantically similar prompts.

Without caching, every request results in:

  • repeated upstream API calls
  • additional token usage
  • higher cost
  • avoidable latency

AI Cost Firewall introduces a two-layer cache:

  1. Exact cache (Redis)
  2. Semantic cache (Qdrant)

The firewall behaves similarly to “nginx for LLM APIs”:

  • applications call AI Cost Firewall
  • the firewall evaluates exact and semantic cache reuse
  • only cache misses reach the upstream provider

Supported OpenAI-compatible providers include:

  • OpenAI
  • Ollama
  • LM Studio
  • vLLM
  • LiteLLM
  • OpenRouter

Production Capabilities

AI Cost Firewall includes:

  • exact Redis caching
  • semantic Qdrant caching
  • OpenAI-compatible request routing
  • configurable cache fail-open/fail-closed behavior
  • structured lifecycle evidence
  • Prometheus metrics and Grafana dashboards
  • readiness, liveness, graceful shutdown, and runtime diagnostics

Optional integrations with separately licensed VCAL modules include:

  • VCAL Privacy Guard
  • VCAL Security Guard
  • VCAL Audit
  • VCAL Compliance

See the latest GitHub release for release-specific changes.


Included Dashboards

AI Cost Firewall includes Grafana dashboards for cost visibility, cache effectiveness, runtime diagnostics, and high-level guard orchestration health.

The dashboards are included in the Docker deployment files and are automatically provisioned by Grafana when using the provided Docker Compose setup.

Detailed Privacy Guard and Security Guard findings remain in their product-specific dashboards. VCAL Audit provides its own dashboard for ingestion volume, durable persistence latency, authentication failures, hash-chain verification, and license status.

Cost Savings Overview

AI Cost Firewall Grafana Dashboard

30-minute cold-cache demo run with local simulated OpenAI-compatible upstream.

The Overview dashboard shows the high-level cost and cache impact of AI Cost Firewall.

It demonstrates:

  • total request volume
  • estimated chat-completion cost
  • gross savings from cache reuse
  • embedding overhead
  • net savings after embedding cost
  • net savings percentage
  • cache hit rate
  • exact and semantic cache activity
  • cache bypass request rate
  • per-model spend and savings
  • savings by cache type

This dashboard is intended for quick validation, demos, and cost-savings reviews.


Semantic Diagnostics

AI Cost Firewall Grafana Dashboard

Semantic diagnostics from the same cold-cache demo run, including readiness, threshold behavior, lookup latency, and cache activity.

The Diagnostics dashboard provides a deeper operational view of semantic-cache behavior and runtime health.

It demonstrates:

  • readiness state
  • semantic lookup volume
  • semantic threshold pass/fail behavior
  • semantic candidate evaluation
  • expired semantic entries skipped during lookup
  • semantic lookup latency
  • upstream and embedding latency
  • embedding overhead by operation
  • gross vs net semantic savings
  • exact vs semantic savings
  • semantic cache misses vs threshold passes
  • semantic store health
  • runtime and provider pressure signals
  • provider error classes
  • guard orchestration outcomes by guard, stage, and result
  • guard orchestration latency

This dashboard is intended for troubleshooting, tuning semantic similarity thresholds, validating fail-open behavior, and understanding runtime cache behavior during pilots.


Deployment Patterns

AI Cost Firewall includes ready-to-run deployment examples under:

deploy/examples/

Available patterns:

Pattern Description
openai-cloud/ Fastest cloud evaluation path
local-ollama/ Fully local OpenAI-compatible deployment
hybrid-openai-local-embeddings/ OpenAI chat + local embeddings
openrouter/ OpenRouter upstream with OpenAI embeddings
local-full-stack/ Fully local stack with dashboards

Each example includes:

  • docker-compose.yml
  • minimal configuration
  • example requests
  • expected behavior
  • expected metrics
  • optional observability overlays

Architecture Overview

AI Cost Firewall Architecture Diagram

Client applications send requests to AI Cost Firewall instead of directly to the LLM provider.

The firewall:

  1. validates requests
  2. optionally calls VCAL Security Guard before cache/upstream processing
  3. optionally calls VCAL Privacy Guard to anonymize or redact sensitive text
  4. checks exact cache
  5. checks semantic cache
  6. forwards only cache misses upstream
  7. optionally scans assistant responses before Privacy Guard restore
  8. emits structured evidence events and Prometheus metrics
  9. optionally delivers evidence batches to VCAL Audit
  10. exposes operational diagnostics

Full architecture documentation:

docs/architecture.md

Quick Start (Docker)

Prerequisites

Install:

  • Docker
  • Docker Compose

Verify installation:

docker --version
docker compose version

Clone the repository

git clone https://github.com/vcal-project/ai-firewall.git
cd ai-firewall

Copy the example configuration:

cp configs/ai-firewall.conf.example configs/ai-firewall.conf

Edit the configuration and add your API key:

nano configs/ai-firewall.conf

Start the stack

The default deployment starts:

  • AI Cost Firewall
  • Redis
  • Qdrant
  • Prometheus
  • Grafana
docker compose pull
docker compose up -d

Validate the deployment

curl http://localhost:8080/healthz
curl http://localhost:8080/readyz
curl http://localhost:8080/version

Expected:

OK
ready

The /version endpoint returns release metadata, including the AI Cost Firewall version, release title, and OpenAI-compatible compatibility model.


Streaming behavior

Streaming chat completions are currently not supported.

Requests with "stream": true are rejected with HTTP 422 before cache, guard, or upstream processing.


Operational Features

AI Cost Firewall includes operational safeguards and observability features designed for real deployments.

Runtime Features

  • readiness and liveness endpoints
  • graceful shutdown with request draining
  • startup dependency validation
  • nginx-style configuration reload (SIGHUP)
  • structured Prometheus metrics
  • semantic cache lifecycle control
  • upstream timeout tracking
  • request size protection
  • runtime diagnostics
  • configurable semantic cache fail-open behavior
  • optional Security Guard and Privacy Guard orchestration
  • configurable guard fail-open/fail-closed behavior
  • structured evidence events with trace correlation
  • exactly one terminal request lifecycle event per received trace
  • optional buffered HTTP evidence delivery to VCAL Audit
  • configurable Audit batching, queue capacity, timeout, and retry behavior

Optional VCAL Modules

VCAL Privacy Guard, VCAL Security Guard, VCAL Audit, and VCAL Compliance are separate commercial products. They are not required to deploy or use AI Cost Firewall.

AI Cost Firewall can optionally orchestrate VCAL Security Guard and VCAL Privacy Guard before forwarding non-streaming chat requests upstream.

Supported deployment modes:

AI Firewall only
AI Firewall + VCAL Security Guard
AI Firewall + VCAL Privacy Guard
AI Firewall + VCAL Security Guard + VCAL Privacy Guard

When both enterprise guards are enabled, the recommended request/response flow is:

Client
  -> AI Cost Firewall
      -> VCAL Security Guard request scan
      -> VCAL Privacy Guard anonymize/redact
      -> exact/semantic cache lookup or upstream LLM
      -> VCAL Security Guard response scan
      -> VCAL Privacy Guard restore
      -> Client

Security Guard can block malicious request-side prompts before Privacy Guard, cache, or upstream processing. Privacy Guard can replace sensitive values with placeholders before cache/upstream processing and restore them in the final response.

Example Privacy Guard transformation:

Original request:
Analyze login from 185.23.10.5 by john@example.com

Sent upstream/cache path:
Analyze login from [IP_1] by [EMAIL_1]

Returned response:
john@example.com logged in from 185.23.10.5

Example Security Guard block returned by AI Firewall:

{
  "error": {
    "code": 403,
    "guard": "security",
    "type": "security_request_blocked",
    "stage": "request",
    "rule_id": "VSG-PA-003"
  }
}

Streaming requests are rejected globally in v0.4.2, regardless of whether guard modules are enabled. Requests with stream=true return HTTP 422 before cache, guard, or upstream processing.

Security Guard and Privacy Guard are disabled by default in configs/ai-firewall.conf.example.


Health Endpoints

Endpoint Purpose
/healthz Process liveness
/readyz Ready to serve traffic

Configuration Validation

Validate configuration statically before startup:

docker compose run --rm firewall \
  --config /configs/ai-firewall.conf \
  --test-config

Expected output:

configuration OK

Semantic Cache Fail-Open Behavior

When semantic_cache_fail_open is enabled, runtime semantic cache lookup or embedding failures skip semantic cache and continue to the upstream LLM endpoint.

This setting applies to runtime semantic cache behavior. It does not bypass startup dependency validation when semantic cache is enabled. If semantic cache is enabled, Qdrant must be reachable during startup and the configured vector size must match the collection.


Print Loaded Configuration

docker compose run --rm firewall \
  --config /configs/ai-firewall.conf \
  --print-config

Secrets are automatically masked.


OpenAI-Compatible Providers

AI Cost Firewall supports practical OpenAI-compatible deployments while keeping a simple flat configuration model.

The current model is:

upstream_provider openai_compatible;
embedding_provider openai_compatible;

This means AI Cost Firewall expects OpenAI-style chat and embedding APIs. It does not yet provide provider-specific configuration blocks or native provider-specific request transformations.

Common OpenAI-compatible deployment patterns include:

Runtime or Gateway Usage Pattern
OpenAI Cloud OpenAI-compatible chat and embeddings
Ollama Local OpenAI-compatible model endpoint
LM Studio Local OpenAI-compatible model endpoint
vLLM Self-hosted OpenAI-compatible serving
LiteLLM Gateway in front of multiple providers
OpenRouter OpenAI-compatible hosted gateway

Example configuration:

upstream_provider openai_compatible;
upstream_base_url https://api.openai.com;
upstream_api_key sk-your-key;

embedding_provider openai_compatible;
embedding_base_url https://api.openai.com;
embedding_api_key sk-your-key;

The upstream provider and embedding provider may use different OpenAI-compatible base URLs.

Important limitations:

  • AI Cost Firewall does not claim universal compatibility with every OpenAI-like API.
  • Native Anthropic, Gemini, Mistral, and Cohere APIs are not currently supported directly.
  • Mistral, Anthropic, Gemini, or other providers may be used only when exposed through an OpenAI-compatible layer such as LiteLLM, OpenRouter, or another compatible gateway.

See:

configs/examples/
deploy/examples/
docs/provider-compatibility.md

Metrics Overview

Metrics are exposed at:

http://localhost:8080/metrics

Example metrics:

aif_requests_total
aif_cache_exact_hits
aif_cache_semantic_hits
aif_cache_hits_total
aif_cache_bypass_requests_total
aif_model_cost_micro_usd_total
aif_gross_saved_micro_usd_total
aif_net_saved_micro_usd_total
aif_embedding_overhead_micro_usd_total
aif_guard_requests_total
aif_guard_latency_seconds
aif_security_blocks_total
aif_privacy_restore_skipped_total

AI Cost Firewall reports:

  • gross chat-completion savings
  • embedding overhead
  • net savings after embedding cost
  • cache hit ratios
  • semantic cache diagnostics
  • per-model traffic and cost metrics
  • guard request counts by guard, stage, and result
  • Security Guard block counts by stage and rule ID
  • Privacy Guard restore-skip counters when response Security blocks occur

Evidence events are emitted through structured application logs and can optionally be delivered to VCAL Audit. Enable evidence logging with:

RUST_LOG=info,vcal_evidence=info

Configuration

AI Cost Firewall uses a simple nginx-style configuration format.

Minimal example:

listen_addr 0.0.0.0:8080;

redis_url redis://redis:6379;

upstream_provider openai_compatible;
upstream_base_url https://api.openai.com;
upstream_api_key sk-your-key;

semantic_cache_enabled true;

Full documentation:

  • docs/config-reference.md
  • docs/provider-compatibility.md
  • docs/quickstart.md

Benchmarks

AI Cost Firewall has been benchmarked with a local simulated OpenAI-compatible upstream provider to isolate gateway behavior, Redis/Qdrant integration, cache effectiveness, and Prometheus metrics without external API cost or provider rate-limit noise.

In a 30-minute cache-effectiveness benchmark, AI Cost Firewall sustained 30 RPS with 0% request failures, p95 latency of 9.03 ms, and a 98.86% aggregate cache-hit rate.

In a single-VM high-load benchmark, AI Cost Firewall sustained approximately 500 RPS for 5 minutes with 0% HTTP failures. Higher RPS values caused instability in the single-VM test environment, so this should be treated as a local benchmark observation, not a universal capacity limit.

See BENCHMARKS.md for benchmark methodology, environment, limitations, and detailed results.


Evidence Events and VCAL Audit

AI Cost Firewall emits structured evidence using:

vcal.evidence.event
schema_version: 1.1

Each request trace uses a stable trace_id to correlate request validation, cache activity, upstream calls, guard decisions, response processing, and the terminal request outcome.

Every received request trace ends with exactly one terminal event:

request.completed

or:

request.failed

AI Cost Firewall also emits structured evidence for VCAL Security Guard and VCAL Privacy Guard activity.

Guard evidence contains operational metadata only. Prompt and response content is not included.

Evidence can be emitted through structured application logs and optionally delivered asynchronously to VCAL Audit.

VCAL Audit provides:

  • authenticated evidence ingestion
  • durable event persistence
  • ordered trace reconstruction
  • NDJSON export
  • integrity verification

See docs/audit-integration.md for configuration, delivery behavior, retry handling, and operational guidance.


Troubleshooting

See:

  • docs/troubleshooting.md
  • docs/provider-compatibility.md
  • docs/operation.md

Common issues include:

  • incorrect upstream base URLs
  • provider TLS/certificate failures
  • embedding dimension mismatches
  • Qdrant vector-size mismatch
  • unsupported provider behavior
  • semantic threshold tuning

Documentation

Document Description
docs/architecture.md System architecture
docs/config-reference.md Configuration directives
docs/faq.md Frequently asked questions
docs/how-it-works.md Request flow and cache logic
docs/metrics-and-costs.md Cost and savings accounting
docs/operation.md Runtime behavior
docs/provider-compatibility.md OpenAI-compatible providers
docs/quickstart.md Extended setup guide
docs/troubleshooting.md Troubleshooting guide
docs/audit-integration.md VCAL Audit evidence delivery and operations

Full documentation:

https://ai-firewall.docs.vcal-project.com/


Build from Source

git clone https://github.com/vcal-project/ai-firewall.git
cd ai-firewall

cargo build --release
cargo run --release

Testing

Run tests:

cargo test

AI Cost Firewall includes tests for:

  • configuration validation
  • request validation
  • semantic cache requirements
  • semantic cache fail-open behavior
  • environment variable parsing
  • request size parsing
  • cost accounting logic
  • guard configuration parsing
  • Privacy Guard orchestration
  • Security Guard orchestration
  • OpenAI-compatible metadata preservation
  • evidence lifecycle completion
  • request and response Security Guard block evidence
  • global streaming rejection
  • buffered Audit evidence delivery configuration
  • evidence batching, retry, and delivery failure handling

Contributing

Contributions are welcome.

Areas where contributions are especially valuable:

  • documentation
  • performance
  • observability
  • provider compatibility
  • deployment examples
  • testing

See:

CONTRIBUTING.md

Integration with VCAL Semantic Cache

AI Cost Firewall can optionally integrate with VCAL Semantic Cache for advanced semantic caching and distributed vector storage.

https://vcal-project.com/vcal-server


Releases

For release-specific features, fixes, compatibility notes, and validation results, see the GitHub Releases page.


License

Apache License 2.0

About

OpenAI-compatible LLM gateway that reduces API costs using Redis exact cache and Qdrant semantic cache.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages