Product · Designed & built in Belgium

The AI-Firewall Gateway

One secure gateway to all your AI models, local and cloud. The gateway automatically detects sensitive data, masks it, blocks prompt injections, routes every request to the right model and keeps an audit log of every request. Your teams simply use AI. Your data stays inside the walls.

OpenAI-compatible API Your models, one endpoint Masks PII automatically No external dependencies

The pipeline

What happens to every request, in less than the blink of an eye

Every request passes through five automatic checks before it reaches a model. This is no brochure flowchart: it is the exact order in which the code runs them.

1

Sensitive-data detection always first

Emails, phone numbers, card numbers, passwords, API keys, public IP addresses, detected with regex patterns and semantic keywords. A hit? The request is forced to a local model, whatever you asked for.

2

Prompt-injection containment

"Ignore all previous instructions" and its friends are detected in user messages. Such an attack is never forwarded: the request is handled locally and safely at once.

3

Intelligent routing five layers

Every request lands on the right model: explicit tags, configurable regex routes and, if nothing matches, semantic similarity via embeddings. Your rules, live-adjustable, without a restart.

4

Automatic PII masking

Sensitive data is masked in the prompt itself before sending: [EMAIL], [PHONE], [CRED]. Internal IP addresses stay readable, so technical questions remain legible.

5

Audit log & traceability

Every request gets a unique ID (or inherits your X-Request-ID). Routed model, classification, latency, token usage and masked text: fully traceable, append-only.

Defense in depth, built in: suppose a route points at a cloud model after all, and the request does contain sensitive data: the request is then forced back to a local model at the very last moment. Sensitive content simply never reaches the cloud. Period.

Features

Everything the gateway does for you

One endpoint to all your models, with a security layer on by default, not as an option.

Core feature

One endpoint, all your models

Your teams, scripts and agents talk to one OpenAI-compatible address. Behind that single door sit your local models (Ollama) and, if you allow it, carefully chosen cloud models. Switching models? One line in the config; no client notices a thing.

Smart routing

Every request automatically goes to the best model: coding questions to your coding model, writing to your writing model, sensitive content always local.

  • Explicit tags in the request itself ([[fast]], [[strong]], your own tags)
  • Regex routes, editable in the UI & hot-reloaded
  • Semantic similarity via local embeddings

PII firewall & masking

Credit cards, emails, phones, keys, passwords, public IPs: all detected and masked, in the prompt and in the logs alike.

  • Sly patterns: cards with spaces/dashes, Amex, API keys
  • Keyword catcher for medical, banking and ID contexts
  • System prompts are never touched

Injection blocking

Jailbreak attempts and "ignore all instructions" tricks are detected and kept out of the cloud, automatically, without you lifting a finger.

Self-healing failover

A model degrades or drops? The gateway automatically switches to your fallback chain, with retries and cooldowns. Your app notices nothing.

  • Per-model fallback chains, defined by you
  • 2 retries per model, cooldown after repeated failures

Full audit logging

Append-only log of every call: masked input & output, routed model, why that route, latency and token usage, all correlatable via request ID.

Built-in SIEM

A detection engine that watches LLM and auth logs live. Define your own use cases with a simple match DSL; alerts go to the UI and to your webhook.

  • Ready out of the box: sensitive-data-to-cloud (critical), login bursts, slow models…
  • Threshold and single-event triggers, with cooldown deduplication

Auth without compromise

Zero-dependency authentication, built the way it should be: scrypt hashing, HMAC sessions, optional TOTP 2FA, service keys for agents, and lockout with exponential backoff.

Your own dashboard

Live monitor, playground to test routes before you roll them out, usage dashboard, rules editor with classification test: no external service, just runs at your place.

Under the hood

This is what a logged request looks like

Every event is written append-only, readable with any tool you already have. This is a real log line, as it comes in:

Note: the client had asked for the fast cloud model. The gateway detected the sensitive data, masked it in the log, and routed the request to the local private model instead. The client noticed nothing, except that it just worked.

Specifications

Technical at a glance

API & integration

API protocolOpenAI-compatible (chat-completions, streaming)
One endpointYour entire model fleet behind one URL
AuthenticationMaster key + service keys per agent/app/script
Request-ID tracingInherit X-Request-ID or automatic

Routing & models

BackendsOllama (local) + optionally permitted cloud models
Routing layersPII → tag → regex rules → semantic (embeddings) → default
Semantic routerLocal embeddings, threshold & examples in the UI
FailoverPer-model fallback chains, retries, cooldown

Security

PIIRegex + keywords + phone + public IPs; masking in prompts and logs
InjectionDetected jailbreak patterns → forced local
Defense in depthSensitive + cloud route → last resort: local model
AuthScrypt + HMAC sessions + optional TOTP + lockout backoff

Management & logging

DashboardRules editor, live monitor, playground, usage, SIEM
Live reloadFirewall rules & SIEM use cases without restart
Log formatAppend-only JSONL, with request-ID correlation
WebhooksSIEM alerts to your own endpoint

Architecture

InstallationDocker Compose, entirely on your hardware
Core componentsLiteLLM proxy with routing hook + zero-dependency Node UI
Dependencies (UI)None. No npm packages, no CDNs
CacheDisk cache + semantic route cache

What you get

On-site installationWe set it up, anywhere in Belgium
TrainingYour team learns to manage the gateway itself
SupportOptional management plan, or simply standalone
OwnershipThe code runs at your place. Full stop.

See the gateway in action, with your own questions.

In the live demo we run your real test questions through the gateway, so you can see with your own eyes how PII is detected, masked and kept local. No obligation.

We listen

Which models do you use today, and where does privacy hit a wall?

Design

Your routing rules, your models, your users, in a concrete proposal.

Installation

On your hardware, on your premises. Tested with your real use cases.

Support

You learn to manage the gateway yourself; we stay a call away.