📡
Core feature
One endpoint, all your models
Your teams, scripts and agents talk to one OpenAI-compatible address. Behind that single door sit your local models (Ollama) and, if you allow it, carefully chosen cloud models. Switching models? One line in the config; no client notices a thing.
🧭
Smart routing
Every request automatically goes to the best model: coding questions to your coding model, writing to your writing model, sensitive content always local.
- Explicit tags in the request itself ([[fast]], [[strong]], your own tags)
- Regex routes, editable in the UI & hot-reloaded
- Semantic similarity via local embeddings
🎭
PII firewall & masking
Credit cards, emails, phones, keys, passwords, public IPs: all detected and masked, in the prompt and in the logs alike.
- Sly patterns: cards with spaces/dashes, Amex, API keys
- Keyword catcher for medical, banking and ID contexts
- System prompts are never touched
🚫
Injection blocking
Jailbreak attempts and "ignore all instructions" tricks are detected and kept out of the cloud, automatically, without you lifting a finger.
🔁
Self-healing failover
A model degrades or drops? The gateway automatically switches to your fallback chain, with retries and cooldowns. Your app notices nothing.
- Per-model fallback chains, defined by you
- 2 retries per model, cooldown after repeated failures
📋
Full audit logging
Append-only log of every call: masked input & output, routed model, why that route, latency and token usage, all correlatable via request ID.
🚨
Built-in SIEM
A detection engine that watches LLM and auth logs live. Define your own use cases with a simple match DSL; alerts go to the UI and to your webhook.
- Ready out of the box: sensitive-data-to-cloud (critical), login bursts, slow models…
- Threshold and single-event triggers, with cooldown deduplication
🔐
Auth without compromise
Zero-dependency authentication, built the way it should be: scrypt hashing, HMAC sessions, optional TOTP 2FA, service keys for agents, and lockout with exponential backoff.
📊
Your own dashboard
Live monitor, playground to test routes before you roll them out, usage dashboard, rules editor with classification test: no external service, just runs at your place.