Laroute
Self-Hosted AI Gateway
One OpenAI-compatible API for cloud and private models, with keys, spending limits, request logs, and usage reports built in.
One
API for all models
OpenAI
Compatible
100%
Self-Hosted
One OpenAI-compatible API for cloud and private models.
Laroute brings keys, limits, logs, and usage reports to your own infrastructure.
Why Laroute?
One API, Every Model
Use one base URL and key format across all connected providers.
Cloud & Private Providers
Connect OpenRouter, OpenAI-compatible services, and private vLLM servers.
Keys & Spending Limits
Create virtual keys with expiry dates and per-period budgets.
Team Sharing
Shared app accounts with their own keys, budgets, and request history.
Full Request Logs
See model, key, result, token use, cost, and response time per request.
AI Usage Analytics
Detect tools like Cursor and Cline, and rank apps by tokens and spend.
Up-to-Date Model List
Refresh provider model lists hourly with admin approval.
Proactive Alerts
Get notified when providers are down or error rates climb.
Self-Hosted
Runs as one Go application with SQLite and an included Docker image.
Built-In AI Workspace
Chat directly with models, with persistent conversations and live budget insight.
Skills & Plugins
Personalize chats with reusable skills and connect external tools.
Administration & Onboarding
Manage keys, models, and users with a guided first-run setup.
Product Tour
Unified API & Every Provider
Use one base URL and key format across all connected providers, with the model names your team already knows.
Single Endpoint
One base URL and key format across all connected providers.
OpenAI-Compatible
Works with clients that support the OpenAI API format.
Chat, Text, Embeddings & Responses
Supports chat, text generation, embeddings, and the OpenAI Responses API.
Live Streaming
Supports regular responses and live streaming.
Stable Model Names
Keeps public model names stable even when provider model names change.
Familiar Errors
Returns errors in the same basic format as OpenAI.
Cloud & Private Providers
OpenRouter, OpenAI-compatible services, private vLLM servers, and Gemini on Google Distributed Cloud.
Custom CA Certificates
Custom CA certificates for private providers.
Loop Protection
Loop protection when one laroute gateway connects to another.
Keys, Budgets & Team Sharing
Create virtual keys for people and apps, keep provider keys hidden, and give teams shared accounts with their own budgets.
Virtual Keys
Separate virtual keys for people and apps.
Hidden Provider Keys
Provider keys stay hidden from users.
Key Expiry
Add expiry dates to keys.
Spending Limits
Set daily, weekly, or monthly spending limits.
HTTP 402 on Exhaustion
Stop new requests when the budget is used up.
Self-Service Keys
Users manage their own keys through an optional API.
Shared App Accounts
Team accounts with their own keys and budgets.
Per-App History
Keep a request history for each app.
Member Visibility
Members see the apps they belong to.
Observe Every Request & Usage
Inspect every call, detect the agents consuming tokens, and keep the model list current.
Request Details
Model, key, result, token use, cost, and response time.
Saved Previews
Open a request to view request and response previews.
Filters
Filter requests by result, model, and date.
Streamed Detection
See whether a response was streamed.
Tool Detection
Detects Cursor, Cline, Aider, OpenWebUI, and common SDKs.
App Rankings
Rank apps by tokens, requests, or spend.
Errors & Speed
Error rates and average response times.
Model Catalogue
Central searchable model list with pricing and context info, synced hourly with admin approval.
Company-Wide View
Switch between personal and company-wide usage.
Experience the AI Workspace
Chat directly with models with persistent conversations, tools, and live budget insight.
Request DemoA Built-In AI Workspace
Use laroute directly as an AI chat workspace, not just an API gateway, with persistent conversations and live budget insight.
Native Chat Workspace
Chat with models directly, alongside the API gateway.
Model & Reasoning Control
Pick models and adjust reasoning effort right from the composer.
Live Usage Indicators
See context usage and budget status while chatting.
Model Price & Context
View pricing and context limits before choosing a model.
Persistent & Resumable Replies
Replies keep streaming through refresh and resume when you return.
Long-Chat Compaction
Compact long conversations to continue past the context limit.
Reply Regeneration
Regenerate answers and keep alternate versions.
Tool Activity & Call Cards
See when the assistant uses a tool and what it returned.
Markdown Tables & Video
Render model tables cleanly and attach videos to conversations.
Manage, Administer & Self-Host
Alerts, MCP, skills, plugins, administration, notifications, retention, and a self-hosted runtime on your own systems.
Proactive Alerts
- Provider downtime alerts
- Error-rate threshold alerts
- Budget-level alerts
MCP Management
- Connect through the MCP endpoint at /mcp
- OAuth sign-in from clients like ChatGPT
- Manage accounts, keys, providers, models, and alerts
- Interactive dashboards in MCP Apps
Employee Portal
- SSO sign-in
- Spend, money left, and next budget reset
- Search and filter request logs
- Masked API keys
- Profile page and light/dark/system themes
Self-Hosted Runtime
- Included Docker image
- Health check at GET /api/health
- Optional OpenTelemetry metrics
- Provider keys, settings, and logs stay on the laroute server
Skills
- Reusable personal skills
- Create and manage skills from the Customize area
- Import a skill from Markdown
- Start a skill with a chat prompt
Plugins & Connectors
- Personal and public workspace plugins
- Guided connector setup with health checks
- OAuth connector sign-in
- Embedded connector interfaces
- Connector failure isolation
Administration & Onboarding
- Administration panel with search across collections
- Request traffic separation and filtering
- Key, provider, and model management
- Guided first-run setup and admin onboarding tour
Notifications
- Notification center
- Per-user notification preferences
- Admin alert subscriptions
- Signup and key-change notices
Privacy & Retention
- Configurable prompt and response retention
- Keep usage history without keeping content forever
- Clear retention status
Licensing
- Licence status in administration
- Licence expiry warnings
- Degraded or inactive licence notices
Get Started with Laroute
Deploy on your own infrastructure: one OpenAI-compatible gateway for every AI model your team uses, with full sovereignty.
Deploy Laroute




