Laroute

Self-Hosted AI Gateway

One OpenAI-compatible API for cloud and private models, with keys, spending limits, request logs, and usage reports built in.

One

API for all models

OpenAI

Compatible

100%

Self-Hosted

One OpenAI-compatible API for cloud and private models.
Laroute brings keys, limits, logs, and usage reports to your own infrastructure.

Why Laroute?

One API, Every Model

Use one base URL and key format across all connected providers.

Cloud & Private Providers

Connect OpenRouter, OpenAI-compatible services, and private vLLM servers.

Keys & Spending Limits

Create virtual keys with expiry dates and per-period budgets.

Team Sharing

Shared app accounts with their own keys, budgets, and request history.

Full Request Logs

See model, key, result, token use, cost, and response time per request.

AI Usage Analytics

Detect tools like Cursor and Cline, and rank apps by tokens and spend.

Up-to-Date Model List

Refresh provider model lists hourly with admin approval.

Proactive Alerts

Get notified when providers are down or error rates climb.

Self-Hosted

Runs as one Go application with SQLite and an included Docker image.

Built-In AI Workspace

Chat directly with models, with persistent conversations and live budget insight.

Skills & Plugins

Personalize chats with reusable skills and connect external tools.

Administration & Onboarding

Manage keys, models, and users with a guided first-run setup.

Product Tour

Unified API & Every Provider

Use one base URL and key format across all connected providers, with the model names your team already knows.

Single Endpoint

One base URL and key format across all connected providers.

OpenAI-Compatible

Works with clients that support the OpenAI API format.

Chat, Text, Embeddings & Responses

Supports chat, text generation, embeddings, and the OpenAI Responses API.

Live Streaming

Supports regular responses and live streaming.

Stable Model Names

Keeps public model names stable even when provider model names change.

Familiar Errors

Returns errors in the same basic format as OpenAI.

Cloud & Private Providers

OpenRouter, OpenAI-compatible services, private vLLM servers, and Gemini on Google Distributed Cloud.

Custom CA Certificates

Custom CA certificates for private providers.

Loop Protection

Loop protection when one laroute gateway connects to another.

Keys, Budgets & Team Sharing

Create virtual keys for people and apps, keep provider keys hidden, and give teams shared accounts with their own budgets.

Virtual Keys

Separate virtual keys for people and apps.

Hidden Provider Keys

Provider keys stay hidden from users.

Key Expiry

Add expiry dates to keys.

Spending Limits

Set daily, weekly, or monthly spending limits.

HTTP 402 on Exhaustion

Stop new requests when the budget is used up.

Self-Service Keys

Users manage their own keys through an optional API.

Shared App Accounts

Team accounts with their own keys and budgets.

Per-App History

Keep a request history for each app.

Member Visibility

Members see the apps they belong to.

Observe Every Request & Usage

Inspect every call, detect the agents consuming tokens, and keep the model list current.

Request Details

Model, key, result, token use, cost, and response time.

Saved Previews

Open a request to view request and response previews.

Filters

Filter requests by result, model, and date.

Streamed Detection

See whether a response was streamed.

Tool Detection

Detects Cursor, Cline, Aider, OpenWebUI, and common SDKs.

App Rankings

Rank apps by tokens, requests, or spend.

Errors & Speed

Error rates and average response times.

Model Catalogue

Central searchable model list with pricing and context info, synced hourly with admin approval.

Company-Wide View

Switch between personal and company-wide usage.

Experience the AI Workspace

Chat directly with models with persistent conversations, tools, and live budget insight.

Request Demo

A Built-In AI Workspace

Use laroute directly as an AI chat workspace, not just an API gateway, with persistent conversations and live budget insight.

Native Chat Workspace

Chat with models directly, alongside the API gateway.

Model & Reasoning Control

Pick models and adjust reasoning effort right from the composer.

Live Usage Indicators

See context usage and budget status while chatting.

Model Price & Context

View pricing and context limits before choosing a model.

Persistent & Resumable Replies

Replies keep streaming through refresh and resume when you return.

Long-Chat Compaction

Compact long conversations to continue past the context limit.

Reply Regeneration

Regenerate answers and keep alternate versions.

Tool Activity & Call Cards

See when the assistant uses a tool and what it returned.

Markdown Tables & Video

Render model tables cleanly and attach videos to conversations.

Manage, Administer & Self-Host

Alerts, MCP, skills, plugins, administration, notifications, retention, and a self-hosted runtime on your own systems.

Proactive Alerts

  • Provider downtime alerts
  • Error-rate threshold alerts
  • Budget-level alerts

MCP Management

  • Connect through the MCP endpoint at /mcp
  • OAuth sign-in from clients like ChatGPT
  • Manage accounts, keys, providers, models, and alerts
  • Interactive dashboards in MCP Apps

Employee Portal

  • SSO sign-in
  • Spend, money left, and next budget reset
  • Search and filter request logs
  • Masked API keys
  • Profile page and light/dark/system themes

Self-Hosted Runtime

  • Included Docker image
  • Health check at GET /api/health
  • Optional OpenTelemetry metrics
  • Provider keys, settings, and logs stay on the laroute server

Skills

  • Reusable personal skills
  • Create and manage skills from the Customize area
  • Import a skill from Markdown
  • Start a skill with a chat prompt

Plugins & Connectors

  • Personal and public workspace plugins
  • Guided connector setup with health checks
  • OAuth connector sign-in
  • Embedded connector interfaces
  • Connector failure isolation

Administration & Onboarding

  • Administration panel with search across collections
  • Request traffic separation and filtering
  • Key, provider, and model management
  • Guided first-run setup and admin onboarding tour

Notifications

  • Notification center
  • Per-user notification preferences
  • Admin alert subscriptions
  • Signup and key-change notices

Privacy & Retention

  • Configurable prompt and response retention
  • Keep usage history without keeping content forever
  • Clear retention status

Licensing

  • Licence status in administration
  • Licence expiry warnings
  • Degraded or inactive licence notices

Get Started with Laroute

Deploy on your own infrastructure: one OpenAI-compatible gateway for every AI model your team uses, with full sovereignty.

Deploy Laroute