LLMProxy

Hex.pmDocumentation

A self-hosted, Elixir-native alternative to LiteLLM. Put one API in front of OpenAI, Anthropic, OpenRouter, Kimi Code, OpenAI Codex, and custom OpenAI-compatible providers—then switch models, add fallbacks, and control spend without changing every application that calls them.

LLMProxy gives applications stable model names such as fast, smart, or cheap. You decide where those names run, how traffic fails over, who may use them, and how much they may spend. Provider credentials and usage data stay in infrastructure you control.

Use LLMProxy directly inside an Elixir application:

{:ok, response} =
LLMProxy.chat("Explain supervision trees in three sentences",
model: "fast",
api_key: llm_proxy_key
)
ReqLLM.Response.text(response.message)

Or run it as an OpenAI-compatible service:

curl http://127.0.0.1:4000/v1/chat/completions \
-H "authorization: Bearer $LLM_PROXY_KEY" \
-H "content-type: application/json" \
-d '{"model":"fast","messages":[{"role":"user","content":"Hello"}]}'

Why LLMProxy

One model name, many ways to serve it

Your clients call fast. LLMProxy can send that traffic to one preferred deployment, balance it across several, or fall back when an upstream is unavailable:

┌─ OpenAI / gpt-4.1-mini
client ── model: fast ───┼─ OpenRouter / google/gemini-2.5-flash
└─ private OpenAI-compatible endpoint

Change the route centrally and every client keeps working. The same public model name is available through OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, LLMProxy.chat/2, and ReqLLM.

Installation

Add LLMProxy to mix.exs:

def deps do
[
{:llm_proxy, "~> 0.1"}
]
end

LLMProxy requires Elixir 1.17 or later. Phoenix, SQLite, Igniter, and Incant integrations are optional dependencies; add only the ones your deployment uses.

Choose a mode before configuring storage and serving:

See Getting Started for the complete setup path.

Library mode

Point LLMProxy at the host application's repo and disable its separate HTTP listener when you only need in-process calls or Phoenix-mounted routes:

# config/runtime.exs
config :llm_proxy,
repo: MyApp.Repo,
http_enabled: false,
master_key: System.fetch_env!("LLM_PROXY_MASTER_KEY"),
providers: %{
"openai" => %{api_keys: System.fetch_env!("OPENAI_API_KEYS")}
},
models: [
fast: [route: [to: :openai, model: "gpt-4.1-mini"]]
]

Install migration aliases and migrate the host repo:

mix igniter.install llm_proxy
mix llm_proxy.migrate

Then call it directly:

{:ok, response} =
LLMProxy.chat("Summarize this incident",
model: "fast",
api_key: System.fetch_env!("LLM_PROXY_MASTER_KEY"),
metadata: %{"incident_id" => "inc_123"},
tags: ["operations"]
)

To expose the same gateway through an existing Phoenix endpoint:

defmodule MyAppWeb.Router do
use MyAppWeb, :router
use LLMProxy.Router
scope "/" do
pipe_through :api
llm_proxy "/llm", setup: false
end
end

This mounts /llm/v1/chat/completions, /llm/v1/messages, /llm/v1/responses, /llm/v1/moderations, model listing, and feedback routes.

See Library Mode for storage, migrations, ReqLLM, Phoenix, and SafeRPC examples.

Standalone mode

The standalone release runs a local HTTP gateway backed by DuckDB through QuackDB. It listens on 127.0.0.1 and is intended to sit behind your reverse proxy or service mesh.

Configure secrets with environment variables:

export MASTER_KEY="replace-with-a-long-random-key"
export OPENAI_API_KEYS="sk-..."
export DATABASE_PATH="/var/lib/llm-proxy/llm_proxy.duckdb"
export PORT="4000"

Configure public providers and model aliases as data in /etc/llm-proxy/config.toml:

[providers.openai-primary]
adapter = "openai"
base_url = "https://api.openai.com/v1"
token_pool = "openai"
[[models]]
name = "fast"
routing = "ordered"
[[models.routes]]
to = "openai-primary"
model = "gpt-4.1-mini"
timeout = 30000

Credentials do not belong in TOML. Bootstrap named pools with LLM_PROXY_PROVIDER_KEYS or manage persisted tokens through the optional admin integration.

The release serves health, model, generation, moderation, and feedback routes. Streaming responses include heartbeats during upstream silence and preserve OpenAI or Anthropic event payloads.

See Standalone Mode and Standalone Deployment for configuration, migrations, release commands, health checks, and shutdown draining.

Routing and providers

A public model can target one or more deployments:

config :llm_proxy,
models: [
fast: [
routing: :lowest_cost,
routes: [
[
to: :openai,
model: "gpt-4.1-mini",
timeout: 15_000,
failure_threshold: 3,
cooldown_ms: 30_000
],
[to: :anthropic, model: "claude-3-5-haiku-20241022", order: 2]
]
]
]

Built-in providers cover OpenAI, Anthropic, OpenRouter, and OpenAI Codex OAuth. For another OpenAI-compatible service, declare a named provider with a ReqLLM adapter, base_url, and isolated token_pool; no provider module is required.

See Providers and Routing for routing semantics, custom endpoints, token pools, Codex OAuth, and native protocol extensions.

Governance and observability

Every authenticated request runs through the same controls:

  1. Resolve the actor and API key.
  2. Check model access, quota windows, and budget limits.
  3. Apply request guardrails and deterministic cache policy.
  4. Select a healthy deployment and credential.
  5. Execute with timeout, retry, fallback, and circuit-breaker handling.
  6. Record usage, estimated cost, latency, trace IDs, messages, and optional bodies.
  7. Emit OpenTelemetry spans and return the trace ID to the caller.

Incant support is optional. When installed, LLMProxy.Admin describes resources for API keys, provider tokens, traces, and messages plus an operations dashboard. A separate Incant host can consume those surfaces over SafeRPC; the public gateway does not expose an admin UI.

See Governance and Observability, Cache and Guardrails, and Admin Integration.

Documentation

Development

mix deps.get
mix ci

License

MIT © 2026 Danila Poyarkov