ExAgent
ExAgent 2.0 is available on Hex. When upgrading from 1.x, read the migration guide for the runtime, model, event and snapshot changes. See supported features and limits for deployment requirements.
An agent framework for Elixir — structured output, tool-calling, streaming, stateful agents, multi-agent sessions and durable persistence, powered by the BEAM.
ExAgent is layered and opt-in: use just the one-shot core, or stack on the
stateful runtime, persistence and coordination as you need them. It is built the
Elixir way — recursion, behaviours, Ecto changesets, cheap concurrency for tools,
supervision/durability, :telemetry, and events that plug straight into LiveView.
Layer 3 ExAgent.Session coordinated multi-agent turns + shared state
Layer 2 ExAgent.Store snapshots: resume after crash / restart
Layer 1 ExAgent.Server a supervised, stateful, event-emitting agent
Layer 0 ExAgent.run/3 the one-shot model ⇄ tools loop
──────────────────────── events (ExAgent.Event) over ExAgent.PubSub
Features
- One-shot agentic loop — a model ⇄ tools recursion built as idiomatic Elixir.
- Type-derived tool schemas — define tools as plain functions; JSON Schema is
generated from
name :: Typeannotations and@docstrings (no hand-written schemas). - Structured output — any Ecto
embedded_schemabecomes the output spec; JSON Schema is derived from the schema and its changeset validations, validated with retry-on-failure. - Streaming — a lazy, demand-driven view of the same tool/output loop, with provisional deltas and one complete result or partial failure.
- Supervised stateful agents — keep history, accumulate usage, thread stateful models across runs, and emit versioned events over PubSub (LiveView-ready).
- Opt-in checkpoints — confirm complete state through a Store and restore a conversation on restart. ETS is in-process; Postgres uses your DB/repo. This does not replay or roll back arbitrary in-flight effects.
- Multi-agent sessions — coordinated turns over shared state with pluggable
turn policies (
round_robin,initiative, or your own). - Orchestration — scoped delegation with ancestor budgets/permissions and hand-off between session participants, durable sequences, routing and bounded parallel flows.
- Persisted human approval — pause at an admitted boundary and resume using versioned continuation records, explicit authority and recovery rules.
- Robustness & safety — context compaction, usage/cost limits, and per-tool
permissions (
allow/ask/deny). - Model-agnostic — one stock ReqLLM backend with explicitly qualified profiles,
custom specs/gateways and the public
ExAgent.Modelextension behaviour. - External tools (MCP) — consume stdio or opt-in Streamable HTTP tools with caller-owned timeouts, pending limits and transport cleanup.
- Observable —
:telemetry, app-levelExAgent.Eventenvelopes and opt-in native OpenTelemetry tracing and metrics. ExAgent owns its execution spans; the ReqLLM bridge supplies metadata without duplicate model spans. Content is disabled by default; the application owns its SDK and exporter. - Offline-first testing — a deterministic
ExAgent.Models.Testmodel drives the full loop with no API key and no network.
Requirements
- Elixir 1.18+ (required by ReqLLM's mandatory
llm_dbdependency) - Runtime targets: Elixir 1.18 / OTP 28 and Elixir 1.20 / OTP 29
See supported features and limits for runtime and dependency requirements. These targets do not cover every patch release or dependency combination.
ReqLLM starts its own supervisor and Finch pool alongside ExAgent. Its default
startup loads .env in the host's working directory. Applications that manage
credentials themselves should set config :req_llm, load_dotenv: false before
startup. ExAgent.Models.ReqLLM.new/1 is the general backend; custom Model/Test
remain extension points. Tools, Ecto-tool output and streaming use an explicit
Chat profile: :chat_tools_v1, or :openrouter_chat_tools_v1 for OpenRouter
provider routing. Native JSON Schema uses a separate opt-in profile. Catalogue
resolution alone does not enable those capabilities. See
models and limits for credentials, profiles,
routing, normalized usage and host retention boundaries.
Installation
Add the published package to your application:
def deps do
[{:exagent, "~> 2.0"}]
end
To develop against a source checkout, replace that dependency with
{:exagent, path: "../exAgent"} and adjust the path to your clone.
The library starts its own supervised ExAgent.Finch HTTP pool, a Registry
(ExAgent.PubSub.Local), a Task.Supervisor, an ExAgent.Store.ETS table and
an ExAgent.AgentSupervisor, so it works out of the box. Tune the Finch pool
with:
config :exagent, :finch_pools, %{:default => [size: 32]}
ExAgentdoes not shadow OTP'sAgentunless you alias it asAgent.
Quick start
The fastest way to try ExAgent is with Mix.install/2 (Livebook or a script) —
using the built-in ExAgent.Models.Test model, no API key needed:
Mix.install([
{:exagent, "~> 2.0"}
])
agent = ExAgent.new(model: "test", instructions: "Be concise.")
{:ok, %{output: text}} = ExAgent.run(agent, "Hello!")
Resolve a stock catalogue string or tuple with explicit instance credentials:
{:ok, model} = ExAgent.Model.resolve({:openai, id: "gpt-4o"},
api_key: System.fetch_env!("OPENAI_API_KEY"))
agent = ExAgent.new(model: model, instructions: "Be concise.")
{:ok, %{output: text}} = ExAgent.run(agent, "Hello!")
Table of Contents
- Layer 0 — the one-shot loop
- Layer 1 — a stateful, supervised agent
- Layer 2 — snapshots & resume
- Layer 3 — multi-agent sessions
- Coordination
- Robustness & safety
- External tools (MCP)
- Events & PubSub
- OpenTelemetry
- Models
- Examples
- Documentation
- Contributing
- License
Layer 0 — the one-shot loop
The core is a small loop: UserPromptNode → ModelRequestNode ⇄ CallToolsNode → End.
agent = ExAgent.new(model: "test", instructions: "Be concise.")
{:ok, %{output: text}} = ExAgent.run(agent, "Hello!")
Operational loop failures return {:error, %ExAgent.RunError{reason: cause, partial: result}}; success returns {:ok, result}. Invalid construction options
may still raise. The
result map carries :output, :messages, :new_messages, :usage
(%{input_tokens:, output_tokens:}), :run_step and the (possibly updated)
:model, plus run/tree/request IDs, status, usage completeness, scoped counters
and estimated :cost_cents / :cost_status. Inspect error.reason for classification
and retain error.partial for reconciliation; never treat partial output as
authorization to act. Live results may contain model credentials: use safe error
projections for logs/events instead of serializing the whole result.
Tools with derived schemas
For tools/streaming, explicitly declare the actual protocol and capabilities of your chosen model/endpoint. These declarations are configuration, not proof that every model in a catalogue supports this profile. This example selects Chat with non-reasoning tools; qualify your exact backend before production use:
chat_model = ExAgent.Models.ReqLLM.new(
model: %{provider: :openai, id: "gpt-4o-mini",
capabilities: %{tools: %{enabled: true}, reasoning: %{enabled: false}},
extra: %{wire: %{protocol: "openai_chat"}}},
api_key: System.fetch_env!("OPENAI_API_KEY"),
tool_profile: :chat_tools_v1
)
Reuse chat_model in the following tool, output, runtime and coordination examples.
For OpenRouter, select tool_profile: :openrouter_chat_tools_v1 and configure
provider routing through provider_options: [openrouter_provider: ...].
Routing is bound to the conversation; a follow-up cannot silently change it.
The model guide shows the full constructor.
Define a tool
defmodule MyApp.Tools do
use ExAgent.Tools
@doc "Get the weather for a city."
deftool get_weather(_ctx, city :: String.t(), days :: integer()) do
{:ok, "#{city}: sunny for #{days} day(s)"}
end
end
agent = ExAgent.new(model: chat_model, tools: MyApp.Tools.tools())
deftool receives the ExAgent.RunContext as its first arg (named ctx by convention);
tool_plain takes only parameters. Each parameter is name :: Type, so the JSON
Schema is derived for you. A tool may return value, {:ok, value} or
{:error, reason}. Arguments are validated locally before invocation, and results
must be JSON-portable. Use ExAgent.ModelRetry explicitly for a correctable
rejection before an uncertain effect; execution exceptions/timeouts are not
automatic permission to replay it. Valid tool schemas are prepared on the reusable
agent definition and checked again if their schema changes.
Structured output
Any embedded_schema becomes the output spec; JSON Schema is derived from the
schema and its changeset validations (validate_inclusion → enum,
validate_number → minimum/maximum, validate_length → minLength/
maxLength), then validated with the changeset, with retry-on-failure.
defmodule WeatherReport do
use Ecto.Schema
embedded_schema do
field :city, :string
field :temp_c, :float
field :condition, Ecto.Enum, values: [:sunny, :rainy, :cloudy]
end
def changeset(s, a) do
s |> Ecto.Changeset.cast(a, [:city, :temp_c, :condition])
|> Ecto.Changeset.validate_required([:city, :temp_c])
|> Ecto.Changeset.validate_number(:temp_c, greater_than: -100, less_than: 100)
end
end
# → the model is told temp_c is a number in (-100, 100) and condition is one of
# the enum values, so it can comply instead of guessing and being retried.
agent = ExAgent.new(model: chat_model, output: WeatherReport)
{:ok, %{output: %{__struct__: WeatherReport}}} =
ExAgent.run(agent, "It's 22 and sunny in Madrid")
Tool mode is the default. For native JSON Schema, set
output_profile: :chat_json_schema_v1 on the qualified Chat model above and use
ExAgent.new(model: native_model, output: WeatherReport, output_mode: :native).
This sends a separate, non-strict schema without an output tool; function tools
still use their validated envelope. Optional fields/defaults are preserved, and
the final changeset remains authoritative. Corrective retries consume the same
host request budget in sync, stream_text and run_stream; deltas are provisional.
No automatic tool fallback or JSON repair is performed.
Stock stream objects become semantic JSON text in history, not original bytes. Stock Chat 1.26 may discard a refusal field beside otherwise valid JSON: that JSON can produce a locally valid output. Exposed refusals, absent/invalid output and incomplete terminals fail; total wire refusal detection is not promised. See the migration guide for exact qualification and limits.
Streaming
ExAgent.run_stream(agent, "count to five")
|> Stream.each(fn
{:delta, t} -> IO.write(t)
{:result, %{usage: u}} -> IO.puts("\n#{u.output_tokens} tokens")
{:error, error} -> IO.puts(Exception.message(error))
end)
|> Stream.run()
ExAgent.run_stream/3 uses the full loop, including tools, hooks, Ecto output
validation and limits. Deltas are provisional across all model requests; the final
result supplies the validated output. Each enumeration is a new run, so enumerate
once. Halting closes owned resources; a deliberately suspended continuation must
be resumed or halted. Custom Model adapters emit a terminal
{:response, response, final_model} or an error.
Serialization / durable runs
The core is DB-free: it doesn't own a database or job queue. It provides best-effort message-history serialization so you can persist a conversation anywhere and resume it:
json = ExAgent.Message.to_json(result.messages) # store this
{:ok, history} = ExAgent.Message.from_json(json) # load it back
ExAgent.run(agent, "follow up", message_history: history)
For persistent job dispatch, an application can use Oban with the public continuation APIs — see the runnable job recipe and its integration guide. A duplicate job inspects the existing record and resumes only a ready/approved boundary. Storing message history alone does not provide mid-run recovery. The application owns its queue, database, authenticated decisions and uncertain-effect recovery.
Layer 1 — a stateful, supervised agent
ExAgent.Server keeps an agent alive across runs: it preserves history,
accumulates usage, threads stateful models, and emits events.
{:ok, dm} =
ExAgent.AgentSupervisor.start_agent(
agent: ExAgent.new(model: chat_model, instructions: "You are a DM."),
agent_id: "dm",
pubsub: :local
)
{:ok, %{output: _}} = ExAgent.Server.chat(dm, "I enter the tavern.") # synchronous
{:ok, %{output: _}} = ExAgent.Server.chat(dm, "I pick the lock.") # sees prior turn
# Async: returns immediately, result arrives as a :run_finished event
{:ok, request_id} = ExAgent.Server.send_message(dm, "describe the room")
ExAgent.Server.abort(dm) # cancel the in-flight run (stays responsive)
ExAgent.Server.health(dm) # %{status: :idle, pending: 0}
While a run is in flight, chat/3 returns {:error, :busy} and send_message/3
enqueues up to max_pending (default 8) then returns {:error, :queue_full}.
Layer 2 — snapshots & resume
Point a Server at a store to confirm complete checkpoints and rehydrate the last confirmed conversation on restart:
ExAgent.AgentSupervisor.start_agent(
agent: agent_template,
agent_id: "dm",
store: :ets # ExAgent.Store behaviour; ETS ships by default
)
The persisted ExAgent.Server.Snapshot carries serializable history, usage and
metadata, not the live model, pids or tool closures. JSON does not detect secrets
introduced as strings by the application. The live model/tools come from a trusted
template on restart. Store is disabled unless configured;
ExAgent.Store.ETS is in-process and depends on its table owner/VM. For a durable DB, use
ExAgent.Store.Postgres (needs ecto_sql + postgrex):
ExAgent.Store.Postgres.migrate(MyApp.Repo) # once
ExAgent.AgentSupervisor.start_agent(
agent: agent_template, agent_id: "dm",
store: {ExAgent.Store.Postgres, MyApp.Repo}
)
A failed save returns ExAgent.CheckpointError while retaining the new state in
memory. Further mutations wait for Server.checkpoint/1 (or Session.checkpoint/1)
to retry only storage, not the model, tools or state-change function. Async queue
admission remains volatile. Restore rejects corrupt/future/mismatched snapshots
instead of starting empty; see the v1/v2 migration guidance.
Layer 3 — multi-agent sessions
ExAgent.Session coordinates participants (agents or humans) taking turns over
a piece of shared state, through a pluggable TurnPolicy. The Session is the
single writer of shared_state.
alias ExAgent.Session
alias ExAgent.Session.Participant
{:ok, game} =
Session.start_link(
shared_state: %{log: []},
policy: {:initiative, order: ["rogue", "fighter"]},
participants: [
Participant.new(id: "rogue", kind: :agent),
Participant.new(id: "fighter", kind: :human)
],
pubsub: :local
)
{:ok, "rogue"} = Session.start(game)
{:ok, world, next} =
Session.take_turn(game, "rogue", fn s -> {:ok, %{s | log: ["rogue acts" | s.log]}} end)
# `next` is now "fighter"; it sees the rogue's change via Session.read_state/1
Tools inside an agent run read/propose state through an
ExAgent.Session.SharedState handle in RunContext.deps, through the Session's
API rather than direct mutation of a shared value. Policies: RoundRobin, Initiative (custom :order),
SupervisorPolicy (a coordinator alternates with workers).
Coordination
ExAgent.Coordination adds the classic orchestration patterns on top of a
Session (levels 2 & 3):
alias ExAgent.Coordination
# Delegation (agent-as-tool): the parent calls a sub-agent; both runs' tokens
# are counted together.
helper = ExAgent.new(model: chat_model, instructions: "You summarize.")
parent =
ExAgent.new(
model: chat_model,
tools: [Coordination.delegation_tool(helper, name: "summarize")]
)
# Hand-off: transfer control between participants directly.
{:ok, "fighter"} = Coordination.handoff(game, "fighter")
Delegation uses one owned execution scope: child rules cannot override ancestor
denial/approval or expand their limits. ExAgent.run_child(context, agent, prompt, opts) is the explicit scoped entry point for custom auxiliary calls. Arbitrary
IO outside that scope is not automatically accounted or sandboxed.
For persisted workflows, ExAgent.Coordination.Composition runs a versioned
sequence and ExAgent.Coordination.Flow runs trusted routing or bounded parallel
branches with ordered merge and explicit failure policy. Both use the same
continuation Store, shared limits and human decision contracts. The
coordination recipes demonstrate typed
extraction, specialist selection and review before an effect; definitions and
model codecs remain trusted host code.
Robustness & safety
Long sessions and cost stay under control, all opt-in:
alias ExAgent.{Compaction, CostGuard, Permissions, UsageLimits}
# Summarize old turns once the context grows (capability hook).
compaction = %Compaction.Capability{
compactor: Compaction.Summary,
opts: [threshold_tokens: 6000, keep_recent: 8, summarize: &MyApp.summarize/1]
}
# Per-tool admission control (allow/ask/deny with globs).
perms = Permissions.new!(rules: [{"*", :deny}, {"read", :allow}, {"bash", :ask}])
agent =
ExAgent.new(
model: chat_model,
capabilities: [compaction],
usage_limits: %UsageLimits{request_limit: 20, tool_calls_limit: 15,
max_budget_cents: 25, accounting: :estimated}
)
ExAgent.run(agent, "go",
permissions: perms,
approve: &MyApp.ask_human/1, # called on :ask
# Illustrative cents per 1K tokens; supply your provider's actual prices.
estimate_cost: CostGuard.estimator(%{input_per_1k_cents: 0.25, output_per_1k_cents: 1.0})
)
Compaction projects request context while retaining canonical history; it does not
bound persisted history size. A caller's one-argument summarizer may perform IO
outside the scope. Cost is estimated per request, retaining fractional cents;
heterogeneous model pricing uses an arity-two (model, usage) estimator. Missing
usage/price is unknown, not zero, and requests already in flight can exceed a
retrospective token/cost threshold. Run options deadline (absolute monotonic
milliseconds) and max_concurrent_requests control admission separately.
External tools (MCP)
Consume a stdio Model Context Protocol server's
tools as plain ExAgent.Tools:
alias ExAgent.MCP.Client
{:ok, fs} =
Client.start_link(
command: "npx",
args: ["-y", "@modelcontextprotocol/server-filesystem", "./data"]
)
{:ok, tools} = Client.tools(fs) # [ExAgent.Tool.t(), ...]
agent = ExAgent.new(model: chat_model, tools: tools)
The client owns the stdio JSON-RPC connection (handshake, tools/list,
tools/call, line buffering); transport exits and errors surface cleanly.
Defaults are 128 pending requests and 8 MiB frames, configurable. A timeout/dead
caller is cleaned up locally; that does not prove rollback of a remote effect.
Streamable HTTP is opt-in and uses an application-owned HTTP1-only Finch pool.
See MCP setup for protocol profiles, continuation binding
and the qualified independent SDK scenarios.
Events & PubSub
Every layer emits versioned ExAgent.Event envelopes (distinct from
:telemetry). Subscribe to drive a UI:
:ok = ExAgent.PubSub.subscribe({ExAgent.PubSub.Local, []}, ExAgent.Event.agent_topic("dm"))
receive do
{:exagent_event, %ExAgent.Event{type: :run_finished, payload: p}} ->
IO.puts("done: #{inspect(p)}")
end
ExAgent.PubSub is a behaviour: None (default, no-op), Local (Registry),
Phoenix (delegates to Phoenix.PubSub dynamically — no hard dependency), or
your own.
OpenTelemetry
Opt into native Erlang/Elixir OTel with an application-owned SDK and exporter:
tracing = ExAgent.Observability.OpenTelemetry.new()
agent = ExAgent.new(model: model, observability: tracing)
Spans cover runs, model requests, tools, delegation, compaction and checkpoints. Content is off by default; enabling it requires explicit redaction before attributes reach the exporter. Request usage and inclusive run totals are kept separate to avoid double counting. The optional bounded processor provides an asynchronous export route with observable saturation and failures.
For applications also tracing standalone ReqLLM calls, attach
ExAgent.Observability.ReqLLM.attach/1 once at host startup instead of the stock
ReqLLM bridge. It enriches the existing ExAgent Model span and preserves standalone
tracing. ExAgent keeps usage, cost, status, privacy and span lifecycle; conflicting
stock bridges reject before provider IO. The host maintains ReqLLM's tracking TTL.
The same bridge can opt into four metric instruments with bounded model labels.
Long-lived applications can supervise the optional
ExAgent.Observability.ReqLLM.Maintenance child to clean up expired tracking.
Neither option installs a metric SDK or exporter for the application.
See observability for application configuration, context propagation, privacy, Langfuse/Opik configuration and export limits. The application supplies its backend credentials and chooses a separate metric destination when needed. Neither backend is a required dependency.
Models
Resolve stock specs with explicit options, or pass a custom Model struct:
{:ok, model} = ExAgent.Model.resolve("openai:gpt-4o",
api_key: System.fetch_env!("OPENAI_API_KEY"))
ExAgent.new(model: model)
ExAgent.new(model: %ExAgent.Models.Test{script: ["offline"]})
Strings alone resolve identity; they do not load credentials or enable tools/Chat.
resolve/2 accepts public stock map/tuple/LLMDB specs and instance options; custom
Model structs remain unchanged through resolve/1. Unknown options reject.
OpenCode Go/Zen require explicit endpoints; stock zai: means ZAI, not the old
Anthropic gateway alias. See the exact migration recipes. Bring your own model by
implementing ExAgent.Model, without private ReqLLM APIs. Automatic retries and
redirects are disabled; stream limits are documented on ExAgent.Models.ReqLLM.
Incomplete/invalid terminal responses fail before effects. Usage is normalized,
provider presence unknown, and costs estimated; strict metric limits reject this
contract, while host request/tool limits remain exact admission counters.
Examples
examples/demo.exs— offline loop with the TestModel.examples/openrouter.exs— live tool-calling via OpenRouter.examples/structured_output.exs— live structured output via Ecto.examples/streaming.exs— live SSE streaming.examples/stateful_agent.exs— supervised stateful agent + events.examples/observability.exs— native OTel with a local exporter; optional--benchmeasures synthetic instrumentation overhead.examples/framework_evals.exs— deterministic offline evaluations in two domains, with typed output, scoped delegation and checkpoint-only recovery;--json <path>writes the machine-readable result.examples/multi_agent_session.exs— two agents, round-robin, shared state.examples/coordination_workflows.exs --run— offline typed extraction, selected specialist and a human-reviewed sequence; run with--no-start.examples/flow_pipeline.exs --self-test— offline router and bounded parallel specialists with a delegated child, also usable as a native tracing callback.examples/dnd_session.exs— a mini D&D round: DM + bot + human over a shared world, coordinated by a Session (SupervisorPolicy), offline.
Run any of them with mix run examples/<name>.exs (live ones need an API key in
the environment).
Documentation
Start with the documentation home or follow the
getting-started tutorial. Task guides cover
tools/output,
models/limits,
runtime/events,
durability/approvals,
coordination and testing.
Coding-agent integration notes map tasks to public APIs;
ExDoc generates llms.txt and Markdown pages from the same sources.
- Full module reference on hexdocs
- Documentation index — task guides and integration recipes.
- Supported features and limits — profiles and deployment boundaries.
- Architecture — layers and ownership.
- Migration — upgrading contracts from 1.x to 2.0.
- Observability — optional native OTel setup and privacy.
- Testing — deterministic tests for your application.
Contributing
Bug reports and pull requests are welcome on GitHub.
For source contributions, follow AGENTS.md in the repository and run
./bin/check using the project-local tooling. Integration issues should include
a minimal public-API example, package and Elixir/OTP versions, and a sanitized
error. Keep provider credentials and application data out of reports.
EXAGENT_OFFLINE=1 MIX_ENV=test mix compile --warnings-as-errors
EXAGENT_OFFLINE=1 MIX_ENV=test mix test --warnings-as-errors
EXAGENT_OFFLINE=1 MIX_ENV=test mix run examples/demo.exs
EXAGENT_OFFLINE=1 skips even the Postgres bootstrap. Real-provider tests are
tagged :integration and excluded by default; enable them only for an explicitly
scoped backend check. Without offline mode, Postgres tests attempt their database
and auto-skip if unavailable. Mock/stdio/local-HTTP tests do not prove acceptance
by external providers or a durable database.
License
Copyright (c) 2025 kukapu
Licensed under the MIT License — see LICENSE.