ExAgent

Hex Version Hex Docs License Release

ExAgent 2.0 is available on Hex. When upgrading from 1.x, read the migration guide for the runtime, model, event and snapshot changes. See supported features and limits for deployment requirements.

An agent framework for Elixir — structured output, tool-calling, streaming, stateful agents, multi-agent sessions and durable persistence, powered by the BEAM.

ExAgent is layered and opt-in: use just the one-shot core, or stack on the stateful runtime, persistence and coordination as you need them. It is built the Elixir way — recursion, behaviours, Ecto changesets, cheap concurrency for tools, supervision/durability, :telemetry, and events that plug straight into LiveView.

Layer 3 ExAgent.Session coordinated multi-agent turns + shared state
Layer 2 ExAgent.Store snapshots: resume after crash / restart
Layer 1 ExAgent.Server a supervised, stateful, event-emitting agent
Layer 0 ExAgent.run/3 the one-shot model ⇄ tools loop
──────────────────────── events (ExAgent.Event) over ExAgent.PubSub

Features

Requirements

See supported features and limits for runtime and dependency requirements. These targets do not cover every patch release or dependency combination.

ReqLLM starts its own supervisor and Finch pool alongside ExAgent. Its default startup loads .env in the host's working directory. Applications that manage credentials themselves should set config :req_llm, load_dotenv: false before startup. ExAgent.Models.ReqLLM.new/1 is the general backend; custom Model/Test remain extension points. Tools, Ecto-tool output and streaming use an explicit Chat profile: :chat_tools_v1, or :openrouter_chat_tools_v1 for OpenRouter provider routing. Native JSON Schema uses a separate opt-in profile. Catalogue resolution alone does not enable those capabilities. See models and limits for credentials, profiles, routing, normalized usage and host retention boundaries.

Installation

Add the published package to your application:

def deps do
[{:exagent, "~> 2.0"}]
end

To develop against a source checkout, replace that dependency with {:exagent, path: "../exAgent"} and adjust the path to your clone.

The library starts its own supervised ExAgent.Finch HTTP pool, a Registry (ExAgent.PubSub.Local), a Task.Supervisor, an ExAgent.Store.ETS table and an ExAgent.AgentSupervisor, so it works out of the box. Tune the Finch pool with:

config :exagent, :finch_pools, %{:default => [size: 32]}

ExAgent does not shadow OTP's Agent unless you alias it as Agent.

Quick start

The fastest way to try ExAgent is with Mix.install/2 (Livebook or a script) — using the built-in ExAgent.Models.Test model, no API key needed:

Mix.install([
{:exagent, "~> 2.0"}
])
agent = ExAgent.new(model: "test", instructions: "Be concise.")
{:ok, %{output: text}} = ExAgent.run(agent, "Hello!")

Resolve a stock catalogue string or tuple with explicit instance credentials:

{:ok, model} = ExAgent.Model.resolve({:openai, id: "gpt-4o"},
api_key: System.fetch_env!("OPENAI_API_KEY"))
agent = ExAgent.new(model: model, instructions: "Be concise.")
{:ok, %{output: text}} = ExAgent.run(agent, "Hello!")

Table of Contents

Layer 0 — the one-shot loop

The core is a small loop: UserPromptNode → ModelRequestNode ⇄ CallToolsNode → End.

agent = ExAgent.new(model: "test", instructions: "Be concise.")
{:ok, %{output: text}} = ExAgent.run(agent, "Hello!")

Operational loop failures return {:error, %ExAgent.RunError{reason: cause, partial: result}}; success returns {:ok, result}. Invalid construction options may still raise. The result map carries :output, :messages, :new_messages, :usage (%{input_tokens:, output_tokens:}), :run_step and the (possibly updated) :model, plus run/tree/request IDs, status, usage completeness, scoped counters and estimated :cost_cents / :cost_status. Inspect error.reason for classification and retain error.partial for reconciliation; never treat partial output as authorization to act. Live results may contain model credentials: use safe error projections for logs/events instead of serializing the whole result.

Tools with derived schemas

For tools/streaming, explicitly declare the actual protocol and capabilities of your chosen model/endpoint. These declarations are configuration, not proof that every model in a catalogue supports this profile. This example selects Chat with non-reasoning tools; qualify your exact backend before production use:

chat_model = ExAgent.Models.ReqLLM.new(
model: %{provider: :openai, id: "gpt-4o-mini",
capabilities: %{tools: %{enabled: true}, reasoning: %{enabled: false}},
extra: %{wire: %{protocol: "openai_chat"}}},
api_key: System.fetch_env!("OPENAI_API_KEY"),
tool_profile: :chat_tools_v1
)

Reuse chat_model in the following tool, output, runtime and coordination examples.

For OpenRouter, select tool_profile: :openrouter_chat_tools_v1 and configure provider routing through provider_options: [openrouter_provider: ...]. Routing is bound to the conversation; a follow-up cannot silently change it. The model guide shows the full constructor.

Define a tool

defmodule MyApp.Tools do
use ExAgent.Tools
@doc "Get the weather for a city."
deftool get_weather(_ctx, city :: String.t(), days :: integer()) do
{:ok, "#{city}: sunny for #{days} day(s)"}
end
end
agent = ExAgent.new(model: chat_model, tools: MyApp.Tools.tools())

deftool receives the ExAgent.RunContext as its first arg (named ctx by convention); tool_plain takes only parameters. Each parameter is name :: Type, so the JSON Schema is derived for you. A tool may return value, {:ok, value} or {:error, reason}. Arguments are validated locally before invocation, and results must be JSON-portable. Use ExAgent.ModelRetry explicitly for a correctable rejection before an uncertain effect; execution exceptions/timeouts are not automatic permission to replay it. Valid tool schemas are prepared on the reusable agent definition and checked again if their schema changes.

Structured output

Any embedded_schema becomes the output spec; JSON Schema is derived from the schema and its changeset validations (validate_inclusion → enum, validate_number → minimum/maximum, validate_length → minLength/ maxLength), then validated with the changeset, with retry-on-failure.

defmodule WeatherReport do
use Ecto.Schema
embedded_schema do
field :city, :string
field :temp_c, :float
field :condition, Ecto.Enum, values: [:sunny, :rainy, :cloudy]
end
def changeset(s, a) do
s |> Ecto.Changeset.cast(a, [:city, :temp_c, :condition])
|> Ecto.Changeset.validate_required([:city, :temp_c])
|> Ecto.Changeset.validate_number(:temp_c, greater_than: -100, less_than: 100)
end
end
# → the model is told temp_c is a number in (-100, 100) and condition is one of
# the enum values, so it can comply instead of guessing and being retried.
agent = ExAgent.new(model: chat_model, output: WeatherReport)
{:ok, %{output: %{__struct__: WeatherReport}}} =
ExAgent.run(agent, "It's 22 and sunny in Madrid")

Tool mode is the default. For native JSON Schema, set output_profile: :chat_json_schema_v1 on the qualified Chat model above and use ExAgent.new(model: native_model, output: WeatherReport, output_mode: :native). This sends a separate, non-strict schema without an output tool; function tools still use their validated envelope. Optional fields/defaults are preserved, and the final changeset remains authoritative. Corrective retries consume the same host request budget in sync, stream_text and run_stream; deltas are provisional. No automatic tool fallback or JSON repair is performed.

Stock stream objects become semantic JSON text in history, not original bytes. Stock Chat 1.26 may discard a refusal field beside otherwise valid JSON: that JSON can produce a locally valid output. Exposed refusals, absent/invalid output and incomplete terminals fail; total wire refusal detection is not promised. See the migration guide for exact qualification and limits.

Streaming

ExAgent.run_stream(agent, "count to five")
|> Stream.each(fn
{:delta, t} -> IO.write(t)
{:result, %{usage: u}} -> IO.puts("\n#{u.output_tokens} tokens")
{:error, error} -> IO.puts(Exception.message(error))
end)
|> Stream.run()

ExAgent.run_stream/3 uses the full loop, including tools, hooks, Ecto output validation and limits. Deltas are provisional across all model requests; the final result supplies the validated output. Each enumeration is a new run, so enumerate once. Halting closes owned resources; a deliberately suspended continuation must be resumed or halted. Custom Model adapters emit a terminal {:response, response, final_model} or an error.

Serialization / durable runs

The core is DB-free: it doesn't own a database or job queue. It provides best-effort message-history serialization so you can persist a conversation anywhere and resume it:

json = ExAgent.Message.to_json(result.messages) # store this
{:ok, history} = ExAgent.Message.from_json(json) # load it back
ExAgent.run(agent, "follow up", message_history: history)

For persistent job dispatch, an application can use Oban with the public continuation APIs — see the runnable job recipe and its integration guide. A duplicate job inspects the existing record and resumes only a ready/approved boundary. Storing message history alone does not provide mid-run recovery. The application owns its queue, database, authenticated decisions and uncertain-effect recovery.

Layer 1 — a stateful, supervised agent

ExAgent.Server keeps an agent alive across runs: it preserves history, accumulates usage, threads stateful models, and emits events.

{:ok, dm} =
ExAgent.AgentSupervisor.start_agent(
agent: ExAgent.new(model: chat_model, instructions: "You are a DM."),
agent_id: "dm",
pubsub: :local
)
{:ok, %{output: _}} = ExAgent.Server.chat(dm, "I enter the tavern.") # synchronous
{:ok, %{output: _}} = ExAgent.Server.chat(dm, "I pick the lock.") # sees prior turn
# Async: returns immediately, result arrives as a :run_finished event
{:ok, request_id} = ExAgent.Server.send_message(dm, "describe the room")
ExAgent.Server.abort(dm) # cancel the in-flight run (stays responsive)
ExAgent.Server.health(dm) # %{status: :idle, pending: 0}

While a run is in flight, chat/3 returns {:error, :busy} and send_message/3 enqueues up to max_pending (default 8) then returns {:error, :queue_full}.

Layer 2 — snapshots & resume

Point a Server at a store to confirm complete checkpoints and rehydrate the last confirmed conversation on restart:

ExAgent.AgentSupervisor.start_agent(
agent: agent_template,
agent_id: "dm",
store: :ets # ExAgent.Store behaviour; ETS ships by default
)

The persisted ExAgent.Server.Snapshot carries serializable history, usage and metadata, not the live model, pids or tool closures. JSON does not detect secrets introduced as strings by the application. The live model/tools come from a trusted template on restart. Store is disabled unless configured; ExAgent.Store.ETS is in-process and depends on its table owner/VM. For a durable DB, use ExAgent.Store.Postgres (needs ecto_sql + postgrex):

ExAgent.Store.Postgres.migrate(MyApp.Repo) # once
ExAgent.AgentSupervisor.start_agent(
agent: agent_template, agent_id: "dm",
store: {ExAgent.Store.Postgres, MyApp.Repo}
)

A failed save returns ExAgent.CheckpointError while retaining the new state in memory. Further mutations wait for Server.checkpoint/1 (or Session.checkpoint/1) to retry only storage, not the model, tools or state-change function. Async queue admission remains volatile. Restore rejects corrupt/future/mismatched snapshots instead of starting empty; see the v1/v2 migration guidance.

Layer 3 — multi-agent sessions

ExAgent.Session coordinates participants (agents or humans) taking turns over a piece of shared state, through a pluggable TurnPolicy. The Session is the single writer of shared_state.

alias ExAgent.Session
alias ExAgent.Session.Participant
{:ok, game} =
Session.start_link(
shared_state: %{log: []},
policy: {:initiative, order: ["rogue", "fighter"]},
participants: [
Participant.new(id: "rogue", kind: :agent),
Participant.new(id: "fighter", kind: :human)
],
pubsub: :local
)
{:ok, "rogue"} = Session.start(game)
{:ok, world, next} =
Session.take_turn(game, "rogue", fn s -> {:ok, %{s | log: ["rogue acts" | s.log]}} end)
# `next` is now "fighter"; it sees the rogue's change via Session.read_state/1

Tools inside an agent run read/propose state through an ExAgent.Session.SharedState handle in RunContext.deps, through the Session's API rather than direct mutation of a shared value. Policies: RoundRobin, Initiative (custom :order), SupervisorPolicy (a coordinator alternates with workers).

Coordination

ExAgent.Coordination adds the classic orchestration patterns on top of a Session (levels 2 & 3):

alias ExAgent.Coordination
# Delegation (agent-as-tool): the parent calls a sub-agent; both runs' tokens
# are counted together.
helper = ExAgent.new(model: chat_model, instructions: "You summarize.")
parent =
ExAgent.new(
model: chat_model,
tools: [Coordination.delegation_tool(helper, name: "summarize")]
)
# Hand-off: transfer control between participants directly.
{:ok, "fighter"} = Coordination.handoff(game, "fighter")

Delegation uses one owned execution scope: child rules cannot override ancestor denial/approval or expand their limits. ExAgent.run_child(context, agent, prompt, opts) is the explicit scoped entry point for custom auxiliary calls. Arbitrary IO outside that scope is not automatically accounted or sandboxed.

For persisted workflows, ExAgent.Coordination.Composition runs a versioned sequence and ExAgent.Coordination.Flow runs trusted routing or bounded parallel branches with ordered merge and explicit failure policy. Both use the same continuation Store, shared limits and human decision contracts. The coordination recipes demonstrate typed extraction, specialist selection and review before an effect; definitions and model codecs remain trusted host code.

Robustness & safety

Long sessions and cost stay under control, all opt-in:

alias ExAgent.{Compaction, CostGuard, Permissions, UsageLimits}
# Summarize old turns once the context grows (capability hook).
compaction = %Compaction.Capability{
compactor: Compaction.Summary,
opts: [threshold_tokens: 6000, keep_recent: 8, summarize: &MyApp.summarize/1]
}
# Per-tool admission control (allow/ask/deny with globs).
perms = Permissions.new!(rules: [{"*", :deny}, {"read", :allow}, {"bash", :ask}])
agent =
ExAgent.new(
model: chat_model,
capabilities: [compaction],
usage_limits: %UsageLimits{request_limit: 20, tool_calls_limit: 15,
max_budget_cents: 25, accounting: :estimated}
)
ExAgent.run(agent, "go",
permissions: perms,
approve: &MyApp.ask_human/1, # called on :ask
# Illustrative cents per 1K tokens; supply your provider's actual prices.
estimate_cost: CostGuard.estimator(%{input_per_1k_cents: 0.25, output_per_1k_cents: 1.0})
)

Compaction projects request context while retaining canonical history; it does not bound persisted history size. A caller's one-argument summarizer may perform IO outside the scope. Cost is estimated per request, retaining fractional cents; heterogeneous model pricing uses an arity-two (model, usage) estimator. Missing usage/price is unknown, not zero, and requests already in flight can exceed a retrospective token/cost threshold. Run options deadline (absolute monotonic milliseconds) and max_concurrent_requests control admission separately.

External tools (MCP)

Consume a stdio Model Context Protocol server's tools as plain ExAgent.Tools:

alias ExAgent.MCP.Client
{:ok, fs} =
Client.start_link(
command: "npx",
args: ["-y", "@modelcontextprotocol/server-filesystem", "./data"]
)
{:ok, tools} = Client.tools(fs) # [ExAgent.Tool.t(), ...]
agent = ExAgent.new(model: chat_model, tools: tools)

The client owns the stdio JSON-RPC connection (handshake, tools/list, tools/call, line buffering); transport exits and errors surface cleanly. Defaults are 128 pending requests and 8 MiB frames, configurable. A timeout/dead caller is cleaned up locally; that does not prove rollback of a remote effect. Streamable HTTP is opt-in and uses an application-owned HTTP1-only Finch pool. See MCP setup for protocol profiles, continuation binding and the qualified independent SDK scenarios.

Events & PubSub

Every layer emits versioned ExAgent.Event envelopes (distinct from :telemetry). Subscribe to drive a UI:

:ok = ExAgent.PubSub.subscribe({ExAgent.PubSub.Local, []}, ExAgent.Event.agent_topic("dm"))
receive do
{:exagent_event, %ExAgent.Event{type: :run_finished, payload: p}} ->
IO.puts("done: #{inspect(p)}")
end

ExAgent.PubSub is a behaviour: None (default, no-op), Local (Registry), Phoenix (delegates to Phoenix.PubSub dynamically — no hard dependency), or your own.

OpenTelemetry

Opt into native Erlang/Elixir OTel with an application-owned SDK and exporter:

tracing = ExAgent.Observability.OpenTelemetry.new()
agent = ExAgent.new(model: model, observability: tracing)

Spans cover runs, model requests, tools, delegation, compaction and checkpoints. Content is off by default; enabling it requires explicit redaction before attributes reach the exporter. Request usage and inclusive run totals are kept separate to avoid double counting. The optional bounded processor provides an asynchronous export route with observable saturation and failures.

For applications also tracing standalone ReqLLM calls, attach ExAgent.Observability.ReqLLM.attach/1 once at host startup instead of the stock ReqLLM bridge. It enriches the existing ExAgent Model span and preserves standalone tracing. ExAgent keeps usage, cost, status, privacy and span lifecycle; conflicting stock bridges reject before provider IO. The host maintains ReqLLM's tracking TTL.

The same bridge can opt into four metric instruments with bounded model labels. Long-lived applications can supervise the optional ExAgent.Observability.ReqLLM.Maintenance child to clean up expired tracking. Neither option installs a metric SDK or exporter for the application.

See observability for application configuration, context propagation, privacy, Langfuse/Opik configuration and export limits. The application supplies its backend credentials and chooses a separate metric destination when needed. Neither backend is a required dependency.

Models

Resolve stock specs with explicit options, or pass a custom Model struct:

{:ok, model} = ExAgent.Model.resolve("openai:gpt-4o",
api_key: System.fetch_env!("OPENAI_API_KEY"))
ExAgent.new(model: model)
ExAgent.new(model: %ExAgent.Models.Test{script: ["offline"]})

Strings alone resolve identity; they do not load credentials or enable tools/Chat. resolve/2 accepts public stock map/tuple/LLMDB specs and instance options; custom Model structs remain unchanged through resolve/1. Unknown options reject. OpenCode Go/Zen require explicit endpoints; stock zai: means ZAI, not the old Anthropic gateway alias. See the exact migration recipes. Bring your own model by implementing ExAgent.Model, without private ReqLLM APIs. Automatic retries and redirects are disabled; stream limits are documented on ExAgent.Models.ReqLLM. Incomplete/invalid terminal responses fail before effects. Usage is normalized, provider presence unknown, and costs estimated; strict metric limits reject this contract, while host request/tool limits remain exact admission counters.

Examples

Run any of them with mix run examples/<name>.exs (live ones need an API key in the environment).

Documentation

Start with the documentation home or follow the getting-started tutorial. Task guides cover tools/output, models/limits, runtime/events, durability/approvals, coordination and testing. Coding-agent integration notes map tasks to public APIs; ExDoc generates llms.txt and Markdown pages from the same sources.

Contributing

Bug reports and pull requests are welcome on GitHub.

For source contributions, follow AGENTS.md in the repository and run ./bin/check using the project-local tooling. Integration issues should include a minimal public-API example, package and Elixir/OTP versions, and a sanitized error. Keep provider credentials and application data out of reports.

EXAGENT_OFFLINE=1 MIX_ENV=test mix compile --warnings-as-errors
EXAGENT_OFFLINE=1 MIX_ENV=test mix test --warnings-as-errors
EXAGENT_OFFLINE=1 MIX_ENV=test mix run examples/demo.exs

EXAGENT_OFFLINE=1 skips even the Postgres bootstrap. Real-provider tests are tagged :integration and excluded by default; enable them only for an explicitly scoped backend check. Without offline mode, Postgres tests attempt their database and auto-skip if unavailable. Mock/stdio/local-HTTP tests do not prove acceptance by external providers or a durable database.

License

Copyright (c) 2025 kukapu

Licensed under the MIT License — see LICENSE.