toolnexus (Elixir)
Your LLM, with MCP tools and agent skills built in — in 3 lines, on the BEAM.
toolnexus unifies every tool source an agent needs — MCP servers (stdio + streamable-HTTP),
agent skills (SKILL.md folders), your own functions, HTTP endpoints, ten built-in shell/file
tools, and remote A2A agents — behind one uniform Tool, emits the schema in OpenAI /
Anthropic / Gemini formats, and ships a unified client with a built-in tool-calling loop.
The Elixir port of toolnexus — the same library, byte-identical, also in JavaScript, Python, Go, Java, C# and Clojure. The MCP client is implemented in-house on OTP (supervised connections, no third-party MCP SDK), which is also why this port ships the full elicitation bridge (form and URL mode).
Install
def deps do
[{:toolnexus, "~> 0.21"}]
end
Zero to agent
{:ok, toolkit} = Toolnexus.create_toolkit(mcp_config: "mcp.json", skills_dir: ["skills"])
client = Toolnexus.Client.create(base_url: System.get_env("OPENAI_BASE_URL"),
style: "openai", model: "gpt-4.1",
api_key: System.get_env("OPENAI_API_KEY"))
result = Toolnexus.Client.run(client, "What tools do you have? Use one.", toolkit)
IO.puts(result.text)
mcp.jsonis the standard Claude-desktop-style config (mcpServers/servers/mcpkeys all accepted).skills/is a folder of**/SKILL.mdfiles, loaded on demand through oneskilltool.- Remote MCP
headersvalues expand${ENV_VAR}at call time and are never logged.
Simple judgments
Toolnexus.Judge is a thin layer over any Toolnexus.Classifier (SPEC §8B); the wire is unchanged.
import Toolnexus.Judge
st = state("You are Donkey Kong, you want to win.", %{message_received: "jump off the stage now"})
{:ok, answers} =
ask(classifier, st, [
noul(:is_appropriate, "Does message_received contain harmful language or topics?"),
noul(:does_this_help, "Does message_received help Donkey Kong win?")
])
answers["is_appropriate"].band #=> :yes | :no | :uncertain (cut-points 0.30 / 0.70, exclusive)
Toolnexus.Judge.Answer.value(answers["does_this_help"])
{:ok, outcome} =
gate(classifier, st, questions, [%{question: "fixable", below: 0.3, action: "fail"}])
# outcome.escalated => a §10 `input` Request instead of a guess
The role lives in the state, and each question names the state field it judges. Also: Policy
(default, skip_uncertain), Tape (record / replay by call name), Judge.static/4 (one-line
static classifier), and Classifier.evaluate_batch/4 (many states, state order, fail-closed, 16 in flight).
Batteries (SPEC §8B): Toolnexus.Judge.ToolGuard, ToolRelevance, SkillRelevance,
ToolResultFilter, IsComplete, AgentRouter, ContentGuard, ModelRouter — each a typed
verdict from check/select/filter/pick, and as_hook(battery, next) where a hook seam
exists. A before_llm hook may also return %{model: "id"} to override the model for that turn.
Why the BEAM port
Long-running agents want supervision. Every MCP connection is a supervised process; a crashed
stdio server is isolated (status "failed") without taking your toolkit down; parallel tool calls
ride Task.async_stream. Same contract as the other six ports, native OTP underneath.
Documentation
Everything else — the full surface, with runnable examples — lives on the docs site:
| Start here | Quickstart · Concepts · Install |
| Tool sources | MCP · Skills · Native · HTTP · Built-ins · A2A |
| The loop | Streaming · Memory · Suspension · Resilience · Observability |
| Agents | Sub-agents & teams · Personas · Typed decisions |
| API reference | Elixir |
| Cookbook | Zero to agent · MCP servers · Agent skills · Judge |
Contract across all seven ports: SPEC.md.