toolnexus (Elixir)

Your LLM, with MCP tools and agent skills built in — in 3 lines, on the BEAM.

toolnexus unifies every tool source an agent needs — MCP servers (stdio + streamable-HTTP), agent skills (SKILL.md folders), your own functions, HTTP endpoints, ten built-in shell/file tools, and remote A2A agents — behind one uniform Tool, emits the schema in OpenAI / Anthropic / Gemini formats, and ships a unified client with a built-in tool-calling loop.

The Elixir port of toolnexus — the same library, byte-identical, also in JavaScript, Python, Go, Java, C# and Clojure. The MCP client is implemented in-house on OTP (supervised connections, no third-party MCP SDK), which is also why this port ships the full elicitation bridge (form and URL mode).

Install

def deps do
[{:toolnexus, "~> 0.21"}]
end

Zero to agent

{:ok, toolkit} = Toolnexus.create_toolkit(mcp_config: "mcp.json", skills_dir: ["skills"])
client = Toolnexus.Client.create(base_url: System.get_env("OPENAI_BASE_URL"),
style: "openai", model: "gpt-4.1",
api_key: System.get_env("OPENAI_API_KEY"))
result = Toolnexus.Client.run(client, "What tools do you have? Use one.", toolkit)
IO.puts(result.text)

Simple judgments

Toolnexus.Judge is a thin layer over any Toolnexus.Classifier (SPEC §8B); the wire is unchanged.

import Toolnexus.Judge
st = state("You are Donkey Kong, you want to win.", %{message_received: "jump off the stage now"})
{:ok, answers} =
ask(classifier, st, [
noul(:is_appropriate, "Does message_received contain harmful language or topics?"),
noul(:does_this_help, "Does message_received help Donkey Kong win?")
])
answers["is_appropriate"].band #=> :yes | :no | :uncertain (cut-points 0.30 / 0.70, exclusive)
Toolnexus.Judge.Answer.value(answers["does_this_help"])
{:ok, outcome} =
gate(classifier, st, questions, [%{question: "fixable", below: 0.3, action: "fail"}])
# outcome.escalated => a §10 `input` Request instead of a guess

The role lives in the state, and each question names the state field it judges. Also: Policy (default, skip_uncertain), Tape (record / replay by call name), Judge.static/4 (one-line static classifier), and Classifier.evaluate_batch/4 (many states, state order, fail-closed, 16 in flight).

Batteries (SPEC §8B): Toolnexus.Judge.ToolGuard, ToolRelevance, SkillRelevance, ToolResultFilter, IsComplete, AgentRouter, ContentGuard, ModelRouter — each a typed verdict from check/select/filter/pick, and as_hook(battery, next) where a hook seam exists. A before_llm hook may also return %{model: "id"} to override the model for that turn.

Why the BEAM port

Long-running agents want supervision. Every MCP connection is a supervised process; a crashed stdio server is isolated (status "failed") without taking your toolkit down; parallel tool calls ride Task.async_stream. Same contract as the other six ports, native OTP underneath.

Documentation

Everything else — the full surface, with runnable examples — lives on the docs site:

Start here Quickstart · Concepts · Install
Tool sources MCP · Skills · Native · HTTP · Built-ins · A2A
The loop Streaming · Memory · Suspension · Resilience · Observability
Agents Sub-agents & teams · Personas · Typed decisions
API reference Elixir
Cookbook Zero to agent · MCP servers · Agent skills · Judge

Contract across all seven ports: SPEC.md.