ExAtlas

Hex.pm Docs License

A composable, pluggable Elixir SDK for infrastructure management. Two concerns under one roof:

  1. GPU / CPU compute across cloud providers. Spawn pods, run serverless inference, orchestrate transient per-user GPU sessions. Swap providers by changing one option.
  2. Fly.io platform operations. First-class deploys, log streaming, and token lifecycle — independent of the compute pipeline. See ExAtlas.Fly and the Fly guide.

Table of contents


Installation

The one-liner — uses the Igniter installer to add the dep, write sensible config, and create storage directories:

mix igniter.install ex_atlas

Or add manually to mix.exs:

def deps do
[
{:ex_atlas, "~> 0.2"}
]
end

…then run mix ex_atlas.install once to wire config defaults, or configure things yourself (see Configuration).

For the optional orchestrator + LiveDashboard features, also include:

{:phoenix_pubsub, "~> 2.1"}, # PubSub broadcasts from the orchestrator
{:phoenix_live_dashboard, "~> 0.8"} # ExAtlas.LiveDashboard.ComputePage tab

ExAtlas declares both as optional: true, so they are not pulled into pure library consumers.

Upgrading

To upgrade atlas and run any version-specific migrations:

mix deps.update atlas
mix ex_atlas.upgrade

The upgrade task is idempotent and runs only the steps needed between your previous and current atlas version.

Architecture at a glance

┌───────────────────────────────────────────────────────────────────────┐
│ ExAtlas (top-level provider-agnostic API) │
│ ExAtlas.spawn_compute/1 · run_job/2 · stream_job/1 · terminate/1 │
└───────────────────────────┬───────────────────────────────────────────┘
│
┌───────────────▼───────────────┐ ┌───────────────────┐
│ ExAtlas.Provider (behaviour) │◄───│ ExAtlas.Spec.* │
└───────────────┬───────────────┘ │ normalized structs│
│ └───────────────────┘
┌─────────┬─────────────┼──────────────┬─────────────┐
│ │ │ │ │
┌──▼───┐ ┌──▼───┐ ┌───────▼────────┐ ┌──▼─────┐ ┌──────▼──────┐
│RunPod│ │ Fly │ │ Lambda Labs │ │ Vast │ │ Mock (test)│
│ v0.1 │ │ v0.2 │ │ v0.2 │ │ v0.3 │ │ v0.1 │
└──────┘ └──────┘ └────────────────┘ └────────┘ └─────────────┘
┌───────────────────────────────────────────────────────────────────────┐
│ ExAtlas.Orchestrator (opt-in supervision tree) │
│ ComputeServer (GenServer/resource) · Registry · DynamicSupervisor │
│ · Reaper · PubSub events │
└───────────────────────────────────────────────────────────────────────┘
┌───────────────────────────────────────────────────────────────────────┐
│ ExAtlas.Auth │
│ Token (bearer mint/verify) · SignedUrl (S3-style HMAC) │
└───────────────────────────────────────────────────────────────────────┘
┌───────────────────────────────────────────────────────────────────────┐
│ ExAtlas.LiveDashboard.ComputePage │
│ Live-refreshing table · per-row Touch/Stop/Terminate │
└───────────────────────────────────────────────────────────────────────┘

Quick start — Fly.io platform ops

ExAtlas gives you a clean Elixir API over fly deploy, the Fly Machines log API, and Fly token lifecycle. Works with or without Phoenix.

Discover apps

ExAtlas.Fly.discover_apps("/path/to/project")
# => [{"my-api", "/path/to/project"}, {"my-web", "/path/to/project/web"}]

Tail logs

ExAtlas.Fly.subscribe_logs("my-api", "/path/to/project")
# In the subscriber:
def handle_info({:ex_atlas_fly_logs, "my-api", entries}, state) do
# entries :: [ExAtlas.Fly.Logs.LogEntry.t()]
...
end

A single streamer runs per app regardless of subscriber count, and stops once all subscribers disconnect. Automatic 401 retry is built in.

Stream a deploy

ExAtlas.Fly.subscribe_deploy(ticket_id)
Task.start(fn ->
ExAtlas.Fly.stream_deploy(project_path, "web", ticket_id)
end)
def handle_info({:ex_atlas_fly_deploy, ^ticket_id, line}, state) do
...
end

Deploys are guarded by a 5 min activity timer (resets on output) and a 30 min absolute cap.

Tokens

ExAtlas.Fly.Tokens resolves tokens via ETS → DETS (durable) → ~/.fly/config.yml → fly tokens create readonly → manual override. You usually don't call it directly — the log client uses it transparently — but you can:

{:ok, token} = ExAtlas.Fly.Tokens.get("my-api")
ExAtlas.Fly.Tokens.invalidate("my-api")
ExAtlas.Fly.Tokens.set_manual("my-api", "fo1_...")

Full docs: Fly guide.

Quick start — transient per-user GPU pod

The motivating use case: a Fly.io-hosted Phoenix app spawns a RunPod GPU per user, hands the browser a preshared key, the browser runs real-time video inference directly against the pod, and ExAtlas reaps the pod when the session ends or goes idle.

# config/config.exs
config :ex_atlas, default_provider: :runpod
config :ex_atlas, :runpod, api_key: System.get_env("RUNPOD_API_KEY")
config :ex_atlas, start_orchestrator: true
# LiveView.mount/3
{:ok, pid, compute} =
ExAtlas.Orchestrator.spawn(
gpu: :h100,
image: "ghcr.io/me/my-inference-server:latest",
ports: [{8000, :http}],
auth: :bearer,
user_id: socket.assigns.current_user.id,
idle_ttl_ms: 15 * 60_000,
name: "atlas-" <> to_string(socket.assigns.current_user.id)
)
Phoenix.PubSub.subscribe(ExAtlas.PubSub, "compute:" <> compute.id)
assign(socket,
inference_url: hd(compute.ports).url, # https://<pod-id>-8000.proxy.runpod.net
inference_token: compute.auth.token # handed straight to the browser
)

Inside the inference server running in the pod:

# Any request from the browser must carry the preshared key.
def authenticated?(conn) do
preshared = System.fetch_env!("ATLAS_PRESHARED_KEY")
case Plug.Conn.get_req_header(conn, "authorization") do
["Bearer " <> token] -> Plug.Crypto.secure_compare(token, preshared)
_ -> false
end
end

Heartbeat while the browser is active:

ExAtlas.Orchestrator.touch(compute.id)

When the user leaves, or after idle_ttl_ms with no heartbeat, the ComputeServer shuts down and terminates the upstream pod automatically. You can also terminate manually:

:ok = ExAtlas.Orchestrator.stop_tracked(compute.id)

Waiting until a pod is usable

spawn_compute/1 and Orchestrator.spawn/1 return the moment the provider accepts the rental — 30–90 seconds before the container answers anything. A LiveView should subscribe and render "starting…"; everything else should block on await_ready/2.

# Tracked. Rides the ComputeServer's existing status poll, so it adds no
# requests to your provider, and returns immediately if it is already up.
case ExAtlas.Orchestrator.await_ready(compute.id, timeout_ms: 120_000) do
{:ok, ready} -> hd(ready.ports).url
{:error, {:dead, reason, _}} -> {:error, reason} # :failed | :vanished | :preempted | …
{:error, {:timeout, last}} -> wait_longer_or_give_up(last)
end
# Untracked — a bare spawn, a script, a mix task. Polls get_compute/2 itself.
ExAtlas.await_ready(id, provider: :runpod, timeout_ms: 120_000, poll_interval_ms: 2_000)

The tracked wait runs in the calling process and consumes that pod's {:atlas_compute, id, _} messages, so run it in a task rather than in a process that is itself subscribed. See the transient pods guide.

Quick start — batch task on a GPU pod

Unattended work that runs to completion, rather than a session a user holds open:

{:ok, _pid, compute} =
ExAtlas.Orchestrator.run_task(
provider: :runpod,
gpu: :rtx_4090,
image: "ghcr.io/acme/trainer:latest",
command: ["/app/train.sh", "--epochs", "3"],
name: "atlas-task-42",
max_runtime_ms: :timer.minutes(90)
)
Phoenix.PubSub.subscribe(ExAtlas.PubSub, "compute:" <> compute.id)
def handle_info({:atlas_compute, _id, {:task, :completed}}, socket) do
{:noreply, put_flash(socket, :info, "Training finished")}
end
def handle_info({:atlas_compute, _id, {:task, :timed_out}}, socket) do
{:noreply, put_flash(socket, :error, "Training hit its 90-minute cap")}
end
def handle_info({:atlas_compute, _id, {:task, {:failed, reason}}}, socket) do
{:noreply, put_flash(socket, :error, "Training failed: #{reason}")}
end

No heartbeats, no idle TTL: the pod ends when the command ends, when the wall-clock deadline fires, or when it dies — and it is always destroyed. See run_task/1 for why exit detection needs the container's cooperation, and why :completed does not mean "succeeded" — unless you add a callback.

Pod callbacks — progress, logs, exit codes

RunPod has no API for pod logs and reports no container state, so for unattended work the only party that knows what is happening inside the container is the container. ExAtlas.Callback is the supported way for it to say so: progress reports, streamed log lines, and an exit code declared before the pod goes.

# config/runtime.exs
config :ex_atlas, :callback,
secret: System.fetch_env!("ATLAS_CALLBACK_SECRET"), # >= 32 bytes, same on every node
base_url: "https://app.example.com/atlas/cb"
# lib/my_app_web/router.ex — outside :browser, outside your auth, and NOT
# through Plug.Parsers (its 8 MB default is the memory vector).
scope "/atlas" do
forward "/cb", ExAtlas.Callback.Plug
end
{:ok, _pid, compute} =
ExAtlas.Orchestrator.run_task(
provider: :runpod,
gpu: :rtx_4090,
image: "ghcr.io/acme/trainer:latest",
command: ["/app/train.sh"],
callback: "https://app.example.com/atlas/cb"
)
def handle_info({:atlas_compute, _id, {:progress, payload}}, socket),
do: {:noreply, assign(socket, pct: payload["pct"])}
def handle_info({:atlas_compute, _id, {:log, payload}}, socket),
do: {:noreply, stream(socket, :lines, payload["lines"])}
def handle_info({:atlas_compute, _id, {:task, {:failed, {:exit_code, n}}}}, socket),
do: {:noreply, put_flash(socket, :error, "Training exited #{n}")}

The container is handed ATLAS_CALLBACK_URL, ATLAS_CALLBACK_TOKEN and ATLAS_TASK_ID; the self-termination wrapper already POSTs /finish with the exit code from its trap, so only progress and logs need code in your image.

Three things worth knowing up front:

Not using Plug? ExAtlas.Callback is framework-free — verify/1 and ingest/3 are all a hand-rolled controller needs, and :plug stays an optional dependency. Full details, container-side examples, rate limits and the security model are in the pod callbacks guide.

Quick start — serverless inference

{:ok, job} =
ExAtlas.run_job(
provider: :runpod,
endpoint: "abc123",
input: %{prompt: "a beautiful sunset"},
mode: :async
)
{:ok, done} = ExAtlas.get_job(job.id, provider: :runpod, endpoint: "abc123")
done.output
# Synchronous with a hard timeout (wrapped in Task.async + Task.yield internally)
{:ok, done} =
ExAtlas.run_job(
provider: :runpod,
endpoint: "abc123",
input: %{prompt: "a beautiful sunset"},
mode: :sync,
timeout_ms: 60_000
)
# Stream partial output
ExAtlas.stream_job(job.id, provider: :runpod, endpoint: "abc123")
|> Enum.each(&IO.inspect/1)

Swapping providers

# Today
ExAtlas.spawn_compute(provider: :runpod, gpu: :h100, image: "...")
# v0.2
ExAtlas.spawn_compute(provider: :fly, gpu: :a100_80g, image: "...")
ExAtlas.spawn_compute(provider: :lambda_labs, gpu: :h100, image: "...")
# v0.3
ExAtlas.spawn_compute(provider: :vast, gpu: :rtx_4090, image: "...")
# Your in-house cloud, today:
ExAtlas.spawn_compute(provider: MyCompany.Cloud.Provider, gpu: :h100, image: "...")

All built-in and user-defined providers implement ExAtlas.Provider.

Configuration

# config/config.exs
# Provider resolution: per-call :provider option > :default_provider > raise
config :ex_atlas, default_provider: :runpod
# API keys: per-call :api_key > :ex_atlas / :<provider> config > env var
config :ex_atlas, :runpod, api_key: System.get_env("RUNPOD_API_KEY")
config :ex_atlas, :fly, api_key: System.get_env("FLY_API_TOKEN")
config :ex_atlas, :lambda_labs, api_key: System.get_env("LAMBDA_LABS_API_KEY")
config :ex_atlas, :vast, api_key: System.get_env("VAST_API_KEY")
# Start the orchestrator (Registry + Task.Supervisor + DynamicSupervisor +
# PubSub + Reaper).
# When false (default), ExAtlas boots no processes.
config :ex_atlas, start_orchestrator: true
# Reaper: periodic orphan reconciliation and idle-TTL enforcement.
config :ex_atlas, :orchestrator,
reap_interval_ms: 60_000,
reap_providers: [:runpod],
reap_name_prefix: "atlas-", # safety switch: only reap resources ExAtlas spawned
reap_grace_ms: 60_000, # spare resources too young to have a tracker yet
# Required on every machine when more than one shares a provider account.
# Unique per machine and stable across its restarts — see "More than one node".
reap_owner: System.get_env("FLY_MACHINE_ID"),
# Durable tracking, so a deploy does not orphan (and then reap) a running
# task. Defaults to the DETS store; set `false` to disable persistence.
tracking_store: ExAtlas.Orchestrator.TrackingStore.Dets,
storage_path: "/data/ex_atlas", # a MOUNTED volume — see "Surviving a deploy"
scrub_keys: [] # extra opt keys to keep off disk

Default environment variable names used when nothing else is set:

Provider Env var
:runpod RUNPOD_API_KEY
:fly FLY_API_TOKEN
:lambda_labs LAMBDA_LABS_API_KEY
:vast VAST_API_KEY

Providers

Provider Module Version shipped Capabilities
:runpod ExAtlas.Providers.RunPod v0.1 :serverless, :network_volumes, :http_proxy, :raw_tcp, :symmetric_ports, :webhooks, :global_networking
:fly ExAtlas.Providers.Fly v0.2 (stub) :http_proxy, :raw_tcp, :global_networking
:lambda_labs ExAtlas.Providers.LambdaLabs v0.2 (stub) :raw_tcp
:vast ExAtlas.Providers.Vast v0.3 (stub) :spot, :raw_tcp
:mock ExAtlas.Providers.Mock v0.1 (tests) :spot, :serverless, :network_volumes, :http_proxy, :raw_tcp, :webhooks

Stub modules return {:error, %ExAtlas.Error{kind: :unsupported}} from every non-capabilities/0 callback so the name is reserved and callers get a clear error — no FunctionClauseErrors.

Canonical GPU atoms

ExAtlas refers to GPUs by stable atoms. ExAtlas.Spec.GpuCatalog maps each atom to each provider's native identifier.

Canonical RunPod Lambda Labs Fly.io Vast.ai
:h200 "NVIDIA H200" — — "H200"
:h100 "NVIDIA H100 80GB HBM3" "gpu_1x_h100_pcie" — "H100"
:a100_80g "NVIDIA A100 80GB PCIe" "gpu_1x_a100_sxm4_80gb" "a100-80gb" "A100_80GB"
:a100_40g "NVIDIA A100-SXM4-40GB" "gpu_1x_a100_sxm4" "a100-pcie-40gb" "A100"
:l40s "NVIDIA L40S" — "l40s" —
:l4 "NVIDIA L4" — — —
:a6000 "NVIDIA RTX A6000" "gpu_1x_a6000" — "RTX_A6000"
:rtx_4090 "NVIDIA GeForce RTX 4090" — — "RTX_4090"
:rtx_3090 "NVIDIA GeForce RTX 3090" — — "RTX_3090"
:mi300x "AMD Instinct MI300X OAM" — — —

See ExAtlas.Spec.GpuCatalog for the full mapping.

The ExAtlas.Provider behaviour

Every provider implements one callback per operation. See ExAtlas.Provider for the full contract.

Callback Purpose
spawn_compute/2 Provision a GPU/CPU resource
get_compute/2 Fetch current status
list_compute/2 List with optional filters
stop/2 / start/2 Pause / resume
terminate/2 Destroy
run_job/2 Submit a serverless job
get_job/2 / cancel_job/2 Job control
stream_job/2 Stream partial outputs
capabilities/0 Declare supported features
list_gpu_types/1 Catalog + pricing

Callers can check ExAtlas.capabilities(:runpod) before relying on an optional feature:

if :serverless in ExAtlas.capabilities(provider) do
ExAtlas.run_job(provider: provider, endpoint: "...", input: %{...})
end

Capability atoms

Atom Meaning
:spot Interruptible/spot instances
:serverless run_job/2 and friends
:network_volumes Attach persistent volumes
:http_proxy Provider terminates TLS on a *.proxy.* hostname
:raw_tcp Public IP + mapped TCP ports
:symmetric_ports internal == external port guarantee
:webhooks Push completion callbacks
:global_networking Private networking across datacenters
:self_terminate Honors :self_terminate — wraps :command so the resource ends itself

Normalized specs (ExAtlas.Spec.*)

Requests and responses flow through normalized structs so callers don't have to know each provider's native shape.

Every spec struct has a :raw field preserving the provider's native response for callers who need fields ExAtlas hasn't yet normalized.

The :provider_opts field on request structs is the escape hatch for provider-specific options ExAtlas doesn't model — values are stringified and merged into the outgoing REST body.

Auth primitives

ExAtlas.Auth.Token and ExAtlas.Auth.SignedUrl are exposed directly if you want them without the rest of the orchestration layer.

Bearer tokens

mint = ExAtlas.Auth.Token.mint()
# %{
# token: "kX9fP...", # hand to client once
# hash: "4c1...", # persist this
# header: "Authorization: Bearer kX9fP...",
# env: %{"ATLAS_PRESHARED_KEY" => "kX9fP..."} # inject into the pod
# }
ExAtlas.Auth.Token.valid?(candidate, mint.hash)

When you pass auth: :bearer to spawn_compute/1, ExAtlas mints a token, adds it to the pod's env as ATLAS_PRESHARED_KEY, and returns the handle in compute.auth — all in one round-trip.

Callback tokens

ExAtlas.Callback.Token is the inbound direction — the credential a pod presents when calling back into your app. It is a stateless signed token (Plug.Crypto.sign/4 over %{task_id, kinds}, verified in constant time), bound to a task id rather than a compute id, and deliberately not the same credential as ATLAS_PRESHARED_KEY: that one is handed to a browser, and a browser-held secret must never also authorize writing into your orchestrator. See the pod callbacks guide.

S3-style signed URLs

For <video src>, <img src>, or any client that can't set request headers:

url =
ExAtlas.Auth.SignedUrl.sign(
"https://pod-id-8000.proxy.runpod.net/stream",
secret: signing_secret,
expires_in: 3600
)
:ok = ExAtlas.Auth.SignedUrl.verify(url, secret: signing_secret)

The signature covers the path + canonicalized query + expiry with HMAC-SHA256; verification uses constant-time comparison.

Orchestrator — lifecycle, events, reaper

ExAtlas.Orchestrator.spawn/1

Spawns the resource via the provider, then starts an ExAtlas.Orchestrator.ComputeServer under ExAtlas.Orchestrator.ComputeSupervisor that:

  1. Registers itself in ExAtlas.Orchestrator.ComputeRegistry under {:compute, id}.
  2. Traps exits — its terminate/2 always calls ExAtlas.terminate/2 on the upstream provider, whether the supervisor shuts it down or it exits on an idle timeout.
  3. Tracks :last_activity_ms and compares against :idle_ttl_ms on every heartbeat tick. If idle, the server stops normally and the upstream resource is destroyed.
  4. Polls the provider every :status_poll_ms and reports what it finds, so a resource that died on the cloud's side is noticed rather than assumed away.

Upstream status polling

The heartbeat answers "does anyone still want this?"; it cannot answer "is it still there?". A pod can die without you asking — a host fails, an image crash-loops, spot capacity is reclaimed — and until the tracker asks the provider, subscribers go on believing the session is healthy and the meter goes on running.

So the ComputeServer runs a second, independent clock. Every :status_poll_ms (default 60s) it calls get_compute/2 and broadcasts the result: a status change while the resource is alive, or the cause of death followed by a normal shutdown.

ExAtlas.Orchestrator.spawn(
gpu: :h100,
image: "ghcr.io/me/trainer:latest",
spot: true,
status_poll_ms: 30_000, # false to disable polling entirely
on_failure: {:respawn, 3} # replace preempted pods, up to 3 times
)
Option Default Meaning
:status_poll_ms 60_000 Poll interval; false disables polling
:on_failure :stop Or {:respawn, max_attempts} for spot workloads

The interval is separate from :heartbeat_ms on purpose: one is paced by your users' activity, the other by what your provider's API tolerates. RunPod's management API publishes no rate limits at all, which is why the default is unhurried and why polls back off exponentially (with jitter, so a fleet of trackers doesn't poll in lockstep) while the provider is failing.

A failed poll is not a death. Only a 404 means the resource is gone. A 500, a rate limit, a socket error, a malformed body, even a bad API key come back as {:poll_failed, error} — the poller backs off and the session continues. Tearing down a live GPU because one request failed is far more expensive than noticing a death a minute late.

Preemption is inferred, not reported. No provider publishes a "you were outbid" signal, and Runpod no longer sells spot pods at all (spot: true on :runpod returns {:error, %ExAtlas.Error{kind: :unsupported}}). So a resource on a provider that still sells spot, spawned with spot: true, that stops, is terminated, or vanishes without you asking is reported as {:status, :preempted}. On-demand resources keep the literal reason (:stopped, :terminated, :vanished).

With on_failure: {:respawn, n} a preempted resource is replaced from the same opts rather than ending the session: the tracker terminates the old resource if the provider still has it, re-keys itself under the new id, and emits {:respawned, new_id} on the old topic. The event carries the id and nothing else — the replacement's URL and bearer token come from ExAtlas.Orchestrator.info(new_id), which is readable by the time the event lands, and a live credential has no business on a PubSub topic. Only preemption is retried — a stopped or terminated resource was ended by someone, and an image that :failed here will fail on the next host too.

ExAtlas.Orchestrator.UpstreamStatus is the primitive underneath, usable on its own if you want the classification without a tracker process:

ExAtlas.Orchestrator.UpstreamStatus.observe(pod_id, provider: :runpod, spot: true)
# {:alive, %ExAtlas.Spec.Compute{}}
# {:dead, :preempted, nil}
# {:poll_failed, %ExAtlas.Error{}}

PubSub events

Every state change is broadcast over ExAtlas.PubSub on the topic "compute:<id>" as {:atlas_compute, id, event}:

Event Emitted when
{:status, status} ComputeServer starts (whatever the resource's status is), and on every upstream status change
{:heartbeat, monotonic_ms} Heartbeat tick (no idle timeout)
{:status, :preempted} A spot: true resource was reclaimed
{:status, :failed | :stopped | :vanished} A poll found the resource dead
{:poll_failed, error} A status poll couldn't reach the provider, or blew up trying
{:respawned, new_id} Preempted resource replaced (sent on the old id)
{:respawn_failed, {reason, error}} Replacement couldn't be spawned
{:task, outcome} A mode: :task session ended: :completed, :timed_out, {:failed, reason}
{:terminating, reason} Server is about to shut down
{:status, :terminated} Upstream provider confirmed termination, or had nothing left to terminate
{:terminate_failed, error} Upstream terminate call returned an error

{:terminating, _} followed by {:status, :terminated} is the end-of-session signal; no individual status is.

Subscribe in a LiveView:

Phoenix.PubSub.subscribe(ExAtlas.PubSub, "compute:" <> compute.id)
def handle_info({:atlas_compute, _id, {:status, :terminated}}, socket) do
{:noreply, put_flash(socket, :info, "Session ended")}
end

ExAtlas.Orchestrator.run_task/1 — run a container to completion

The third compute shape, alongside interactive per-user pods and serverless jobs: run this image with this command until it exits, then tell me the outcome and stop the meter.

{:ok, pid, compute} =
ExAtlas.Orchestrator.run_task(
provider: :runpod,
gpu: :rtx_4090,
image: "ghcr.io/acme/trainer:latest",
command: ["/app/train.sh", "--epochs", "3"],
name: "atlas-task-42",
max_runtime_ms: :timer.minutes(90)
)
Phoenix.PubSub.subscribe(ExAtlas.PubSub, "compute:" <> compute.id)
Option Default Meaning
:command nil Overrides the image's start command (cmd on Runpod)
:self_terminate true Wrap :command so the resource destroys itself on exit
:max_runtime_ms 60 min Wall-clock deadline from spawn; then DELETE
:ready_timeout_ms 15 min Fail as :never_ready if still provisioning when it fires

Task mode is the same ComputeServer, so everything above still applies — the status poll, the backoff, the guaranteed DELETE on teardown, the respawn option. Three things differ:

Why exit detection needs the container's help

Runpod's REST v2 API exposes no exit code and no restart count: the Pod body has a runtime object (uptime, GPUs, ports), but status stays RUNNING when your command exits, and Runpod restarts the container. A live probe (2026-09-30) read RUNNING 180 seconds after a 30-second command, which had started 6 times. The GPU stays reserved, and no amount of polling can tell that the work is done.

The only party that knows is the container. With self_terminate: true ex_atlas wraps your command in a shell that deletes the pod when it ends:

atlas_self_terminate() {
curl -sS -X DELETE -H "Authorization: Bearer $RUNPOD_API_KEY" \
"https://api.runpod.io/v2/pods/$RUNPOD_POD_ID"
}
trap atlas_self_terminate EXIT INT TERM
/app/train.sh --epochs 3

RUNPOD_POD_ID and the pod-scoped RUNPOD_API_KEY are injected by RunPod, so no secret of yours travels to the pod, and trap … EXIT fires on a crash and on a signal as well as on a clean finish. The orchestrator sees the resulting 404 and reports {:task, :completed}.

Pass self_terminate: false for an image with no shell or no curl, or when you want the resource kept up for inspection. Such a task will always end at :max_runtime_ms and report :timed_out, which is honest.

Self-termination and the deadline are both required — they cover disjoint failures. Self-termination is the only source of a normal-exit signal; the deadline is the only cover for a SIGKILL or OOM kill that runs no cleanup, a hung process, an image that never pulled, and self_terminate: false.

:completed does not mean "succeeded"

It means the container ended and the resource is gone. trap … EXIT fires on a crash too, so success and failure arrive as the same 404, and the exit code dies with the pod. If you need the difference, have the container report it before it exits — a status file on a network volume, a call to your own webhook.

Spot tasks

With spot: true a resource that vanished is reported as :preempted, and a self-terminating container also makes it vanish — so on spot capacity the two are indistinguishable through the API. on_failure: {:respawn, n} can therefore re-run a task that had actually finished. Use it only where re-running is harmless (checkpoint-resuming training, which is what the option was added for); the carried deadline bounds the total spend either way.

Reaper

ExAtlas.Orchestrator.Reaper runs periodically (configurable, default 60s) and:

  1. Lists each configured provider's running resources.
  2. Compares against the resources tracked by the local ComputeRegistry and recorded in the tracking store (see "Surviving a deploy").
  3. Terminates any orphan whose :name starts with :reap_name_prefix (default "atlas-") and that is older than :reap_grace_ms.

The prefix is a safety switch so ExAtlas never touches pods created by other tools on the same cloud account. Set it to "" to disable.

More than one node

Every machine lists the whole provider account, and knows only its own pods. Without an owner, machine B deletes machine A's live pods once they pass the grace window. Every deployment with more than one machine on one provider account must set :reap_owner on every machine, unique per machine and the same across that machine's restarts:

# config/runtime.exs
config :ex_atlas, :orchestrator, reap_owner: System.get_env("FLY_MACHINE_ID")

Use a-z and 0-9, 1 to 32 characters. On Fly, FLY_MACHINE_ID fits. Do not use node(): the Phoenix Fly template puts the image ref in RELEASE_NODE, so it changes on every deploy.

With an owner set:

The Reaper reaps nothing and logs an error when the owner is invalid, when a machine with no owner is connected to other nodes, or when a connected node reports the same owner. Those checks see connected nodes only:

Upgrading a cluster from v0.7.0 takes two deploys, because a v0.7.0 node running beside a v0.8.0 node deletes the new node's pods:

  1. Deploy v0.7.0 with reap_providers: [], so no node reaps.
  2. Deploy v0.8.0 with :reap_owner set and reap_providers restored.

Pods named by v0.7.0 carry no owner after the upgrade and are left alone. Their trackers still end them: adopted tasks by deadline, live sessions by idle TTL.

The Reaper and the status poller look in opposite directions: the Reaper asks "is anything running that nothing is tracking?" (provider → local), while each tracker asks "is the thing I track still alive?" (local → provider). Between them, a resource can neither outlive its tracker nor be believed alive after it dies.

They do meet in one place. A resource is created upstream before its tracker is registered — by spawn/1, and again by every respawn — so for the duration of that provider call it looks exactly like an orphan. The Reaper therefore leaves resources younger than :reap_grace_ms (default: one reap interval) alone, and only resources reporting no created_at at all skip the grace. Shorten the grace and you shorten how long a genuine orphan bills; shorten it below your provider's worst-case spawn latency and the Reaper starts killing brand-new resources.

Surviving a deploy: the tracking store

The Registry is in memory and the trackers are restart: :temporary, so a deploy empties them while your pods keep running. Every one of those pods is then old, prefix-matching and untracked — an orphan by the rules above — and the Reaper terminates it on the first tick. :reap_grace_ms cannot save it: the window is keyed off created_at, and a pod that has been training for three hours is not young.

ExAtlas.Orchestrator.TrackingStore is the fix. Opt a task in with persist: true:

ExAtlas.Orchestrator.run_task(
provider: :runpod,
gpu: :h100,
image: "ghcr.io/acme/trainer:latest",
command: ["/app/train.sh"],
name: "atlas-train-#{run.id}",
max_runtime_ms: :timer.hours(6),
persist: true
)

At the next boot ExAtlas.Orchestrator.Adopter reads the store, asks the provider whether each id still exists, and starts a tracker for the ones that do. The Reaper waits for it: with a store configured it reaps nothing until adoption has settled, and if the store could not be read it reaps nothing for the whole boot — a node that cannot tell which pods are its own must never issue a DELETE.

What an adopted task keeps:

Carried across the restart Why
:max_runtime_ms deadline Recomputed from a wall-clock anchor, so a 90-minute task that was down for two hours ends immediately instead of starting a second 90 minutes.
on_failure: {:respawn, n} budget A restart must not refill it.
A landed finish report So work that already reported is never re-run.
The callback task_id In-flight pod callbacks stop answering 410 Gone.

Four things to know before you rely on it:

Nothing here is on by default. With persist: false — the default — the orchestrator behaves exactly as it did before the store existed.

Phoenix LiveDashboard integration

If your Phoenix app already mounts Phoenix.LiveDashboard, adding an ExAtlas tab is a one-liner — the library ships ExAtlas.LiveDashboard.ComputePage:

# lib/my_app_web/router.ex
import Phoenix.LiveDashboard.Router
live_dashboard "/dashboard",
metrics: MyAppWeb.Telemetry,
allow_destructive_actions: true, # required for Stop/Terminate buttons
additional_pages: [
atlas: ExAtlas.LiveDashboard.ComputePage
]

Visit /dashboard/atlas to see a live-refreshing table of every tracked compute resource with per-row Touch, Stop, and Terminate controls. The page is only compiled when :phoenix_live_dashboard is in your deps (both LiveDashboard and LiveView are declared as optional: true in ExAtlas, so library-only users pay nothing).

HTTP layer + telemetry

Every provider uses Req under the hood:

Telemetry events

Every request emits [:ex_atlas, <provider>, :request]:

Measurement Value
status HTTP status code
Metadata Value
api :management / :runtime
method :get / :post / :delete / ...
url Full request URL

Wire into your existing telemetry pipeline:

:telemetry.attach(
"atlas-http-logger",
[:ex_atlas, :runpod, :request],
fn _event, measurements, metadata, _ ->
Logger.info("ExAtlas → RunPod #{metadata.method} #{metadata.url} → #{measurements.status}")
end,
nil
)

Per-call Req overrides

Any option accepted by Req.new/1 can be passed via req_options::

ExAtlas.spawn_compute(
provider: :runpod,
gpu: :h100,
image: "...",
req_options: [receive_timeout: 60_000, max_retries: 5, plug: MyPlug]
)

Error handling

All provider callbacks return {:ok, value} or {:error, %ExAtlas.Error{}}. The error struct has a stable :kind atom you can pattern-match on:

Kind When it happens
:unauthorized Bad or missing API key (HTTP 401)
:forbidden API key lacks permission (HTTP 403)
:not_found Resource doesn't exist (HTTP 404)
:rate_limited Provider 429
:timeout Client-side timeout (e.g. run_sync over cap)
:unsupported Provider lacks this capability
:validation ExAtlas-side validation (e.g. missing :endpoint)
:provider Provider-reported 4xx/5xx with no finer bucket
:transport HTTP/socket failure
:unknown Anything else
case ExAtlas.spawn_compute(provider: :runpod, gpu: :h100, image: "...") do
{:ok, compute} -> ...
{:error, %ExAtlas.Error{kind: :unauthorized}} -> rotate_key()
{:error, %ExAtlas.Error{kind: :rate_limited}} -> backoff()
{:error, err} -> Logger.error(Exception.message(err))
end

Writing your own provider

defmodule MyCloud.Provider do
@behaviour ExAtlas.Provider
@impl true
def capabilities, do: [:http_proxy]
@impl true
def spawn_compute(%ExAtlas.Spec.ComputeRequest{} = req, ctx) do
# translate `req` into your cloud's native payload,
# POST it with Req, normalize the response into %ExAtlas.Spec.Compute{}
end
# ... implement the other callbacks ...
end
# Use it without any further configuration:
ExAtlas.spawn_compute(provider: MyCloud.Provider, gpu: :h100, image: "...")

Register it with a short atom by mapping it in your own code — ExAtlas accepts modules directly, so the atom is a convenience:

defmodule MyApp.ExAtlas do
defdelegate spawn_compute(opts), to: ExAtlas
# Or wrap ExAtlas and inject a default provider module
end

Testing

The ExAtlas.Test.ProviderConformance macro runs a shared ExUnit suite against any provider implementation:

defmodule MyCloud.ProviderTest do
use ExUnit.Case, async: false
use ExAtlas.Test.ProviderConformance,
provider: MyCloud.Provider,
reset: {MyCloud.TestHelpers, :reset_fixtures, []}
end

For unit tests that don't actually talk to a cloud, use the built-in ExAtlas.Providers.Mock:

setup do
ExAtlas.Providers.Mock.reset()
:ok
end
test "my code is provider-agnostic" do
{:ok, compute} = MyApp.do_work(provider: :mock)
assert compute.status == :running
end

RunPod tests against the live cloud are tagged @tag :live and are excluded from mix test by default — set RUNPOD_API_KEY and run mix test --only live to enable them.

Security considerations

Troubleshooting & FAQ

Q: (RuntimeError) ExAtlas.Orchestrator is not started You didn't set config :ex_atlas, start_orchestrator: true. The orchestrator is opt-in.

Q: {:error, %ExAtlas.Error{kind: :unauthorized}} on every RunPod call Your API key is missing or wrong. Check the resolution order: per-call api_key: → config :ex_atlas, :runpod, api_key: → RUNPOD_API_KEY env var.

Q: get_job/2 returns {:error, :validation, message: "requires :endpoint"} RunPod's serverless API is scoped to an endpoint id. Pass it: ExAtlas.get_job(job.id, provider: :runpod, endpoint: "abc123").

Q: My pod died on RunPod but my app never noticed. Upstream polling is on by default (:status_poll_ms, 60s). If you set it to false, nothing asks the provider anything and the session only ends on idle TTL. Note that a {:poll_failed, _} event means the poll failed, not the pod — check that event before assuming the resource is gone.

Q: My LiveDashboard ExAtlas tab is empty. Either the orchestrator isn't running, or nothing has been spawned with ExAtlas.Orchestrator.spawn/1. Non-tracked resources (spawned via ExAtlas.spawn_compute/1 directly) don't show in the table — they're not under supervision.

Q: Stop/Terminate buttons don't show. Set allow_destructive_actions: true on the live_dashboard call.

Q: I want to use ExAtlas with httpc / Mint / Finch directly instead of Req. Rewrite the provider module, or pass a custom Req adapter via req_options: [adapter: my_adapter]. The ExAtlas.Provider contract doesn't mandate Req — it's an implementation choice of the bundled providers.

Roadmap

All future providers will be additive; adding a provider never breaks existing call sites.

Contributing

PRs welcome. Before opening:

mix format
mix compile --warnings-as-errors
mix test
mix docs # verify docstrings render

For new providers, the shared conformance suite (test/support/provider_conformance.ex) must pass against your module.

License

Apache-2.0. See LICENSE.