Tay
Durable background jobs without a database or message broker.
Tay is a durable, single-node background-job engine with its own append-only storage. It can run embedded inside an Elixir application or as standalone infrastructure in a production OCI container.
No DataBase. No Redis. No RabbitMQ. Tay owns its durable job state itself.
Application / Worker
│
│ local Unix socket
▼
┌─────────────────────┐
│ Tay │
│ │
│ queues · retries │
│ schedules · cron │
│ cancellation │
│ crash recovery │
└──────────┬──────────┘
│
▼
durable local disk
Tay can be used in three ways:
- Elixir library — add
tayfrom Hex and run the Engine inside your supervision tree. - Standalone container — run Tay as infrastructure beside your application, without installing Elixir or Erlang on the host.
- Python client — execute Python tasks through the supported
tay-clientSDK while Tay owns scheduling, retries, persistence, and recovery.
The optional Tay Dashboard adds a Phoenix LiveView operational UI and is also published as a ready-to-run container image.
Tay is intentionally single-node and single-writer. It is designed for cases where durable background execution is needed without operating a separate database or distributed broker. It is not a replacement for a distributed multi-node message bus.
Quick start with Docker
The easiest way to try Tay does not require Elixir or Erlang on the host.
Pull the standalone image:
docker pull ghcr.io/andriisydorenko1904/tay:1.0.0-rc.9
For a ready-made Tay deployment example:
docker compose up --build
To add the optional local Python socket worker, use
docker compose --profile local-worker up --build.
The Tay container owns durable state at /var/lib/tay. Applications and
language workers can use the shared /run/tay Unix socket or the optional
HTTP/JSON API with mTLS. Set TAY_SOCKET_PATH=off for HTTP-only standalone use.
Restarting or replacing the Tay container does not discard jobs as long as the data volume is preserved.
Want the operational UI as well?
TAY_ENABLE_DASHBOARD=true docker compose up --build
Then open:
http://localhost:4000/tay
The same image contains the optional dashboard. Enable it with
TAY_ENABLE_DASHBOARD=true:
docker pull ghcr.io/andriisydorenko1904/tay:1.0.0-rc.9
Why Tay?
A durable background-job system needs somewhere to keep pending, running, scheduled, and retryable work.
Many systems delegate that responsibility to DB, Redis, or a message broker.
Tay takes a different approach:
the job engine includes its own crash-safe local store.
That makes Tay useful when you want:
- durable queues without provisioning a separate database or broker;
- retries, scheduling, cron, cancellation, and queue concurrency in one runtime;
- recovery of durable job state after process or container restarts;
- a small infrastructure footprint for a single-machine or single-node deployment;
- Elixir-native embedding when the application is already written in Elixir;
- a standalone container when the application is written in another language;
- Python workers without moving persistence and orchestration into Python.
Tay provides at-least-once execution. External effects must therefore be idempotent or reconciled by the application.
Architecture
Tay separates durable orchestration from application code.
When running standalone, Tay can be treated as a small infrastructure component:
┌──────────────────────┐
│ Application / Worker │
│ │
│ Python / other │
│ local runtime │
└──────────┬───────────┘
│
│ Unix socket
│
▼
┌──────────────────────┐
│ Tay │
│ │
│ durable job engine │
│ queues │
│ schedules / cron │
│ retries │
│ cancellation │
│ recovery │
└──────────┬───────────┘
│
▼
persistent volume
Tay remains a single-node durable engine, while producers and workers may
connect through Bandit HTTP/JSON from another host. Local workers can still use
the Unix-domain socket. See
docs/protocol.md.
This keeps the single-node trust and failure model explicit while allowing the Engine to be packaged and operated independently from application workers.
For Elixir applications, no service boundary is required at all. The same Engine can run directly inside the application's supervision tree.
Elixir integration
Use Elixir 1.20 and Erlang/OTP 29 with a C11 compiler (cc) available when
building.
Add Tay to your application's Mix dependencies:
defp deps do
[
{:tay, "1.0.0-rc.9"}
]
end
Or use the local checkout:
{:tay, path: "../tay"}
Then run:
mix deps.get
mix compile
For a development run, choose a dedicated absolute directory and initialize it once. Initialization refuses existing history and is never a repair command:
export TAY_DATA_DIR="$PWD/var/tay"
mix tay.storage.init --data-dir "$TAY_DATA_DIR" --durability write
Define a worker with a stable key:
defmodule MyApp.HelloWorker do
use Tay.Worker, key: "hello.v1", queue: :default
@impl true
def perform(%Tay.Job{args: %{"name" => name}}) do
IO.puts("Hello, #{name}!")
:ok
end
end
Add Tay to your application's supervisor.
The :write mode below is for development only and has no power-loss
durability guarantee:
children = [
Tay.child_spec(
data_dir: System.fetch_env!("TAY_DATA_DIR"),
durability: :write,
workers: %{"hello.v1" => MyApp.HelloWorker},
queues: [default: 2]
)
]
Supervisor.start_link(children, strategy: :one_for_one)
Explicit initialization remains the default.
An embedding application may instead pass:
initialize: :if_missing
to Tay.child_spec/1.
That option only initializes a genuinely missing storage root before the normal full recovery path. It never repairs, replaces, truncates, or reinitializes existing storage.
Once the Engine is ready, submit and inspect a job:
%{state: :ready} = Tay.status()
{:ok, intent} =
MyApp.HelloWorker.new(%{"name" => "Ada"})
{:ok, job} =
Tay.insert(intent)
{:ok, current} =
Tay.get_job(job.id)
IO.inspect(current.state)
List and summarize jobs through the bounded public inspection API:
{:ok,
%{
jobs: jobs,
total_count: total,
previous_cursor: previous,
next_cursor: next,
last_cursor: last
}} =
Tay.jobs(
states: [:retryable, :discarded],
queues: [:default],
limit: 50
)
{:ok, counts} = Tay.stats()
{:ok, queues} = Tay.queues()
The tay package includes a Phoenix LiveView dashboard for these APIs.
Its modules live under lib/tay/dashboard/; see the
dashboard guide for router mounting and access control.
Keep the original intent until an insertion outcome is known. A lost reply may follow a durable write; reconcile by job ID or resubmit the same intent, never a newly generated ID.
External worker effects may run again after a crash, so make them idempotent or reconcile them in the application.
Python integration
Python is a first-class external-worker integration.
Tay remains responsible for:
- durable job state;
- queues;
- scheduling;
- retries;
- cancellation;
- crash recovery.
Python processes execute application code.
When Tay is embedded in Elixir, the Engine starts a local Unix-domain socket automatically.
When Tay runs as a standalone container, local Python workers may share the socket volume. Workers on another host instead connect to the mTLS HTTP API; no shared socket is needed.
The same tay-client SDK is used in both cases.
Install it from PyPI:
python -m pip install tay-client
Or from this checkout:
python -m pip install ./clients/python
Then define and submit tasks:
from tay import Tay
tay = Tay(capacity=4)
@tay.task(name="billing.capture.v1")
def capture(invoice_id: str) -> dict:
return {"invoice_id": invoice_id}
async def submit() -> None:
await tay.start()
job = await capture.enqueue(
"inv-42",
submission_id="capture:inv-42",
)
print(await job.status())
This gives Python applications a simple split:
Python
│
│ execute application code
▼
tay-client
│
│ Unix socket
▼
Tay
│
├── persistence
├── queues
├── scheduling
├── retries
└── recovery
The Python application does not need to implement its own durable queue or persistence layer.
Python schedules
The Python client accepts standard five-field Unix cron expressions.
Schedules use UTC by default; pass a fixed numeric offset when needed.
The API accepts:
catch_up="latest"
catch_up="all"
and defaults to:
overlap="skip"
Example:
# At 08:00 on weekdays in UTC+02.
await tay.schedule(
"billing.capture.v1",
cron="0 8 * * 1-5",
timezone="+02",
)
# Start in ten minutes, then repeat every fifteen minutes.
await tay.every(
"billing.capture.v1",
minutes=15,
delay=600,
)
Use start_at=<UTC milliseconds> for an absolute first run instead of
delay; those options cannot be combined.
The listener schedule registry belongs to the live execution generation. A connected Python client retains successful declarations and recreates them as part of every reconnect, including after compaction. A client-process restart still requires application startup to declare schedules again. Missed-time catch-up and enforced overlap policies are not yet implemented. Once a timer is due, temporary admission pressure delays that occurrence instead of dropping it; its deterministic job ID makes retries idempotent.
The optional Bandit HTTP/JSON API exposes producer and worker operations over TCP. It is disabled by default; plaintext binds only to loopback and remote access requires mTLS. Tay's storage remains single-node; remote workers do not turn the engine into a distributed Store.
See the protocol contract for discovery, security, request types, and result retention.
Standalone container
The release workflow publishes a self-contained Linux image for amd64 and
arm64:
ghcr.io/andriisydorenko1904/tay:1.0.0-rc.9
It includes the Erlang VM and Tay runtime.
The host therefore needs Docker or another OCI runtime, not Elixir or Erlang.
Pull it directly:
docker pull ghcr.io/andriisydorenko1904/tay:1.0.0-rc.9
A typical non-Elixir deployment runs the image beside an application worker:
┌────────────────────┐
│ Application worker │
└─────────┬──────────┘
│
│ /run/tay
│
┌─────────▼──────────┐
│ Tay container │
│ │
│ durable job engine │
└─────────┬──────────┘
│
│ /var/lib/tay
▼
persistent volume
The application and Tay share only the /run/tay socket volume.
Durable state belongs on:
/var/lib/tay
The Compose example starts Tay; the local worker uses an optional profile:
docker compose up --build
Run docker compose --profile local-worker up --build to include that worker.
The container is:
- single-node;
- single-writer;
- non-root;
- compatible with a read-only root filesystem.
The Tay container does not execute user job code.
Deleting its data volume deletes Tay's durable state.
Deleting its socket volume does not.
See the standalone runtime guide for volumes, permissions, configuration, health checks, restart behavior, and deployment details.
Dashboard
Tay Dashboard is an optional operational web UI built with Phoenix LiveView.
The tay image contains Tay core and the small Phoenix host in one release.
The web endpoint is disabled by default and enabled with:
TAY_ENABLE_DASHBOARD=true
Start the repository example:
TAY_ENABLE_DASHBOARD=true docker compose up --build
Then open:
http://localhost:4000/tay
The example binds only to host loopback and leaves optional Basic authentication disabled.
The single Tay container owns the Store and optionally serves the dashboard.
See the complete dashboard container guide for configuration, optional Basic authentication, reverse-proxy guidance, persistence, and upgrade rules.
Project status
Tay 1.0 is currently a release candidate with a deliberately constrained, target-validated operating profile. Release-candidate feedback may lead to compatible fixes and documentation changes before the final 1.0.0 release; the documented persistent-format contracts are already fixed.
The storage format, recovery behavior, supported filesystems, and durability requirements are documented explicitly.
Production :sync requires Linux and an explicitly validated supported local
filesystem.
Before production use, read:
- Operations
- Storage contract
- Compatibility
- Idempotency and external effects
- Standalone runtime
- Worker protocol
Operate safely
Production :sync requires Linux, an explicitly validated supported local
filesystem, and:
validated_filesystem: true
macOS supports explicit development :write only.
Tay evaluates obsolete history automatically through bounded-retention, online Store-v2 compaction. Candidate construction keeps admission and execution live; a fenced, validated switch publishes the replacement epoch. Initial Store-v1 adoption still drains and recovers once.
It has no:
- automatic tail repair;
- live backup;
- exactly-once external effects.
A corrupt or unsupported history refuses writable startup and preserves the evidence.
Do not delete individual log files to recover capacity.
See operations for initialization, inspection, cold backup/restore, and incident response.
The storage contract describes persistent formats and recovery.
Compatibility covers supported platforms, upgrade boundaries, and finite release limits.
The standalone guide covers the official Docker/OCI runtime.
License
The Tay engine is source-available under the Elastic License 2.0.
Internal use, including use inside an ordinary commercial SaaS product, is free.
The license does not permit offering Tay itself, or a substantial set of its functionality, to third parties as a hosted or managed service.
Commercial terms for that use are available separately; see commercial licensing.
The separately distributed Python client is licensed under MIT.