Elixir kalku

The native kalku for Elixir: it finds sites with Elixir's own parser, casts wekufe into a warm BEAM node with the project loaded, and runs only the tests that cover each one. It speaks the kalku protocol on stdio, to kalku.

This package is half of the tool. It runs inside the project it measures — that is how a wekufe is loaded into a warm BEAM instead of rebuilt — and the kalku binary drives it from outside.

Install

# mix.exs
{:kalku_elixir, "~> 0.1", only: :test, runtime: false}

Then install the kalku binary (how) and, from the root of your project:

mix deps.get
kalku init # detects Elixir, writes .kalku.toml and .kalku/summon
kalku run lib/thing.ex # measure one file

Status

Message Served
hello → ready yes
sites → sites_found yes, all six spells
prepare → prepared yes
baseline → baseline_done yes, with per-test coverage
shutdown → bye yes
cast → cast_done yes
abort → aborted yes
reset → reset_done yes
reload → reloaded yes

ready announces only cast (the native kind); no capability is claimed before it works.

Requirements

Running

MIX_BUILD_PATH=<reni>/build bin/kalku-elixir

Run from the root of the project being measured. bin/kalku-elixir is the supported way to summon this kalku, and it is a shell script because every part of it is load-bearing:

Summoning mix kalku.serve directly still works once the dependencies are built, but on a cold start it puts the compiler's output on the protocol channel. The script exists so nobody has to remember that; it refuses to start at all without a MIX_BUILD_PATH.

Preparing

prepare compiles the project into the reni with protocol consolidation off — a consolidated protocol is built from every implementation at once, so a wekufe cast into a defimpl would be silently ignored and counted as a survivor.

It answers prepared {duration_ms, modules}, or a fatal error:

Code When
reni_not_isolated the build path is not inside the reni, or no reni was given
wrong_env MIX_ENV is not test
prepare_failed the project does not compile, naming the file and line

MIX_TEST_PARTITION is read from hello.env, so each kalku can have its own test database.

The baseline

baseline runs the suite once inside the kalku's own runtime, not in a mix test subprocess: a subprocess would take its results with it and leave the runtime cold for the casts that follow.

It answers baseline_done with status (green or red), every test named by where it is written (test/green_test.exs:4, relative to the project, so every kalku in a pool calls the same test the same thing), how long each took, and the failures quoted from ExUnit's own words.

Two things it does that are easy to miss:

Per-test coverage

baseline reports which tests execute each line, which is what lets a wekufe be cast against the handful of tests that reach its line instead of the whole suite. Speed is the product, and most of it comes from here.

It costs a second pass over the suite, and that is not an oversight. :cover counts per line, not per test, and ExUnit delivers its formatter events asynchronously — a test_started can arrive after the test it announces has already run, so clearing counters there clears the next test's lines. The only honest attribution is to run each test on its own, with the counters cleared before it. That happens once, in the baseline.

Only the project's own modules are instrumented: nobody mutates a dependency, so counting its lines would cost time and say nothing.

Coverage travels inline when it fits under hello.inline_limit_bytes, and otherwise is written to coverage.json in the reni and reported as coverage_path — a megabyte of JSON per worker is a cost the protocol lets us decline.

A line no test runs has no entry at all: the question is which tests cover a line, and for an uncovered line the honest answer is none, which the kaikai side reads as no_coverage rather than as a hole.

The kalku declares OTP's :tools application, which is where :cover lives.

Casting

cast splices the site's span into the module's source in memory, compiles from a string, and loads the result into the warm runtime. Nothing is written to the project: the file on disk is the one the developer left there, before the cast and after it.

The original modules are kept before anything is compiled and reloaded on every path out, including a compile error. A wekufe that outlived its cast would be attributed to the next one, and the next one's result would be a lie.

Outcome When
equivalent the compiled code is the original's, so no test could notice; no test is run
killed a selected test failed, and killed_by names it
survived every selected test ran and none noticed
compile_error the wekufe does not compile, with the first line of why

Equivalence is proved rather than guessed: the wekufe's modules and the original's are compared by their BEAM MD5s, which is the runtime's own answer to "is this the same code". A kalku never reports an equivalence it cannot demonstrate.

Only the tests in cast.tests run, selected by file and line the way mix test path:line does, stopping at the first failure. Those are the tests the baseline's coverage says reach the changed line; running the rest would cost time and could not change the answer.

Aborting

abort stops a cast where it stands and keeps the kalku warm, which is the whole point: a warm kalku is the most expensive thing kalku owns.

For it to be possible at all, the loop reads while it works. Reading happens in one process and the cast in another, so a loop that read one line, answered it, and only then read again could never receive an abort — the message only matters in the middle of the cast it stops.

Killing the casting process is not enough. ExUnit runs each test in a process it monitors rather than links, so a test looping forever outlives the cast that started it and would burn a core for the rest of the run — measured, not assumed. So an abort also stops everything unnamed that appeared while the cast ran. That is coarse on purpose: a wekufe is the reason any of it is there.

Casts are served one at a time, in the order they arrive. A kalku has exactly one runtime, so two casts at once would measure each other; one that arrives early waits rather than being refused. shutdown finishes what is under way before saying bye, since leaving without it would have the kaikai side report a measured wekufe as crashed.

Reloading, and what depends on what

Between runs a developer edits, and a kalku that stayed warm is holding the code from before. reload {files} recompiles them — and whatever is stitched into them at compile time.

A module that uses another's macro has the expansion baked in. Recompiling only the macro's own module leaves the caller running the old one, so a wekufe cast there is loaded but not running anywhere a test can reach it. It comes back survived, and that is a false survivor: a hole reported where the measurement never happened. The fixture shows it as a test — the same wekufe, the same test, survived with reload: "module" and killed with reload: "dependents".

So sites marks each site with the reload its file needs, and a cast on a dependents site recompiles the dependents too and restores them afterwards. The graph comes from mix xref, the compiler's own answer rather than a guess of ours; its output is captured rather than printed, and the quiet shell the kalku runs under is lifted for the length of the question, since a quiet shell answers nothing.

reload also makes the kalku forget the suite it loaded, since a test file may be among the changed ones.

Dirty state and resetting

State one cast leaves behind is read as the next cast's doing, and the next result is a lie. So each cast is weighed against a mark of what a clean runtime looks like, and one that moved anything answers dirty: true.

What is watched is named ETS tables and the project's application env. Anonymous tables belong to whoever holds them and vanish with it; processes come and go under a supervisor without anything being wrong. Watching those would call a healthy runtime dirty on every cast — measured, not assumed: a snapshot of the two that are watched is identical across runs.

The mark is taken again after the baseline, and that one is what counts. The suite creates named tables the first time it runs, so a mark from before them would have reset delete the test framework's own state and leave the kalku unable to run a test at all. That is exactly what happened when the mark was taken in prepare: the reset succeeded, reported clean: true, and every cast after it survived because nothing could run.

reset deletes the tables the project added, restores the application env entry by entry — both directions, since restoring only known keys would leave the additions — and restarts the application. It answers clean with whether the runtime matches the mark afterwards; a reset that did not work says so, and the kaikai side recycles the kalku rather than trusting it.

Nothing may touch the runtime while a cast is using it, so a reset that arrives mid-cast waits for it. A reset in the middle of a cast cleans up the very state that cast was about to be judged on, and both answers come out wrong.

Sites

Sources are parsed with Code.string_to_quoted/2 (columns, token_metadata, and a literal encoder that keeps positions on literals). The AST decides what and where; the source text only confirms that the expected token sits at the reported position, and a node whose text cannot be pinned down exactly yields no site.

Spell Proposes
arm delete one clause of case, cond, with … else, receive, fn, or a multi-clause function; whole lines, only when every clause starts its own line; never the only clause
compare >↔>=, <↔<=, ==↔!=, ===↔!==, in expressions and guards
connect and↔or, &&↔||
negate if↔unless; drop not / !
literal integer n→n+1, true↔false, :ok↔:error, a non-empty plain string → ""
call drop one pipe stage; f(x, …) → x when x is a variable or a literal

Never proposed: test files and scripts (test/, anything but .ex); @moduledoc, @doc, typespecs, and other directive attributes; anything inside calls matched by exclude_calls; literals inside raise; if/unless whose branches are identical; calls in patterns, guards, or a pipe's right side.

Every site carries its span (line, column in codepoints, byte offset), original, replacement, enclosing (Module.fun/arity), and ordinal. A candidate whose wekufe would not parse is dropped and counted on stderr.

Tests

mix test

test/fixtures/ is excluded from mix format: those files are written a particular way on purpose.