barrel_ngram
Exact substring and regex search over Barrel documents. A byte-level trigram index gives the lexical recall that semantic search misses: identifiers, error strings, config keys, punctuation-heavy literals.
Documentation | HexDocs | Repository
A corpus is bound to a barrel_docdb database. Indexing is driven by the
database's changes feed, and every query result is confirmed against the
real document text, so a trigram false positive is never returned.
Use
%% Open a corpus over a database and build its index.
ok = barrel_ngram:open(<<"code">>, #{db => <<"mydb">>}),
{ok, _Summary} = barrel_ngram:index(<<"code">>),
%% Exact substring search. Each hit carries the document id and the
%% match spans within its indexed text.
{ok, Hits} = barrel_ngram:search(<<"code">>, <<"connect_timeout">>),
%% Regex search (PCRE syntax).
{ok, More} = barrel_ngram:regex(<<"code">>, <<"connect_\\w+timeout">>).
%% Case-insensitive, either mode.
{ok, Ci} = barrel_ngram:search(<<"code">>, <<"connect_timeout">>, #{case_sensitive => false}).
Documentation
- Getting started - open, index, search, regex.
- Selectors - what gets indexed, phase-2 tuning.
- Regex - patterns, what accelerates.
- Sharding - spread a corpus across N shards.
- Operations - refresh, compact, recovery, deletes.
- MCP tool - the
ngram_searchserver tool. - Design - how the index is built and stored.
How it works
- Byte-level trigrams over a 2^24 gram space, direct-addressed through a
flat
u32offset table. - Delta+varint posting lists of local document ordinals, intersected by
galloping search; a corpus can opt into roaring bitmaps
(
postings => roaring) for a native intersection AND on large dense corpora. - Immutable segment files read with
file:pread; a local ordinal maps back to its document key through a sidecar. - A query turns the literal into its overlapping trigrams, intersects
their posting lists, then fetches the candidates and runs the real
substring match (trigram presence is necessary, not sufficient). A
second, content-defined (sparse) index narrows further to a candidate
byte position, and, with a
sourceconfigured, confirms by reading just a window of the document instead of the whole thing. - Literals shorter than a trigram fall back to a scan of the live set.
case_sensitive => falseon eithersearchorregexfor case-insensitive matching.
Status
Substring and regex search over a dense trigram index, across multiple
immutable segments kept live by a push subscription to the changes feed
(updates and deletes reflected via the confirm pass), with crash-safe
manifest recovery and compaction that evicts superseded and deleted
entries. A corpus can be sharded across N nodes by rendezvous hashing
(open option shards => N), and barrel_server exposes it as the
read-only ngram_search MCP tool (mode literal or regex).
Every corpus also builds a second, content-defined (sparse) positional
index alongside the dense one (tuned with phase2_selector_opts): it
narrows candidates to specific byte positions and, with a source
configured, verifies a match by reading a small window instead of the
whole document, for both substring and (a bounded subset of) regex
queries. See selectors.
Segments, manifests and corpus.meta are fsynced before the rename that
commits them. The manifest (version 3) records each segment's sha256 and
size, and open/2 checks them (verify_segments => checksum, the default,
or layout for a header check only); a damaged segment fails open with
{corrupt_segment, Path, Detail}. A query leases its segment files, so a
compaction never deletes a file under a running query, and a read error
fails the query with {segment_read_failed, Path, Reason} instead of
answering "no match". A corpus written by 0.10 or earlier is rebuilt once
when opened with on_legacy => reindex (see
operations).
License
Apache 2.0. See LICENSE.