TFLiteBEAM

TensorFlow Lite BEAM bindings with optional EdgeTPU support.

Hex.pmCoverage Status

OSArchABIBuild StatusHas Precompiled Library
Ubuntu 20.04x86_64gnuCIYes
Ubuntu 20.04arm64gnuCIYes
Ubuntu 20.04armv7lgnueabihfCIYes
Ubuntu 20.04armv6gnueabihfCIYes
Ubuntu 20.04riscv64gnuCIYes
macOS 15 Sequoiax86_64darwinCIYes
macOS 14 Sonomaarm64darwinCIYes

Delegates

tflite_beam_interpreter_builder:build/2 attaches an XNNPACK delegate for you, unless you have attached one yourself. TfLite would otherwise apply XNNPACK on its own, invisibly, inside allocate_tensors/1 -- with a thread count nothing could reach and no way to decline it. The acceleration is the same; where it happens is now visible, and set_num_threads/2 still reaches it.

{ok, Resolver} = tflite_beam_ops_builtin_builtin_resolver:new(),
{ok, Builder} = tflite_beam_interpreter_builder:new(Model, Resolver),
%% your own delegate instead of the default one
{ok, Delegate} = tflite_beam_delegate:xnnpack(#{num_threads => 4}),
ok = tflite_beam_interpreter_builder:add_delegate(Builder, Delegate),
ok = tflite_beam_interpreter_builder:build(Builder, Interpreter).

tflite_beam_delegate:available/0 lists the delegate kinds this build can create: XNNPACK on every target except armv6 and armv7l, where nothing is attached and inference runs as it always has.

To go back to TfLite delegating by itself, ask the resolver for it:

{ok, Resolver} = tflite_beam_ops_builtin_builtin_resolver:new(#{apply_default_delegates => true}),

A delegate must outlive every interpreter built from the builder it was added to, so there is no way to detach or free one: the builder and each interpreter hold it for as long as they need it, and it goes when they do.

Delegates from a shared library

Anything implementing TfLite's delegate plugin interface -- tflite_plugin_create_delegate and tflite_plugin_destroy_delegate -- can be loaded at runtime, which covers Edge TPU, a GPU delegate built elsewhere, and vendor delegates this library knows nothing about:

{ok, Delegate} = tflite_beam_delegate:external("/opt/lib/libvendor_delegate.so",
#{device => 0, precision => fp16}),
ok = tflite_beam_interpreter_builder:add_delegate(Builder, Delegate).

Options are handed to the plugin as strings, which is the whole of that ABI, so atoms and integers are converted and at most 256 pairs fit. What the keys mean is the plugin's business. The path is resolved to an absolute one before loading, because the loader is asked for exactly the file named -- a bare libfoo.so would otherwise be searched for wherever the system looks, which is rarely where anyone means. The library is not unloaded afterwards.

Threading

An interpreter, and any delegate attached to it, belongs to one process at a time. TfLite documents tflite::Interpreter as not thread-safe and leaves serialising access to the caller, and invoke/1 runs on a dirty scheduler, so two processes sharing one interpreter really do run it on two OS threads at once.

The direct API mirrors the C API, which means feeding an interpreter, running it and reading the result back are three separate calls -- and nothing in the C API says they have to be treated as one. Two processes taking turns badly get each other's answers: measured on a real model, 147 wrong results in 400 calls, silently and without a crash.

If you want that handled for you, use tflite_beam_interpreter_server:

{ok, Server} = tflite_beam_interpreter_server:start_link(ModelPath),
Output = tflite_beam_interpreter_server:predict(Server, [Input]).

The interpreter lives inside that process, so feeding, running and reading back is one step nothing can interleave with, and concurrent callers each get the answer to their own input. Use with/2 for the sequences predict/2 does not cover.

The direct API is unchanged and stays available. Two things guard it:

Delegates are the same story: nothing documents a TfLiteDelegate as safe to back two interpreters simultaneously, and XNNPACK's demonstrably is not.

Coral Support

libedgetpu is itself a TfLite delegate plugin, so an Edge TPU can be attached like any other delegate -- which means it composes with set_num_threads/2 and with whatever else is on the builder:

{ok, Delegate} = tflite_beam_coral:edge_tpu_delegate(),
ok = tflite_beam_interpreter_builder:add_delegate(Builder, Delegate),
ok = tflite_beam_interpreter_builder:build(Builder, Interpreter).

tflite_beam_coral:make_edge_tpu_interpreter/2 still works and is unchanged. It builds its own interpreter internally, though, so nothing set on a builder reaches it; the delegate above is the composable route. Asking for a device that is not there is an ordinary {error, Reason} from edge_tpu_delegate/1.

Both routes have been checked to produce identical output on a USB Coral accelerator, running mobilenet_v2_1.0_224_inat_bird_quant_edgetpu.tflite against libedgetpu 0.1.14 on macOS arm64.

Dependencies

For macOS

# only required if not using precompiled binaries
# for compiling libusb
brew install autoconf automake

For some Linux OSes you need to manually execute the following command to update udev rules, otherwise, libedgetpu will fail to initialize Coral devices.

bash "3rd_party/cache/${TFLITE_BEAM_CORAL_LIBEDGETPU_RUNTIME}/edgetpu_runtime/install.sh"

Compile-Time Environment Variable

Installation

Add tflite_beam to your list of dependencies in rebar.config:

{deps, [
{tflite_beam, "0.3.12"}
]}

Documentation is published on HexDocs.

Tests

rebar3 ct

The model fixtures live in test/models/, so the suite needs no network and runs against a precompiled install as well as a build from source.

Releasing

The precompiled tarballs only exist once the v* tag has been pushed and the precompile matrix has finished, and the manifest that verifies them has to be inside the published package -- so it is generated in between:

git tag -a vX.Y.Z -m "vX.Y.Z" && git push origin vX.Y.Z # matrix builds the 7 targets
scripts/generate_checksums.sh X.Y.Z # writes checksum.term
rebar3 hex publish

checksum.term is not tracked in git and does not need to be: it is packaged from the working directory, and the tarballs it lists do not exist until the tag has been built -- so a tracked copy would always be one release out of date.

Skipping the middle step publishes a package that cannot check what it downloads, which it says out loud on install rather than doing quietly.

Upstream Dependencies