Elixir

BindingsRustPythonNode.jsWASMJavaGoC#PHPRubyElixirDartKotlinSwiftZigC FFIDockerHelm chartLicenseDocumentationHugging Face
Join DiscordLive DemoGitHub Stars

Extract text, tables, images, metadata, and code intelligence from 101 file formats and 306 programming languages including PDF, Office documents, images, and audio/video transcripts where native transcription is available. Elixir bindings with native BEAM concurrency, OTP integration, and idiomatic Elixir API.

What This Package Provides

Installation

Package Installation

Add to your mix.exs dependencies:

def deps do
[
{:xberg, "~> 1.0.8"}
]
end

Then run:

mix deps.get

System Requirements

Quick Start

Basic Extraction

Extract text, metadata, and structure from any supported document format:

# Basic document extraction workflow
# Load file -> extract -> access results
{:ok, output} = Xberg.extract(input: %Xberg.ExtractInput{kind: :uri, uri: "document.pdf"}, config: nil)
result = List.first(output.results)
IO.puts("Extracted Content:")
IO.puts(result.content)
IO.puts("\nMetadata:")
IO.puts("Format: #{inspect(result.metadata.format)}")
IO.puts("Tables found: #{length(result.tables)}")

Common Use Cases

Extract with Custom Configuration

Most use cases benefit from configuration to control extraction behavior:

With OCR (for scanned documents):

Table Extraction

See Configuration Guide for table extraction options.

Processing Multiple Files

Async Processing

For non-blocking document processing:

Next Steps

Features

Supported File Formats (101 formats · 115 file extensions) 101 formats across 115 file extensions in 8 major

categories with intelligent format detection and comprehensive metadata extraction. #### Office Documents | Category | Formats | Capabilities | |----------|---------|--------------| | Word Processing | .docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6 | Full text, tables, images, metadata, styles | | Spreadsheets | .xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers | Sheet data, formulas, cell metadata, charts | | Presentations | .pptx, .pptm, .ppt, .ppsx, .potx, .potm, .pot, .odp, .key | Slides, speaker notes, images, metadata | | PDF | .pdf | Text, tables, images, metadata, OCR support | | eBooks | .epub, .fb2 | Chapters, metadata, embedded resources | | Database | .dbf | Table data extraction, field type support | | Hangul | .hwp, .hwpx | Korean document format, text extraction | #### Images (OCR-Enabled) | Category | Formats | Features | |----------|---------|----------| | Raster | .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif | OCR, table detection, EXIF metadata, dimensions, color space | | Advanced | .jp2, .jpx, .jpm, .mj2, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm | OCR via hayro-jpeg2000 (pure Rust decoder), JBIG2 support, table detection, format-specific metadata | | HEIC family | .heic, .heics, .heif, .avif, .avcs | EXIF metadata, optional libheif pixel decoding | | Vector | .svg | DOM parsing, embedded text, graphics metadata |

Audio & Video | Category | Formats | Features | |----------|---------|----------| | Audio | .mp3, .mpga

.m4a, .wav, .webm | Whisper transcription when native transcription is available | | Video audio track | .mp4, .mpeg, .webm | Audio-track transcription only |

Web & Data | Category | Formats | Features | |----------|---------|----------| | Markup | .html, .htm

.xhtml, .xml, .svg | DOM parsing, metadata (Open Graph, Twitter Card), link extraction | | Structured Data | .json, .yaml, .yml, .toml, .csv, .tsv | Schema detection, nested structures, validation | | Text & Markdown | .txt, .md, .markdown, .djot, .mdx, .rst, .org, .rtf | CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode | #### Email & Archives | Category | Formats | Features | |----------|---------|----------| | Email | .eml, .msg, .pst | Headers, body (HTML/plain), attachments, threading | | Archives | .zip, .tar, .tgz, .gz, .7z | Recursive extraction of nested archives, file listing, metadata, zip-bomb protection |

Academic & Scientific | Category | Formats | Features | |----------|---------|----------| | Citations | .bib

.ris, .nbib, .enw | Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX, CSL JSON by MIME type | | Scientific | .tex, .latex, .typ, .typst, .jats, .ipynb | LaTeX, Typst, Jupyter notebooks, PubMed JATS | | Publishing | .fb2, .docbook, .dbk, .docbook4, .docbook5, .opml | FictionBook, DocBook XML, OPML outlines | | Documentation | MIME-only POD, mdoc, troff | Technical documentation formats | #### Code Intelligence (306 Languages) | Feature | Description | |---------|-------------| | Structure Extraction | Functions, classes, methods, structs, interfaces, enums | | Import/Export Analysis | Module dependencies, re-exports, wildcard imports | | Symbol Extraction | Variables, constants, type aliases, properties | | Docstring Parsing | Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats | | Diagnostics | Parse errors with line/column positions | | Syntax-Aware Chunking | Split code by semantic boundaries, not arbitrary byte offsets | Powered by tree-sitter-language-packdocumentation. Complete Format Reference ### Key Capabilities - Text Extraction - Extract all text content with position and formatting information - Metadata Extraction - Retrieve document properties, creation date, author, etc. - Table Extraction - Parse tables with structure and cell content preservation - Image Extraction - Extract embedded images and render page previews

OCR Support

Xberg supports multiple OCR backends for extracting text from scanned documents and images:

OCR Configuration Example

Async Support

This binding provides full async/await support for non-blocking document processing:

Plugin System

Xberg supports extensible post-processing plugins for custom text transformation and filtering.

For detailed plugin documentation, visit Plugin System Guide.

Plugin Example

alias Xberg.Plugin
# Word Count Post-Processor Plugin
# This post-processor automatically counts words in extracted content
# and adds the word count to the metadata.
defmodule MyApp.Plugins.WordCountProcessor do
@behaviour Xberg.Plugin.PostProcessor
require Logger
@impl true
def name do
"WordCountProcessor"
end
@impl true
def processing_stage do
:post
end
@impl true
def version do
"1.0.0"
end
@impl true
def initialize do
:ok
end
@impl true
def shutdown do
:ok
end
@impl true
def process(result, _options) do
content = result["content"] || ""
word_count = content
|> String.split(~r/\s+/, trim: true)
|> length()
# Update metadata with word count
metadata = Map.get(result, "metadata", %{})
updated_metadata = Map.put(metadata, "word_count", word_count)
{:ok, Map.put(result, "metadata", updated_metadata)}
end
end
# Register the word count post-processor
Plugin.register_post_processor(:word_count_processor, MyApp.Plugins.WordCountProcessor)
# Example usage
result = %{
"content" => "The quick brown fox jumps over the lazy dog. This is a sample document with multiple words.",
"metadata" => %{
"source" => "document.pdf",
"pages" => 1
}
}
case MyApp.Plugins.WordCountProcessor.process(result, %{}) do
{:ok, processed_result} ->
word_count = processed_result["metadata"]["word_count"]
IO.puts("Word count added: #{word_count} words")
IO.inspect(processed_result, label: "Processed Result")
{:error, reason} ->
IO.puts("Processing failed: #{reason}")
end
# List all registered post-processors
{:ok, processors} = Plugin.list_post_processors()
IO.inspect(processors, label: "Registered Post-Processors")

Embeddings Support

Generate vector embeddings for extracted text using the built-in ONNX Runtime support. Requires ONNX Runtime installation.

Embeddings Guide

Batch Processing

Process multiple documents efficiently:

Configuration

For advanced configuration options including language detection, table extraction, OCR settings, and more:

Configuration Guide

Documentation

Contributing

Contributions are welcome! See Contributing Guide.

Part of Xberg.dev

License

MIT License — see LICENSE for details.

Support