Xmlixir

Binary-native XML scanning for Elixir, modeled after the node and scanner concepts exposed by Erlang/OTP's xmerl.

{:ok, document} = Xmlixir.parse("<root id=\"1\"><child>value</child></root>")
document.name
#=> "root"

Parser entry points accept UTF-8, BOM-marked UTF-16LE/UTF-16BE, and declared ISO-8859-1 binaries, converting accepted input to UTF-8 internally. Other source encodings must be converted by the caller first. Xmlixir.parse/1 returns a strict, single-document result. Xmlixir.scan/1 returns {element, remainder} like :xmerl_scan.string/1. Xmlixir.file/1 reads and parses the file.

Elements expose binary name, attributes, content, namespace metadata, ordered structured attribute nodes, and OTP-style parent/position metadata. Use Xmlixir.children/2, Xmlixir.descendants/2, Xmlixir.attribute/2, and Xmlixir.text/1 for traversal.

Parsed nodes can be serialized back to XML iodata with Xmlixir.to_iodata/1. Serialization preserves the parsed document structure and escapes text and attribute values, but it does not reproduce source formatting, XML declarations, or DTDs that are not part of the DOM node.

Xmlixir is a binary-native DOM parser with focused traversal helpers. XPath-style selection is available through Xmlixir.XPath.select/2, including the native namespace:: axis, for example:

Xmlixir.XPath.select(document, "//product[@id='1']/name/text()")

Xmlixir exposes only binary-native structs and values; it does not expose or require Erlang/xmerl records or charlists. XPath variables can be supplied to Xmlixir.XPath.select/4 as a map: Xmlixir.XPath.select(document, "$items[1]", [], variables: %{items: nodes}).

For XPath mappings, Xmlixir.Query provides the string-based ~x sigil, xpath/3, xmap/3, parse/2, stream_tags/3, stream/2, and stream_events/2 APIs. Query results and mapped values are returned as binaries.

Querying and mappings

Import Xmlixir.Query to use the ~x sigil:

import Xmlixir.Query
xml = "<catalog><product id=\"1\"><name>Pen</name><price>2.50</price></product></catalog>"
name = xpath(xml, ~x"//product/name/text()"s)
#=> "Pen"
product = xpath(xml, ~x"//product"e)
#=> %Xmlixir{name: "product", ...}

The modifiers use a compact, familiar query shape. Xmlixir returns binaries rather than charlists, and native nodes rather than xmerl records.

Modifier Result
e the native Xmlixir node/value
l all matches as a list
k a keyword-list mapping
o nil when the path is absent
s, S strict or soft string
i, I strict or soft integer
f, F strict or soft float

Modifiers can be combined, for example ~x"//price/text()"il for a list of integers or ~x"//score/text()"Fo for a soft optional float.

Mappings use relative paths and can be nested:

import Xmlixir.Query
xpath(xml, ~x"//product"l,
id: ~x"./@id"s,
name: ~x"./name/text()"s,
price: ~x"./price/text()"f
)
#=> [%{id: "1", name: "Pen", price: 2.5}]
xmap(xml, products: [~x"//product"l, name: ~x"./name/text()"s])
#=> %{products: [%{name: "Pen"}]}

Use transform_by/2 for reusable result transformations:

uppercase = ~x"//product/name/text()"s |> transform_by(&String.upcase/1)
xpath(xml, uppercase)
#=> "PEN"

Namespaces

For namespace-aware queries, parse with namespace_conformant: true, create a query, and attach the prefixes used by the query. Query prefixes are independent of the prefixes present in the source document:

import Xmlixir.Query
doc = parse("<catalog xmlns:p=\"urn:products\"><p:product id=\"1\"/></catalog>",
namespace_conformant: true
)
spec = ~x"//p:product/@id"s |> add_namespace("p", "urn:products")
xpath(doc, spec)
#=> "1"

The same mapping works for qualified attributes. Qualified wildcards and the native namespace:: axis are available through Xmlixir.XPath.

XML Observatory

This repository includes a small CLI test harness for collecting real-world XML and replaying it against Xmlixir. Captures preserve the raw response in .xml files and write URL, status, headers, timestamp, size, and SHA-256 metadata in adjacent .metadata.term files.

mix xml_observatory.fetch google_news
mix xml_observatory.parse 'priv/captures/**/*.xml'
mix xml_observatory.compare 'priv/captures/**/*.xml'

Configured sources include google_news, deterministic arXiv Atom API (max_results=1), and github; an explicit HTTPS URL can also be passed to fetch. Fetching is bounded to 10 MiB by default, allows only HTTPS, follows at most three HTTPS redirects, and leaves external entities disabled when parsing. The capture directory is intentionally append-only so live observations can become deterministic regression fixtures.

The observatory modules are kept independent of Phoenix: XmlObservatory.Fetcher handles retrieval, XmlObservatory.Capture handles persistence, XmlObservatory.ParserRun reports parse metrics, and XmlObservatory.Comparator provides a small Xmlixir/xmerl shape comparison.

Streaming

Use Xmlixir.Query.stream_tags/3 when only selected elements are needed. The source may be a binary, iodata, or a chunked enumerable such as a file stream:

import Xmlixir.Query
File.stream!("catalog.xml", [], 4_096)
|> stream_tags(:product, discard: [:large_unused_subtree])
|> Stream.map(fn {_tag, product} -> xpath(product, ~x"./name/text()"s) end)
|> Enum.to_list()

stream_tags!/3 provides a strict stream entry point. stream/2 emits every element, while stream_events/2 emits parser events. Streams are lazy: malformed input and other stream errors are raised when consumed, not necessarily when constructed. Chunk boundaries may split XML tags.

Internal entity expansion requires allow_entities: true. Xmlixir.file/2 resolves relative external DTDs and referenced external entities; Xmlixir.parse/2 does not perform filesystem I/O. namespace_conformant: true resolves namespace-expanded names, and validation: :dtd enables internal-DTD validation of declared elements, required/declared attributes, enumerations, content models, and ID/IDREF constraints. Unsupported DTD constructs are handled conservatively; this is not a claim of complete xmerl DTD feature parity. For callers that do not need parent paths, positions, namespace metadata, or structured attribute nodes, metadata: false skips constructing those optional fields and leaves them at their empty defaults (the attributes map is still populated). Namespace-conformant queries should retain the default metadata setting. The same option applies to node/event streams. External subsets and entities can be supplied through a binary-native fetch_fun callback returning {ok, {:string, data}, state} or {ok, {:file, path}, state}.

Chunked streams support ordinary XML plus internal general and parameter entities when allow_entities: true. External DTDs and external entities can also be resolved by supplying external_base_dir: path alongside allow_entities: true; the base directory is explicit because an arbitrary enumerable has no source filename. Without that option, chunked streams raise an argument error rather than silently changing entity semantics. Malformed chunked input raises ArgumentError when consumed, matching the query stream API's existing behavior for invalid binary input.

DTDs and untrusted XML

The dtd option supports :none, :all, :internal_only, and [only: entities]. Entity expansion still requires allow_entities: true. Xmlixir.parse/2 never performs filesystem I/O; use Xmlixir.file/2 for relative external resources, or provide fetch_fun to control resource loading. For a chunked enumerable, external resolution requires an explicit external_base_dir. DTD validation is available on the binary path; chunked streams have narrower validation and entity behavior.

Do not enable external entities for untrusted XML unless resource access is deliberately restricted. The core parser keeps names as binaries. Tag streams return existing atoms for known names and preserve unknown names as binaries; pass a finite atom vocabulary when atom keys are required for hostile input.

Limitations

Xmlixir.XPath implements XPath 1.0 paths, axes, node tests, predicates, namespaces, variables, all core functions, arithmetic, comparisons, and XPath node-set/string/number conversions. Custom host-defined functions are outside XPath 1.0 and are not provided. Xmlixir.file/2 is the direct file-parsing API; File.stream!/3 is the lazy-streaming input API.

DTD validation is intentionally partial and does not provide complete xmerl DTD feature parity. Chunked parsing may be expensive when a large token is split across many very small chunks. The stream!/2 and stream_tags!/3 functions are currently compatibility aliases for their non-bang counterparts, and scan/2 has a different success return shape than parse/2. Entity- and DTD-heavy inputs also require additional validation and allocation compared with ordinary XML.

Installation

If available in Hex, the package can be installed by adding xmlixir to your list of dependencies in mix.exs:

def deps do
[
{:xmlixir, "~> 0.1.0"}
]
end

Documentation can be generated with ExDoc and published on HexDocs. Once published, the docs can be found at https://hexdocs.pm/xmlixir.

License and contributing

Xmlixir is released under the Apache License 2.0. Contributions are welcome.

OTP test corpus

The copied OTP xmerl test sources, archives, and XML conformance fixtures are under test/otp/xmerl/. The fixtures include the W3C XML Test Suite profile catalog and the 2013 W3C XML Test Suite archive. test/otp_conformance_test.exs runs the extracted XML well-formedness cases as native ExUnit tests, while test/w3c_conformance_test.exs verifies the W3C catalog and archive metadata. The upstream suite is maintained at https://www.w3.org/XML/Test/.

test/feature_parity_test.exs is the local normalized parity harness. It uses a shared corpus to compare Xmlixir and xmerl for parse results and XPath output, then tests entities/DTDs, namespaces, query modifiers and mappings, streaming, and error categories separately. The W3C XML Test Suite and XQuery Test Suite remain external conformance sources; the latter targets XQuery/XPath semantics broader than Xmlixir's documented XPath subset.

Benchmark

Install the development dependency and run:

mix deps.get
mix run benchmarks/xml_parser.exs

The benchmark compares binary Xmlixir.parse/1 with OTP :xmerl_scan, both including the binary-to-charlist conversion and with that conversion excluded. This measures parser throughput and memory, not feature equivalence: xmerl also provides XPath-oriented tooling and an established record-based representation, while Xmlixir returns a binary-native DOM with its own public structs and helpers. It defaults to the sibling OTP checkout at ../otp; set XMERL_EBIN when using a different OTP build, for example:

XMERL_EBIN=/path/to/otp/lib/xmerl/ebin mix run benchmarks/xml_parser.exs