Behaviour for extracting text from files of any format.
Arcana ships native handling for plain text and markdown, and a Poppler-backed PDF parser. Any other format works by registering a parser for its extension:
config :arcana,
file_parsers: %{
".docx" => {MyApp.DocxParser, endpoint: "http://extract.internal"}
}A single parser can also handle everything Arcana doesn't natively support — useful for extraction services that cover many formats:
config :arcana, fallback_parser: {MyApp.ExtractionService, []}Exact extension matches win over the fallback, and the built-in
handling for .txt/.md/.pdf can be overridden by registering
those extensions explicitly. Registering one as false disables it
outright, fallback included.
Implementing a parser
defmodule MyApp.DocxParser do
@behaviour Arcana.FileParser
@impl true
def parse(path_or_binary, opts) do
{:ok, extracted_text}
end
# Optional: accept binary content, enabling `Arcana.ingest_binary/2`
# for this format (default: false, path only). Path-only parsers
# make ingest_binary/2 return {:error, {:binary_unsupported, mod}},
# unless they are unavailable too — then unavailability is what
# gets reported, since a path retry would fail just the same.
@impl true
def supports_binary?, do: true
# Optional: report whether the parser can run right now, e.g. a
# required CLI is installed (default: true). A parser reporting
# `false` is not invoked at all: parsing returns
# `{:error, {:parser_unavailable, module}}`.
@impl true
def available?, do: true
endThe gate lives in parse/3, so every dispatch path gets it: this
module, Arcana.FileParser.PDF.parse/3, and everything routed through
Arcana.Parser. The one exception to the error shape is the built-in
Arcana.FileParser.PDF.Poppler, which reports :poppler_not_available
for backwards compatibility — see unavailable_reason/1.
Positional metadata
parse/2 may return {:ok, text, meta} instead of {:ok, text}.
Arcana understands %{pages: [%{number: 1, start: 0, end: 1234}, ...]},
where the offsets are byte positions in the returned text. Parsers
that don't know about pages return {:ok, text} as before.
During ingestion Arcana intersects those ranges with each chunk's own
byte range and stores the result as "page_start"/"page_end" in the
chunk's metadata, where Arcana.search/2 surfaces it:
{:ok, _doc} = Arcana.ingest_file("manual.pdf", repo: Repo)
[result | _] = Arcana.search("warranty", repo: Repo)
result.metadata["page_start"]
#=> 12A chunk straddling a page break reports the two different pages.
Summary
Functions
Whether a {module, opts} parser can run right now.
Invokes a {module, opts} parser, normalizing its return to
{:ok, text, meta}.
Whether a {module, opts} parser accepts binary content.
The error reported for a parser that can't run.
Types
@type meta() :: map()
Callbacks
Functions
Whether a {module, opts} parser can run right now.
Defaults to true for parsers that don't implement available?/0, but
a module that can't be loaded (a typo in config) or doesn't implement
parse/2 is never available: reporting true there only postpones the
failure to an UndefinedFunctionError at parse time.
Invokes a {module, opts} parser, normalizing its return to
{:ok, text, meta}.
Parsers predating positional metadata return {:ok, text}; those come
back here as {:ok, text, %{}} so callers only handle one shape.
Metadata that isn't a map, or whose :pages aren't
%{number: _, start: integer, end: integer} maps, comes back as
{:error, {:invalid_parser_metadata, module, meta}}. Ingestion reads
those offsets long after the document row is inserted, so catching the
shape here is what keeps a misbehaving parser from raising out of
library internals and orphaning a :processing document.
A parser reporting itself unavailable (available?/0) is never
invoked: this returns {:error, unavailable_reason(parser)} instead.
The check lives here rather than in a caller so that every route into
a parser carries the same guarantee.
Whether a {module, opts} parser accepts binary content.
Defaults to false: a parser that doesn't say otherwise is assumed to
need a file path.