Importing existing HTML into a document.
This is the migration path. An application that already stores rich text as
HTML — in a :string column, from another editor, from a feed — has to get
that content into the schema before any of the rest of Coelho applies to
it:
{:ok, document, _warnings} = Coelho.HTML.from_html(post.body_html)
post
|> Ecto.Changeset.change(%{body: document})
|> Repo.update()Requires Floki, declared as an optional dependency: the parser is only needed on this path, and the document core has no dependencies at all.
What the import does with markup it does not know
Importing foreign HTML is not validation, and failing on the first surprise would make it useless. The rules are deliberate:
- an element the schema has no rule for is transparent — it
disappears and its children take its place, so a
<div>wrapper or a<span class="fancy">does not cost you the text inside it <script>,<style>,<head>,<template>and<noscript>are dropped with their content- an element whose attributes fail the schema's validators — an
<img>with nosrc, an<a href="javascript:…">— is treated as unknown, so the link text survives while the link does not - inline content in a place that demands blocks is wrapped in the
schema's first suitable block, which is how a bare
Helloat the top level becomes a paragraph - whitespace is collapsed as HTML collapses it, and runs of whitespace between blocks are dropped
What comes out is a validated, normalised document, or {:error, errors}
if what remained still does not fit the schema.
Teaching a schema to import
Each node and mark declares the tags it comes from, in the same spirit as
the parseDOM rules on the browser side:
paragraph: [content: "inline*", group: "block", parse: ["p"]]
heading: [
content: "inline*",
group: "block",
parse: [{"h1", %{"level" => 1}}, {"h2", %{"level" => 2}}]
]
link: [parse: [{"a", &Coelho.HTML.take(&1, ~w(href title))}]]A rule is a tag, optionally paired with the attributes to give the node: a
fixed map, a function of the element's HTML attributes, or a function of
those and the element's text. Rules are tried in declaration order, nodes
before marks, and a rule whose attributes fail the schema does not match —
which is how <span data-user-id="7"> becomes a mention while every other
span stays a span.
Summary
Functions
Converts HTML into a validated document, and says what it left behind.
Keeps the named HTML attributes, dropping those the element does not carry.
Types
@type warning() :: %{ kind: :unknown_element | :rejected_element, tag: String.t(), count: pos_integer() } | %{ kind: :dropped_attribute, tag: String.t(), attribute: String.t(), count: pos_integer() }
Functions
@spec from_html(String.t(), Coelho.Schema.t()) :: {:ok, map(), [warning()]} | {:error, term()}
Converts HTML into a validated document, and says what it left behind.
The import is lenient by design: markup the schema has no rule for is dropped and the text inside it is kept, because the alternative — refusing the paste — loses more. But silence about it is its own problem. Someone importing terms and conditions out of a word processor gets a document back with the tables gone and no way to know, and finds out from a reader.
So the third element of the result says what was removed:
{:ok, document, warnings} = Coelho.HTML.from_html(html, schema)
#=> warnings: [
#=> %{kind: :dropped_attribute, tag: "p", attribute: "style", count: 12},
#=> %{kind: :rejected_element, tag: "a", count: 1},
#=> %{kind: :unknown_element, tag: "table", count: 3}
#=> ]:unknown_element— no node or mark in the schema parses that tag; the element is gone and its text was lifted into its parent:rejected_element— the schema has a rule for the tag, but the attributes the element carried failed their validators, so the rule did not apply.<a href="javascript:alert(1)">is this one: the text stays, the link does not:dropped_attribute— the element was kept, and this attribute is not one its rule extracts
Warnings are counts per tag, in a stable order, and they describe the HTML: an element the schema knows but that could not fit where it appeared — a list item outside a list — is repaired by the import rather than reported here. A mark refused because of where it sat, such as bold inside a code block, is not reported either.
Keeps the named HTML attributes, dropping those the element does not carry.
Handy inside a parse rule: {"a", &Coelho.HTML.take(&1, ~w(href title))}.