Our ProcessWire module works on the finished page: it drops the whitespace that only the templates need, applies typography and runs replacements. In 0.4.0 cleanup is rewritten so that it cannot change an attribute value — and so that any failure sends the page out exactly as it came in.
OutputTransformer hooks Page::render and works on the rendered HTML, not on the templates. The templates stay readable — indentation, newlines, comments and all. The browser gets what a browser needs.
We have run it on our own sites, this one included, since 2025. The code is now open: github.com/theoretic/OutputTransformer. Questions and discussion are in the ProcessWire forum topic.
Three transformations
Cleanup. In a template, HTML is written for a human. A browser needs far less of it: cleanup drops the extra spaces, newlines and comments without changing what the page means.
Typography. Dashes, proper nested quotes, non-breaking spaces. Russian goes through atispro/emt-php8, every other language through mundschenk-at/php-typography.
Replacements. A list of find-and-replace pairs, one per line, through mb_ereg_replace. For the small things that are not worth a template edit.
Each transformation is its own Page::render hook, with its own priority and its own list of excluded templates. Any one of them is switched off on its own: clear its priority.
Why 0.4.0 is a rewrite
Up to 0.3.0 cleanup treated the page as plain text: the same rules ran over tags, text and scripts alike. On most pages that goes unnoticed — until a place turns up where a space is the data.
That is how it surfaced. An admin form sent value=" | " and got "| " back: the title separator changed itself on every save. The same rules turned x < y in running text into the start of a tag, and joined script lines into their own // comments, commenting out everything that followed.
In 0.4.0 the page is split into segments first — comments, tags, text, CDATA and raw elements — and each kind has its own rules. An attribute value is never changed: only class, rel, srcset and sizes are trimmed and collapsed as lists, and style loses its trailing semicolon. The contents of script, pre, textarea and CDATA are not touched at all.
What is left of the markup
<!-- before -->
<ul class=" nav main ">
<li>
<a href="/about/" title=" About us ">About</a>
</li>
</ul>
<!-- after -->
<ul class="nav main"><li><a href=/about/ title=" About us ">About</a></li></ul> The class list is trimmed and collapsed, quotes are dropped where HTML allows it, whitespace between tags is gone. The title value is byte-identical: the spaces around it are data. The comments are kept here for the reader: cleanup removes them, except conditional ones and noindex.
What it actually saves
Being honest about the numbers
Measured on this site on 18 September 2026: each page fetched twice, with the module on and with it off, HTML only, gzip -9 for the compressed figure.
The byte win is modest: 4–10% before compression, about 4% after. That alone is no reason to install anything — gzip and brotli have already done the heavy lifting. The point is elsewhere: the page goes out without the debris, it reads well in View Source, and typography and replacements reach the browser on their own, without a change in every template.
Keeping a piece untouched
Wrappers. The tags <no-cleanup>, <no-typografy> and <no-replace> keep their contents out of that transformation. The wrapper tags themselves are removed from the output.
Excluded templates. Each transformation has its own list of templates whose pages it leaves alone.
Your own patterns. “Keep untouched by…” takes one PCRE per line, delimiters included. Whatever matches stays exactly as it is. A pattern that does not compile is skipped, logged, and reported when the settings are saved.
Everything protected is swapped for private-use characters (U+F8E0–U+F8F1) while the transformation runs. If anything goes wrong — a regex error, an exception, a placeholder lost or duplicated, a page that already contains those characters — the whole transformation is skipped for that page, the page goes out as it came in, and the reason is logged to outputtransformer. No corrupted page, no blank page.
The two escape hatches
<!-- in a template: cleanup would collapse the newlines inside -->
<no-cleanup><pre><code><?=htmlspecialchars($block->code)?></code></pre></no-cleanup>
<!-- in the settings, Keep untouched by cleanup: one PCRE per line -->
~<input\b[^>]*name="?title_separator[^>]*>~i The first is for your own markup, the second for what a wrapper cannot reach: another module, an admin field, a third-party widget.