Open source · MIT · command line

breaklint checks the page that will actually be printed.

When you generate PDFs from HTML, the mistakes only appear after pagination: a heading stranded at the foot of a page, one line of a paragraph carried alone onto the next, a block that fits on no page at all. breaklint measures every page after the break and reports, for each finding, the measured value next to the threshold it failed.

Source on GitHub Package on npm

Film · 0:39 · Instrumental music, no voice-over

What the film shows
  1. 0:00 Perfect in the browser. Broken on page 47. · Heading alone at the bottom · Page 48: two lines. That's it. · Table fits on no page · Client: “What happened on page 47?” · Fri 17:58 · off to the printer · Chart label cut off · Did anyone check the PDF? · It looked fine in the browser. · Page: 1 → 212 / 212
  2. 0:06 212 pages. Check them by hand? Every single build?
  3. 0:10 One command. Every page measured.
  4. 0:13 Open source · MIT · breaklint · Layout checks for HTML to PDF, measured after pagination.
  5. 0:17 It tells you where it breaks. · Real output
  6. 0:21 Page 1: heading stranded. · Sample · page 1
  7. 0:25 Page 4: 5 % filled. · Sample · page 4
  8. 0:29 Every finding shows its numbers. · Widows & orphans · Stranded headings · Blocks too tall to fit · SVG text overflow
  9. 0:33 Catch it before your client does. · Free · open source · MIT · npx breaklint --demo

Try it in one line

npx breaklint --demo

That command needs no browser and no configuration. It ends with exit 1, because the bundled fixture contains findings on purpose — a demo that ends 0 never shows you what a finding looks like. The block below shows one complete finding from the bundled demo; only the long lines are wrapped to fit the reading width. The rule chain running in it is the real one; the page it judges is a hand-written snapshot, built so that every rule path is reachable without a browser — and the report states which kind of fixture it read in its own source field.

error svg/text-overflows-viewport  page 5
  measured   72 px; threshold 0 px (uncalibrated)
  detail     This text extends 72.00 px beyond the
             SVG viewport and is not drawn.
             Coordinates are normalised through
             getScreenCTM().
  remedy     Text rendered inside an SVG extends
             outside the SVG viewport bounds and is
             clipped. Enlarge the SVG 'viewBox' or
             its width/height, or adjust the <text>
             coordinates ('x', 'y', 'text-anchor').
             'overflow: visible' on the container
             also clears the finding, but it does not
             move the text: the viewport then no
             longer clips, the target becomes
             non-applicable and this rule stops
             measuring it. Use that only where the
             overflow is intended.
             untested: no trigger/remedied pair in
             this package shows this advice removing
             this finding
  source     unknown (node produced by the paginator)
  render     unknown (no evidence produced)

Why this exists

Prose linters never see a page. Typesetting systems see the page but do not know German typographic convention. For documents generated from HTML to PDF, neither exists.

The reason for measuring the rendered page rather than the source is a case that source-level checking did not catch: a chart passed an XML validity check and a geometry check, and still came out of the renderer with colliding labels. Whether a page works is decided after pagination — in the renderer, with the fonts that were actually available.

What is checked

Thirteen rules, twelve of them active by default. Two can fail a build; ten more are advisory until you ask for more with --fail-on warn. The thirteenth — layout/half-empty-page — is experimental, never moves an exit code, and since 0.6.0 is not active in the default profile: it fired on 37 of 40 documents of a corpus built to exercise it. It stays registered and reachable with --profile strict or --only. That split is not caution, it is the burden of proof: only two rules compare directly measured quantities against a structural boundary.

Pagination

Unbreakable blocks taller than the page, widows and orphans, headings at the foot of a page, orphaned continuation pages, and hyphenation across a break. Half-empty pages only when you ask for them.

SVG

Text that runs past the viewport. Clip/mask removal and device-pixel collision are not released rules yet: production has no safe ink pass for them.

Typography

A hyphen where a dash belongs, typewriter quotes in typeset prose, short last lines, and word gaps torn open by justification.

Artefacts

file: URIs and build-machine paths left behind in the delivered document.

What remains uncalibrated

All thirteen released rules are still labelled uncalibrated. The two ink definitions remain explicitly unreleased M3 research. The technical foundation for later calibration is now in place: runs in the real renderer, a rights- and privacy-reviewed process pilot with three real documents, separate data for development, threshold tuning and the held-back final test, plus blind review packets and controls that reject invalid evidence.

That is not empirical calibration yet. It still requires a sufficiently large corpus of real documents, two independent blind human judgements, a documented resolution when they disagree, an externally verified final test set that has remained unchanged, and one final evaluation defined in advance. Until then, every rule remains calibrated: false.

The tests and renderer evidence show that breaklint performs the measurements of the thirteen released rules and catches known counterexamples. A mandatory real-document gate binds an unchanged, rights- and privacy-reviewed third-party Project Gutenberg HTML file plus a first-party packaging case, requires exact page and measured-rule counts, and rejects infrastructure drift; that proves robustness, not calibration. How often heuristic boundaries are correct across real documents remains unknown.

What it does not do

No PDF standard conformity

Use veraPDF or pdfcpu for that.

No visual image comparison

Use a visual-diff tool for pixel or image differences. breaklint's conservative document-repair comparison instead requires a complete, verified compatible target.

No prose or style checking

Use vale or typopo for that.

PDF is not an input format

PDF is produced and rasterised here to make evidence, never read as input.

Requirements

Current release: 0.8.0; 0.7.0 is its immediate predecessor. The current version ships thirteen public rules instead of fifteen definitions; twelve are active by default. All rules remain uncalibrated. A failed image becomes the non-fatal image-content-unavailable diagnostic only when positive authored width and height attributes exactly equal the measured browser box; replacement text and CSS drift are insufficient.

New in 0.8.0: each document has a default time budget of 600000 ms (ten minutes). Set documentTimeoutMs in breaklint.config.json, or override it for a run with --document-timeout-ms <n>. The value must be an integer from 30000 to 1800000 ms. Invalid values exit 2. A run that reaches the budget exits 3 and names the effective value. The budget limits duration; it does not speed up pagination or reduce the memory use of large documents.

0.8.0 refuses two known paths to silent content loss before measurement: a body with computed column-count or column-width (including column-count: 1), and a nonempty heading with display: contents. Both exit 3 instead of reporting clean. Repeated SVG text in ordinary page content is assigned by occurrence. Inline SVG text in running margin boxes remains an exception: when page copies share a target ID across pages, the check ends in checker-crashed (exit 3). This does not establish a general PDF content-completeness check.

For a narrow class of visible strokes around SVG text, breaklint proves that a conservative envelope lies strictly inside the SVG viewport. This also requires supported stroke properties and transforms. If the envelope touches the viewport edge, the target remains unmeasured and svg/text-overflows-viewport exits 4 because it requires full coverage. This proves containment, not exact painted bounds or an overflow finding. Masks and filters remain unmeasured.

0.7.0 moves the snapshot to version 5; Report 5, agent context 2 and Configuration 1 stay. Several rules judge unchanged documents differently: findings can appear or disappear, and anyone who tracks page findings by their fingerprint needs a new baseline. layout/unbreakable-block-too-tall, one of the two rules that fail a build by default, sums a block's height over its fragments once the paginator has split it into three or more pieces; a too-tall block split into exactly two pieces is still not reported. A split block whose fragments cannot be joined by their source is no longer measured at its first fragment; the rule declines it and counts that against coverage, so such a run can end with exit 4 where it used to end clean. Content in the page margins, such as running headers and footers from position: running(...), no longer counts as text flow and is judged by no rule. Whether a page boundary was forced is now read from the paginator's own break decision, including in regions with named pages. A report written to a pipe arrives whole; through 0.6.0 it could be cut at a multiple of the pipe buffer, on Linux usually at 64 KiB. If it cannot be written completely, the run ends with exit 3. For produced documents, breaklint names an editable original location only when a host-controlled producer has captured the original bytes and proved the source relation. The same canonical JSON produces an offline human view and bounded context for AI systems. A repair comparison marks a finding resolved only when the target is verified and identical, and the revision, rules, configuration, renderer environment, fonts and resources are compatible; unknown sources and partial coverage remain explicit.

The GitHub Action introduced in 0.7.0 can be included as uses: godarg/breaklint@v0.8.0. It installs the published npm release, checks the given HTML paths in one run, writes SARIF, JUnit and a Markdown summary from the one canonical JSON report, and ends the step with breaklint's exit code; exits 2, 3 and 4 always fail it. docs/ci-recipe.md has a workflow to copy. The Action is tested on ubuntu-latest; macOS runners are untested. It needs bash 4 or newer; with macOS's /bin/bash 3.2 it stops with exit 3.

Separately, the public checkPage API checks an already prepared Playwright Page; it does not take over navigation, authentication or network policy. In our own operations it checks 45 states in an internal dashboard's test suite: nine views at five widths. The command-line path still accepts only .html/.htm; standalone SVG, PDF and Markdown files end with usage Exit 2. Node 22.13 or newer, on macOS or Linux. A run over real documents additionally needs a Chromium-based browser and pagedjs@0.4.3; pdfjs-dist rasterises the produced PDF to bind evidence to findings. Windows is not supported — process termination here rests on POSIX process groups. The measurement chain is exercised on Linux; process termination and profile cleanup have empirical evidence on macOS only.

npm i -D breaklint
npm i -D puppeteer-core@^25.8.0 pagedjs@0.4.3 pdfjs-dist@6.2.108

breaklint executes foreign HTML, scripts included. Offline mode intercepts page requests but is not an egress control: it does not cover WebSocket, WebTransport or WebRTC connections a document opens, nor browser traffic through DNS over HTTPS or the component updater. Since 0.8.0, breaklint refuses to launch the renderer when PUPPETEER_DANGEROUS_NO_SANDBOX or PUPPETEER_TEST_EXPERIMENTAL_CHROME_FEATURES is present; it removes CHROME_EXTRA_FLAGS from Chrome's environment. Other wrapper variables have not been generally audited. Untrusted documents belong in a container or network namespace with no egress.

Still open: Documents with Paged.js footnotes exit 3 before the rules run. Individual rules decline to measure multi-column blocks. An oversized unbreakable block split into exactly two fragments still goes unreported.

breaklint uses no language model at check time. It performs no inference and contains no API client.

Open community test · voluntary · unpaid

Where do the measurement and a human eye disagree?

Those are exactly the cases we are looking for. Anyone can test public or purpose-built HTML — no application, selection or prior experience required. The short guide takes about ten minutes from installation to a structured report.

Use only material you have the right to publish. Your GitHub identity, report and reproduction are public; confidential documents, personal data and security findings do not belong in the form. Structured intake is automated, but submitted content is never executed.

This is additional product QA. Public testing is not a blind study and does not replace the empirical calibration that the rules still need.

Where it comes from

breaklint is built at Dargel Solutions and is free to use under the MIT licence, including commercially. Parts of the repository were written with the help of large language models; the rule set, the thresholds and their sources were chosen by a person. Every rule that cites the German orthography ruleset was checked against the published text of that ruleset, not against a model's summary of it.

Generative AI is used to create and maintain this page. Factual and visual publication approval remains with a person.

breaklint is listed in the PeerPush directory.