Pagination
Unbreakable blocks taller than the page, widows and orphans, headings at the foot of a page, orphaned continuation pages, and hyphenation across a break. Half-empty pages only when you ask for them.
When you generate PDFs from HTML, the mistakes only appear after pagination: a heading stranded at the foot of a page, one line of a paragraph carried alone onto the next, a block that fits on no page at all. breaklint measures every page after the break and reports, for each finding, the measured value next to the threshold it failed.
Film · 0:39 · Instrumental music, no voice-over
npx breaklint --demo
That command needs no browser and no configuration. It ends with exit 1, because the bundled fixture contains findings on purpose — a demo that ends 0 never shows you what a finding looks like. The block below shows one complete finding from the bundled demo; only the long lines are wrapped to fit the reading width. The rule chain running in it is the real one; the page it judges is a hand-written snapshot, built so that every rule path is reachable without a browser — and the report states which kind of fixture it read in its own source field.
error svg/text-overflows-viewport page 5
measured 72 px; threshold 0 px (uncalibrated)
detail This text extends 72.00 px beyond the
SVG viewport and is not drawn.
Coordinates are normalised through
getScreenCTM().
remedy Text rendered inside an SVG extends
outside the SVG viewport bounds and is
clipped. Enlarge the SVG 'viewBox' or
its width/height, or adjust the <text>
coordinates ('x', 'y', 'text-anchor').
'overflow: visible' on the container
also clears the finding, but it does not
move the text: the viewport then no
longer clips, the target becomes
non-applicable and this rule stops
measuring it. Use that only where the
overflow is intended.
untested: no trigger/remedied pair in
this package shows this advice removing
this finding
source unknown (node produced by the paginator)
render unknown (no evidence produced)
Prose linters never see a page. Typesetting systems see the page but do not know German typographic convention. For documents generated from HTML to PDF, neither exists.
The reason for measuring the rendered page rather than the source is a case that source-level checking did not catch: a chart passed an XML validity check and a geometry check, and still came out of the renderer with colliding labels. Whether a page works is decided after pagination — in the renderer, with the fonts that were actually available.
Thirteen rules, twelve of them active by default. Two can fail a build; ten more are advisory until you ask for more with --fail-on warn. The thirteenth — layout/half-empty-page — is experimental, never moves an exit code, and since 0.6.0 is not active in the default profile: it fired on 37 of 40 documents of a corpus built to exercise it. It stays registered and reachable with --profile strict or --only. That split is not caution, it is the burden of proof: only two rules compare directly measured quantities against a structural boundary.
Unbreakable blocks taller than the page, widows and orphans, headings at the foot of a page, orphaned continuation pages, and hyphenation across a break. Half-empty pages only when you ask for them.
Text that runs past the viewport. Clip/mask removal and device-pixel collision are not released rules yet: production has no safe ink pass for them.
A hyphen where a dash belongs, typewriter quotes in typeset prose, short last lines, and word gaps torn open by justification.
file: URIs and build-machine paths left behind in the delivered document.
All thirteen released rules are still labelled uncalibrated. The two ink definitions remain explicitly unreleased M3 research. The technical foundation for later calibration is now in place: runs in the real renderer, a rights- and privacy-reviewed process pilot with three real documents, separate data for development, threshold tuning and the held-back final test, plus blind review packets and controls that reject invalid evidence.
That is not empirical calibration yet. It still requires a sufficiently large corpus of real documents, two independent blind human judgements, a documented resolution when they disagree, an externally verified final test set that has remained unchanged, and one final evaluation defined in advance. Until then, every rule remains calibrated: false.
The tests and renderer evidence show that breaklint performs the measurements of the thirteen released rules and catches known counterexamples. A mandatory real-document gate binds an unchanged, rights- and privacy-reviewed third-party Project Gutenberg HTML file plus a first-party packaging case, requires exact page and measured-rule counts, and rejects infrastructure drift; that proves robustness, not calibration. How often heuristic boundaries are correct across real documents remains unknown.
Use veraPDF or pdfcpu for that.
Use a visual-diff tool for pixel or image differences. breaklint's conservative document-repair comparison instead requires a complete, verified compatible target.
Use vale or typopo for that.
PDF is produced and rasterised here to make evidence, never read as input.
Current release: 0.8.0; 0.7.0 is its immediate predecessor. The current version ships thirteen public rules instead of fifteen definitions; twelve are active by default. All rules remain uncalibrated. A failed image becomes the non-fatal image-content-unavailable diagnostic only when positive authored width and height attributes exactly equal the measured browser box; replacement text and CSS drift are insufficient.
New in 0.8.0: each document has a default time budget of 600000 ms (ten minutes). Set documentTimeoutMs in breaklint.config.json, or override it for a run with --document-timeout-ms <n>. The value must be an integer from 30000 to 1800000 ms. Invalid values exit 2. A run that reaches the budget exits 3 and names the effective value. The budget limits duration; it does not speed up pagination or reduce the memory use of large documents.
0.8.0 refuses two known paths to silent content loss before measurement: a body with computed column-count or column-width (including column-count: 1), and a nonempty heading with display: contents. Both exit 3 instead of reporting clean. Repeated SVG text in ordinary page content is assigned by occurrence. Inline SVG text in running margin boxes remains an exception: when page copies share a target ID across pages, the check ends in checker-crashed (exit 3). This does not establish a general PDF content-completeness check.
For a narrow class of visible strokes around SVG text, breaklint proves that a conservative envelope lies strictly inside the SVG viewport. This also requires supported stroke properties and transforms. If the envelope touches the viewport edge, the target remains unmeasured and svg/text-overflows-viewport exits 4 because it requires full coverage. This proves containment, not exact painted bounds or an overflow finding. Masks and filters remain unmeasured.
0.7.0 moves the snapshot to version 5; Report 5, agent context 2 and Configuration 1 stay. Several rules judge unchanged documents differently: findings can appear or disappear, and anyone who tracks page findings by their fingerprint needs a new baseline. layout/unbreakable-block-too-tall, one of the two rules that fail a build by default, sums a block's height over its fragments once the paginator has split it into three or more pieces; a too-tall block split into exactly two pieces is still not reported. A split block whose fragments cannot be joined by their source is no longer measured at its first fragment; the rule declines it and counts that against coverage, so such a run can end with exit 4 where it used to end clean. Content in the page margins, such as running headers and footers from position: running(...), no longer counts as text flow and is judged by no rule. Whether a page boundary was forced is now read from the paginator's own break decision, including in regions with named pages. A report written to a pipe arrives whole; through 0.6.0 it could be cut at a multiple of the pipe buffer, on Linux usually at 64 KiB. If it cannot be written completely, the run ends with exit 3. For produced documents, breaklint names an editable original location only when a host-controlled producer has captured the original bytes and proved the source relation. The same canonical JSON produces an offline human view and bounded context for AI systems. A repair comparison marks a finding resolved only when the target is verified and identical, and the revision, rules, configuration, renderer environment, fonts and resources are compatible; unknown sources and partial coverage remain explicit.
The GitHub Action introduced in 0.7.0 can be included as uses: godarg/breaklint@v0.8.0. It installs the published npm release, checks the given HTML paths in one run, writes SARIF, JUnit and a Markdown summary from the one canonical JSON report, and ends the step with breaklint's exit code; exits 2, 3 and 4 always fail it. docs/ci-recipe.md has a workflow to copy. The Action is tested on ubuntu-latest; macOS runners are untested. It needs bash 4 or newer; with macOS's /bin/bash 3.2 it stops with exit 3.
Separately, the public checkPage API checks an already prepared Playwright Page; it does not take over navigation, authentication or network policy. In our own operations it checks 45 states in an internal dashboard's test suite: nine views at five widths. The command-line path still accepts only .html/.htm; standalone SVG, PDF and Markdown files end with usage Exit 2. Node 22.13 or newer, on macOS or Linux. A run over real documents additionally needs a Chromium-based browser and pagedjs@0.4.3; pdfjs-dist rasterises the produced PDF to bind evidence to findings. Windows is not supported — process termination here rests on POSIX process groups. The measurement chain is exercised on Linux; process termination and profile cleanup have empirical evidence on macOS only.
npm i -D breaklint
npm i -D puppeteer-core@^25.8.0 pagedjs@0.4.3 pdfjs-dist@6.2.108
breaklint executes foreign HTML, scripts included. Offline mode intercepts page requests but is not an egress control: it does not cover WebSocket, WebTransport or WebRTC connections a document opens, nor browser traffic through DNS over HTTPS or the component updater. Since 0.8.0, breaklint refuses to launch the renderer when PUPPETEER_DANGEROUS_NO_SANDBOX or PUPPETEER_TEST_EXPERIMENTAL_CHROME_FEATURES is present; it removes CHROME_EXTRA_FLAGS from Chrome's environment. Other wrapper variables have not been generally audited. Untrusted documents belong in a container or network namespace with no egress.
Still open: Documents with Paged.js footnotes exit 3 before the rules run. Individual rules decline to measure multi-column blocks. An oversized unbreakable block split into exactly two fragments still goes unreported.
breaklint uses no language model at check time. It performs no inference and contains no API client.
Those are exactly the cases we are looking for. Anyone can test public or purpose-built HTML — no application, selection or prior experience required. The short guide takes about ten minutes from installation to a structured report.
Use only material you have the right to publish. Your GitHub identity, report and reproduction are public; confidential documents, personal data and security findings do not belong in the form. Structured intake is automated, but submitted content is never executed.
This is additional product QA. Public testing is not a blind study and does not replace the empirical calibration that the rules still need.
breaklint is built at Dargel Solutions and is free to use under the MIT licence, including commercially. Parts of the repository were written with the help of large language models; the rule set, the thresholds and their sources were chosen by a person. Every rule that cites the German orthography ruleset was checked against the published text of that ruleset, not against a model's summary of it.
Generative AI is used to create and maintain this page. Factual and visual publication approval remains with a person.
breaklint is listed in the PeerPush directory.