A different starting point, a different split method
For scanned stacks of paper documents, a physical separator sheet between the pages helps — more on that at Create a barcode separator sheet. A bundled file exported from a system has no paper and no separator sheet though, just one continuous PDF. It can still be split — using text markers.
The key difference: PDFs generated digitally — from accounting software, payroll systems, ERP systems — almost always contain real, searchable text, not images. StackSplitter can search that text and start a new file wherever a pattern you define shows up.
Upload PDF
How should the stack be split?
Case-sensitive — "Invoice" does not match "invoice". Marker pages are kept as the first content page of the new document.
Analyse stack
Page overview
Magenta line = split · +/× between pages = add/remove split · Click page = mark for removal
Detected documents
Edit file names, then download as ZIP.
What makes a good marker
Invoice number prefix. Almost every system prints a recurring word directly before the running number on each invoice, e.g. “Invoice No.” or “Rechnung Nr.” — entering exactly that word as the marker reliably splits at every invoice start.
“Page 1 of”. Systems that reset the page counter per document often print “Page 1 of 3” or similar on the first page of each sub-document — a reliable marker regardless of content.
Combining two markers (AND logic). If a single word occurs too often (e.g. “Invoice” also appears in body text), a second pattern can be added that only occurs on the first page of each document together with the first — reducing mis-splits.
A marker that happens to reappear somewhere mid-document is not a problem: StackSplitter only starts a new document at a match on what it has already identified as the first page of a block — see the preview grid, which shows every detected split and lets you correct it manually.
What if the file was scanned instead of generated digitally?
Text markers only work with searchable text. A scanned stack without OCR has none — every page there is an image, not text. In that case, physical separation is the only option: insert a separator sheet between the papers before scanning (see above), or, if your scanner offers OCR, enable it before export.
Documents of different lengths
Sub-documents don't need to have the same number of pages — StackSplitter detects the start of each new document by its marker, not by a fixed page count. A two-page and a five-page invoice in the same bundled PDF are both correctly recognised as their own complete file, as long as each starts with a match.
Frequently asked questions
How does the tool recognise the start of a new document?
By a text pattern you define that appears on the first page of each sub-document — such as the invoice number label or “Page 1 of”. Two patterns can be combined with AND logic.
What if the documents have different lengths?
No problem. StackSplitter splits at the detected marker, not at a fixed page count — each document can have as many pages as it actually has.
Does this work with payslips?
Yes, as long as the file contains searchable text — which is practically always the case for digitally generated payslips. A recurring field like the employee number or “Statement for” works well as a marker.
Does it work on scanned files?
Only if OCR has already been applied. Plain scans without OCR contain no searchable text — a physical separator sheet at scan time helps there instead.