Split PDF by text

Start a new file whenever a piece of text changes.

Drag & drop your files here

Supported: PDF

Files are handled privately and removed automatically. Interface preview - processing is not enabled yet.

About Split PDF by text

Ideal for statement runs and invoice batches: mark the area where the customer name or invoice number sits, and a new document begins each time that value changes.

  • Automatic invoice and payslip splitting
  • Names files from the detected text
  • Handles hundreds of documents

How to split pdf by text

  1. 1

    Upload the batch

    One long PDF holding many documents.

  2. 2

    Select the text area

    Draw a box over the value that identifies each document.

  3. 3

    Split automatically

    Each change of value starts a new file.

A practical guide to split pdf by text

Bursting a statement or invoice run

This tool exists for the moment a system exports thousands of statements as one enormous PDF that then has to become one file per customer. Draw a box over the field that changes per recipient, typically the account number or name, and each change starts a fresh document. The key to a clean result is choosing a value that is truly unique to each document and appears reliably on its first page. A running date or a page number would split in the wrong places, so pick the identifier a human would use to tell two statements apart, and the machine will follow the same logic.

Naming the output from the page

Beyond just cutting the batch, the detected text can name each resulting file, so a statement run comes out as the customer names or account numbers rather than anonymous part numbers. That turns an unmanageable pile into a set you can search and route immediately. Before running the full batch, split a small sample and check the filenames, since a selection box that catches an extra word or a stray character will carry that into every name. Adjusting the box on a few pages first is far quicker than renaming hundreds of files afterward, and it confirms the split points are landing correctly at the same time.

Preparing scans so the text is readable

Split by text depends on the document containing real, selectable text, which digitally generated statements and invoices already have. Scanned paper does not, so a batch of scans must go through the OCR tool first to add a readable text layer; without it, the selection box finds nothing to compare and the split fails. After OCR, confirm the identifying value is being recognized cleanly, because a smudged or low-resolution scan can misread characters and cause a document to split in the wrong spot or not at all. When a page genuinely lacks the value, it is attached to the preceding document rather than dropped, so nothing goes missing.

Frequently asked questions

How does the tool know where one document ends and the next begins?
You draw a box over the value that identifies each document, such as an account number or a customer name, and a new file begins every time the text inside that box changes. As long as that value is consistent and repeats on the first page of each document, the batch splits automatically.
What if the identifying text sits in a slightly different spot on some pages?
Draw the selection box a little larger than the text so minor shifts still fall inside it. If the position varies wildly between documents, the detection can miss, so it works best on batches produced by the same template, like a statement or payslip run where the layout is fixed.
Can it split invoices where each invoice is several pages long?
Yes. Because a new file starts only when the tracked value changes, all the pages sharing one invoice number stay together, and the split falls at the point where the number becomes the next invoice's. Multi-page documents are exactly what this tool is built for.
How many documents can I split out of one file?
There is no fixed cap; the tool is designed for large batches and routinely separates hundreds of statements or invoices from a single long PDF. Bigger files take a little longer to scan for the text, and everything is processed for the session then deleted automatically.