Paperless-ngx Tips: Matching, Workflows, Storage Paths, and Backups That Hold Up
← All posts
GuideOct 6, 2026· 10 min read

Paperless-ngx Tips: Matching, Workflows, Storage Paths, and Backups That Hold Up

If you searched for Paperless-ngx tips, your instance is probably running and the first few hundred documents are in. The hard part now is keeping it useful: tags that mean something, correspondents that get assigned without babysitting, files on disk you can still recognize, and a backup you have actually tested. The short version: give every new document an inbox tag, start matching rules on Exact or Any before you trust Auto, let consume subfolders and ASN barcodes carry information from the scanner, use workflows for anything conditional, set a storage path early, and run document_exporter on a schedule.

The tips below assume a current 3.x release and link to the official Paperless-ngx documentation for each setting. If you are still setting up the stack, start with our Paperless-ngx Docker Compose archive tutorial or the Synology Container Manager guide and come back once documents are flowing.

1. Make an inbox tag the first thing you create

Paperless-ngx lets you mark a tag as an inbox tag, and every newly consumed document gets it. The project’s own recommended workflow starts here for a reason: it turns “did anyone check this?” into a saved view.

A routine that works for small offices:

  • Create a tag called inbox, tick the inbox option, and pin a saved view filtered on it to the dashboard.
  • Open each new document, fix the title, correspondent, document type, and date, then remove inbox.
  • Add a todo tag for documents that need an action (pay, reply, sign) and a second saved view for it.

The inbox tag has a second benefit. The Auto matcher learns from documents that are not in the inbox, so a document you have not reviewed yet does not teach it the wrong label.

2. Pick matching algorithms on purpose

Every tag, correspondent, document type, and storage path has a match string and a matching algorithm. The choices are None, Any, All, Exact, Regular expression, Fuzzy match, and Auto. They behave very differently on OCR text:

Algorithm Good for Watch out for
Exact A supplier name or account number that is always printed the same way One OCR misread breaks the match
Any Several spellings of the same sender, for example "Bank of America" BofA Short or common words match far too much
All Requiring two signals, such as a company name plus “invoice” Word order is ignored, so it can still catch unrelated pages
Regular expression Structured identifiers like a customer number pattern Easy to write a pattern that matches every page
Fuzzy match Names that OCR mangles slightly Less predictable, test on real scans
Auto Collections with enough labeled history Needs training data and retrains periodically, not instantly

Two habits save a lot of cleanup. First, match on something only that sender prints, such as a VAT number, IBAN fragment, or customer ID, rather than the company name that also appears in your own outgoing letters. Second, when you use Any or All with a multi-word phrase, wrap it in double quotes so it is treated as one term, which the docs show with the Bank of America example.

Auto is worth switching on once a label has a few dozen correctly tagged documents behind it. Before that, it has little to learn from and will guess. The docs note it retrains on a schedule (hourly by default), so a fix you make now shows up in suggestions later, not on the next upload.

3. Let the consume folder carry metadata

The scanner already knows things Paperless has to guess. Two settings let you pass that knowledge along:

Set up one scan profile per destination folder on the multifunction printer and the person scanning does the first round of classification with one button press. Recursive watching must be on for the subfolder tags to work.

4. Use ASN barcodes if you keep any paper

The archive serial number (ASN) is the number that links a Paperless record to a physical binder. With PAPERLESS_CONSUMER_ENABLE_ASN_BARCODE=true, Paperless reads a barcode such as ASN00123 off the scan and sets the ASN automatically. The prefix defaults to ASN and can be changed with PAPERLESS_CONSUMER_ASN_BARCODE_PREFIX.

Print a sheet of sequential ASN labels, stick one on each original you must keep, and file the paper in a single binder ordered by ASN only. Anything without a label can be shredded after scanning, which is the approach the official workflow recommends.

If you batch scan, PAPERLESS_CONSUMER_ENABLE_BARCODES splits one long PDF into separate documents at barcode separator pages. Page splitting happens before the ASN is read, so both features work together.

5. Move conditional logic into workflows

Workflows replaced the older consumption templates in v2.3. Each workflow has triggers and actions, and they run in sort order. The current workflow triggers are Consumption Started, Document Added, Document Updated, and Scheduled.

The trigger you choose decides what you can filter on:

  • Consumption Started runs before the document exists, so you can filter by source (consume folder, API, mail), file name, file path, or mail rule, but not by content.
  • Document Added runs after OCR and matching, so you can filter on content, tags, document type, correspondent, and storage path.
  • Scheduled runs against a date (added, created, updated, or a date custom field) with an offset, which suits reminders such as “contract ends in 30 days”.

A few workflows that pay for themselves quickly:

  • Assign an owner and permissions based on the upload folder, using a Consumption Started trigger with a file path filter.
  • Tag everything from a specific mail rule as mail and give it a document type.
  • Send a webhook when a document is added with a given tag, so another system can pick it up. The webhook action supports JSON bodies and placeholders such as the correspondent and document type.

Remember that later assignment actions override earlier ones for single-value fields like document type, while tags, custom fields, and permissions are merged. If two workflows fight over a correspondent, check the sort order before anything else.

6. Decide your file layout before the archive gets big

By default Paperless stores files as 0000123.pdf in the media directory. That is fine until someone needs to browse the archive outside the web UI, or you need to hand an auditor a folder. PAPERLESS_FILENAME_FORMAT and storage paths fix that.

A format like this one gives you year and correspondent folders:

PAPERLESS_FILENAME_FORMAT={{ created_year }}/{{ correspondent }}/{{ title }}

Storage paths go further: each one is its own format string, assigned with the same matching algorithms as tags. One path can sort normal correspondence by year and sender while another keeps insurance documents in a flat folder with the full date in the name. Set PAPERLESS_FILENAME_FORMAT_REMOVE_NONE=true if you do not want none folders for documents that have no correspondent yet.

Choose the layout early. Changing it later is supported, but renaming tens of thousands of files is slower and harder to verify than getting it right on the first thousand.

7. Use custom fields for the values you search by

Tags are for grouping. When you need a value, such as an invoice number, amount, due date, or policy number, use a custom field instead. Supported types include Text, Boolean, Date, URL, Integer, Number, Monetary (an ISO 4217 currency plus an amount), Document Link, and Select.

Two practical notes. The data type cannot be changed after the field is created, so think about whether “Amount” should be Monetary or Number before you save it. And a date custom field can drive a Scheduled workflow, which is how you get a reminder before a warranty or contract expires without a separate calendar.

8. Tune OCR for your documents, not for the default

Paperless runs Tesseract through OCRmyPDF. Three settings matter most:

  • PAPERLESS_OCR_LANGUAGE takes three-letter codes such as eng or deu+eng. Only list languages you really receive, because every extra language costs CPU time on each page.
  • PAPERLESS_OCR_MODE defaults to auto in current releases, which skips OCR when a PDF already has a text layer. Use redo if your scanner embeds poor OCR you want replaced, and force only when nothing else works, since it rasterizes text and makes files larger. Older 2.x installs use different mode names, so check the docs for your version.
  • PAPERLESS_OCR_PAGES limits OCR to the first N pages. On a small NAS or a mini PC, 1 keeps long scans from tying up the box, at the cost of full-text search beyond page one.

Also decide how you want duplicates handled. Since v3.0, Paperless consumes duplicate files by default and shows a Duplicates tab on the detail page. If you prefer the old behavior of rejecting them, set PAPERLESS_CONSUMER_DELETE_DUPLICATES=true.

9. Back up with the exporter, then test a restore

Volume snapshots are good, but the application-level backup is document_exporter. It writes originals, archive files, thumbnails, and a manifest.json with all your metadata (tags, correspondents, custom fields, and so on) to a folder you choose. With the standard Docker Compose setup:

docker compose exec -T webserver document_exporter ../export -z

-z zips the result. -f names the exported files using your filename format, which makes the export readable on its own. Run it from cron or your NAS task scheduler, copy the export off the machine, and once a quarter import it into a throwaway instance with document_importer. An export you have never restored is a hope, not a backup.

10. Know where built-in OCR stops

Tesseract gives you a searchable text layer. It does not decide that a page is a supplier invoice from a particular vendor dated 14 March, and matching rules only get you part of the way when layouts and senders keep changing. That gap is where most of the remaining manual work in a Paperless setup lives.

This is the step OCRskill is built for. In the setup from our archive tutorial, a workflow sends each new document to OCRskill, which returns a document type, the issue date printed on the page, the correspondent, and a review flag when the document is unreadable or does not fit. Paperless then applies those values and your storage path files the document. If you want to pull specific values into custom fields as well, our guide to structured OCR to typed JSON shows what that extraction looks like.

Keep the inbox tag in place while you introduce it. Spot-check the first few hundred documents, fix the patterns that repeat, and only then let classification run without review.

Where to start

You do not need all ten at once. If your archive feels messy today, the highest return is usually the inbox tag, a storage path, and a scheduled export, in that order. Matching rules and workflows come next, once you can see which labels you keep fixing by hand. OCR tuning and automated classification are the last layer, because they only help when the structure underneath them is already sound.