Packing List OCR: Read Line Items and Reconcile Them Against the PO
If you searched for packing list OCR, you probably have a receiving desk where supplier packing lists arrive taped to pallets, stuffed in carton pockets, or attached to an email, and someone keys item numbers and quantities into the ERP or WMS before the goods can be booked in. Short answer: packing list OCR works when you treat the packing list as a table, not as a form. Run OCR to get the whole page as Markdown, parse the line-item table in your own code, compare each line with what is still open on the purchase order, check any SSCC numbers with their GS1 check digit, and route every mismatch to a person before the receipt is posted. Header fields such as seller and buyer can come back as typed JSON. The line items are where the value is, and they need a parser you control.
This post is about the extraction and reconciliation step at receiving and incoming inspection. It is not an archive design. If you need document types, intake folders, and a Paperless-ngx archive for ASNs, packing lists, and receiving tickets, the distribution paperless document workflow covers that. If your packing lists arrive inside a forwarder’s shipment file next to bills of lading and air waybills, read freight forwarding OCR for the shipment side.
What a packing list actually tells you
A packing list is the shipper’s statement of what went into each package. The US International Trade Administration describes it as a document that itemizes the contents of each package with weights, measurements, and detailed lists of the goods, and notes that freight forwarders use it to work out weights and freight costs while customs officials may use it to check a specific carton.
For export shipments, the ITA’s common export documents page lists what a fuller export packing list carries: seller, buyer, shipper, invoice number, date of shipment, mode of transport, carrier, quantities, descriptions, package type (box, crate, drum, carton), number of packages, net and gross weight, and package marks and dimensions. The same page says a packing list is not a substitute for a commercial invoice, and that the invoice should reflect what the packing list shows.
Domestic packing lists and packing slips are usually thinner, but the receiving desk cares about the same core values:
- References: packing list number, the buyer’s PO number, sometimes a delivery note or invoice number
- Parties: the supplier (seller or shipper) and the ship-to or buyer
- Line items: item or part number, description, quantity shipped, unit of measure
- Packaging: number of cartons or pallets, weights, and often a package or pallet ID
The important point for OCR: almost everything a receiving clerk checks lives in the line-item table, not in the header.
Packing list OCR versus the ASN you may already have
Many suppliers also send an advance ship notice. In North American EDI that is usually the X12 856 Ship Notice/Manifest, which X12 describes as listing the contents of a shipment along with order information, product descriptions, packaging, markings, and carrier details. When a supplier sends a clean 856, use it. Structured EDI beats OCR on every count.
Packing list OCR earns its place for the rest: smaller suppliers without EDI, overseas vendors, drop shipments, returns, and the shipments where the paper in the carton does not match the ASN that arrived yesterday. In those cases the packing list is the only document that travelled with the goods, so it is the one you should read.
Two outputs: Markdown for the table, typed JSON for the header
OCRskill gives you two contracts on the same Bearer-token auth. POST /ocr returns Markdown for any page. POST /ocr.json returns typed JSON for the fields you list. Both accept images (PNG, JPEG, WebP, GIF, BMP, TIFF) and documents including PDF, Word, Excel, and CSV, up to 20 MB per request. A free key from get-key.json is enough to test.
Start with Markdown for every packing list:
curl https://api.ocrskill.com/get-key.json
export API_KEY="sk-your-key-here"
curl https://api.ocrskill.com/ocr \
-H "Authorization: Bearer $API_KEY" \
-F "file=@packing-list.pdf" > packing-list.md
The header text comes back as plain lines, and the line-item grid comes back as a table. On a clean PDF the table can arrive as an HTML <table> block inside the Markdown rather than as a pipe table, so your parser should accept both shapes. An abbreviated example of what that looks like (values are fictional):
# PACKING LIST
Seller: Example Fasteners Ltd., 12 Mill Road, Leeds, UK
Packing List No: PL-77120 Date: 2026-09-30
Customer PO: 4500018823 Invoice No: INV-55102
<table>
<tr><th>Line</th><th>Item No.</th><th>Description</th><th>Qty</th><th>UoM</th></tr>
<tr><td>1</td><td>HX-M8-40</td><td>Hex bolt M8x40 zinc</td><td>2000</td><td>PCS</td></tr>
<tr><td>3</td><td>WSH-M8</td><td>Flat washer M8</td><td>1500</td><td>PCS</td></tr>
</table>
For the parties, the invoice fields in the documented catalog fit well. seller_name and buyer_name usually map to the supplier and the receiving company:
curl "https://api.ocrskill.com/ocr.json?fields=seller_name,buyer_name?" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@packing-list.pdf"
{
"seller_name": "Example Fasteners Ltd.",
"buyer_name": "Example Assembly GmbH"
}
The trailing ? marks buyer_name as optional, so a packing list without a clear buyer block does not fail the request. Treat an optional field that is missing or null the same way: not found. A required field that cannot be found returns 400 with the extracted input_text, which your review screen can show.
Be honest about what the catalog does not have. There is no po_number, packing_list_number, quantity, or line-item field today, and the endpoint rejects unknown field names with 400. That is a good thing. It means you will not get a confident-looking JSON object for values nobody defined. Pull those values from the Markdown with code you can test, as the next section shows. The OCR JSON API reference has the full field list.
Reconcile the packing list against the PO in code
A receiving check compares three things: what you ordered (the PO), what the supplier says it shipped (the packing list), and what the dock actually counted. OCR only gives you the second one. The useful automation is to compare it with the first one before anyone starts counting, so the clerk knows which lines to look at.
The script below reads the Markdown from /ocr, finds the first table that has an item column and a quantity column, finds the PO reference in the header text, and compares each line with a CSV of open PO lines. It also checks any 18-digit number in the page as a possible SSCC. GS1 uses the Serial Shipping Container Code to identify a logistic unit such as a case, pallet, or parcel, and its final digit is a mod-10 check digit calculated with alternating weights of 3 and 1, as shown on GS1’s manual check digit page. A failed check digit on an OCR-read SSCC almost always means a misread character.
import csv
import re
import sys
from html.parser import HTMLParser
QTY_HEADERS = {"qty", "quantity", "qty shipped", "shipped", "pcs"}
ITEM_HEADERS = {"item", "item no.", "item no", "part", "part no.",
"part number", "sku", "article"}
class TableParser(HTMLParser):
"""Collect rows from HTML tables inside OCR Markdown."""
def __init__(self):
super().__init__()
self.tables, self.row, self.cell, self.in_cell = [], None, "", False
def handle_starttag(self, tag, attrs):
if tag == "table":
self.tables.append([])
elif tag == "tr":
self.row = []
elif tag in ("td", "th"):
self.in_cell, self.cell = True, ""
def handle_endtag(self, tag):
if tag in ("td", "th") and self.row is not None:
self.row.append(self.cell.strip())
self.in_cell = False
elif tag == "tr" and self.row is not None and self.tables:
self.tables[-1].append(self.row)
self.row = None
def handle_data(self, data):
if self.in_cell:
self.cell += data
def pipe_tables(markdown):
"""Collect rows from Markdown pipe tables, skipping the |---| divider."""
tables, current = [], []
for line in markdown.splitlines():
if line.strip().startswith("|"):
cells = [c.strip() for c in line.strip().strip("|").split("|")]
if not all(re.fullmatch(r":?-{3,}:?", c) for c in cells):
current.append(cells)
elif current:
tables.append(current)
current = []
if current:
tables.append(current)
return tables
def line_items(markdown):
parser = TableParser()
parser.feed(markdown)
for table in parser.tables + pipe_tables(markdown):
header = [h.lower() for h in table[0]]
item_col = next((i for i, h in enumerate(header) if h in ITEM_HEADERS), None)
qty_col = next((i for i, h in enumerate(header) if h in QTY_HEADERS), None)
if item_col is None or qty_col is None:
continue
items = {}
for row in table[1:]:
try:
qty = float(row[qty_col].replace(",", ""))
except (IndexError, ValueError):
qty = None
items[row[item_col]] = qty
return items
return {}
def sscc_ok(code):
"""GS1 mod-10: weights 3,1,3,... from the left over the first 17 digits."""
body, check = code[:17], int(code[17])
total = sum(int(d) * (3 if i % 2 == 0 else 1) for i, d in enumerate(body))
return (10 - total % 10) % 10 == check
def main(markdown_path, po_csv_path):
markdown = open(markdown_path, encoding="utf-8").read()
with open(po_csv_path, newline="", encoding="utf-8") as f:
po = {row["item"]: float(row["open_qty"]) for row in csv.DictReader(f)}
issues = []
po_ref = re.search(
r"\bP\.?O\.?(?:\s*(?:No\.?|Number|#))?\s*[:#]?\s*(\d{6,12})", markdown, re.I
)
if not po_ref:
issues.append("no PO reference found on the packing list")
shipped = line_items(markdown)
if not shipped:
issues.append("no line-item table with item and quantity columns found")
for item, qty in shipped.items():
if qty is None:
issues.append(f"{item}: quantity unreadable")
elif item not in po:
issues.append(f"{item}: not on the PO")
elif qty > po[item]:
issues.append(f"{item}: shipped {qty:g}, open on PO {po[item]:g}")
for item in po.keys() - shipped.keys():
issues.append(f"{item}: on the PO but not on the packing list")
for code in re.findall(r"\b\d{18}\b", markdown):
if not sscc_ok(code):
issues.append(f"SSCC {code}: check digit fails")
print(f"PO reference: {po_ref.group(1) if po_ref else 'missing'}")
print(f"Line items read: {len(shipped)}")
print("Route to review:" if issues else "No discrepancies found.")
for issue in issues:
print(f" - {issue}")
if __name__ == "__main__":
main(sys.argv[1], sys.argv[2])
Export the open lines for the PO from your ERP as a two-column CSV:
item,open_qty
HX-M8-40,2000
NUT-M8,2000
WSH-M8,1000
PIN-4,500
Then run python3 check_packing_list.py packing-list.md po-4500018823.csv. Against a three-line test packing list with a mistyped SSCC, the output looks like this:
PO reference: 4500018823
Line items read: 3
Route to review:
- WSH-M8: shipped 1500, open on PO 1000
- PIN-4: on the PO but not on the packing list
- SSCC 354123450000000013: check digit fails
That is exactly the list a receiving clerk wants before opening cartons: one over-shipment to confirm or refuse, one back-ordered line to expect later, and one pallet ID to read again from the label. Lines that match need a count, not a debate.
A few things to adapt before you trust it. The header synonyms are a starting set, so add the column names your suppliers actually use (Menge, Qté, Ship Qty). Item numbers on the packing list are often the supplier’s part numbers, not yours, so you will need a cross-reference table before the comparison means anything. Units matter too: 20 boxes of 100 and 2,000 pieces are the same shipment written two ways, so convert units of measure before comparing. And the 18-digit pattern can match other long numbers, so only trust it near an SSCC label or a (00) application identifier when your pages have other long references.
A receiving and incoming inspection pipeline that stays small
Each step should be easy to explain to the person who gets the review task:
- Capture. Scan the packing list at the dock, or save the supplier’s emailed PDF, into a staging folder named by PO or delivery. Keep the original unchanged and record a content hash.
- OCR to Markdown. Run
/ocron every page and store the Markdown next to the original so the receipt is searchable later. - Header fields. Run
/ocr.jsonforseller_nameand an optionalbuyer_name, and check the seller against the supplier on the PO. - Reconcile. Run the line-item comparison against open PO lines. Over-shipments, unknown items, unreadable quantities, and failed SSCC checks go to review.
- Count and inspect. The dock counts the goods. For production material, incoming inspection then checks the parts and their certificates against the spec. The packing list check only tells inspectors which lines already look wrong.
- Post. Book the receipt in the ERP or WMS with the counted quantity, not the packing list quantity, and link the original scan to the receipt.
For manufacturing sites, step 5 is where the material certificates, CoA and CoC packets, and inspection records start. The manufacturing paperless document workflow covers how to type, name, and archive those so a lot trace finds them later.
For higher volumes, two paid-key features help. A document URL lets OCRskill fetch a file already sitting in your object storage, and an OCR callback returns 202 right away and delivers the result to your endpoint later. Free keys get 403 for both.
Limitations to plan for
Packing lists look simple and still cause trouble in a pilot:
- Phone photos at the dock. Skewed, shadowed, or crumpled pages raise error rates and can break table structure. Ask for a flat scan when the paper allows it, and expect more review on photos.
- Multipage tables. Long packing lists continue the table on the next page, sometimes without repeating the header row. Run OCR per document, then join the tables before reconciling, or the script will only see page one.
- Merged and nested layouts. Some suppliers print one row per carton with items nested underneath, or merge cells for pallet totals. Simple grids survive OCR well. Nested ones need a supplier-specific parser or a person.
- Handwritten corrections. A crossed-out quantity with a handwritten number beside it is common on short shipments. OCR may return the printed value, the handwritten one, or both. Flag pages with visible corrections for review.
- Packing list versus reality. OCR reads what the supplier wrote. It cannot tell you whether the carton holds what the paper says, which is why the counted quantity is the one you post.
Do not publish accuracy figures you have not measured. Take 50 real packing lists from last month across your top suppliers, run them through the pipeline, and track how many lines a clerk still corrects. That number tells you which suppliers are ready for automation and which still need a person on every page.
Where OCRskill fits
OCRskill handles the reading: files in, Markdown or typed JSON out, with a clear 400 when a required field is missing or a field name is unknown. It does not hold your PO data, count cartons, replace EDI 856 messages from suppliers who can send them, or decide whether an over-shipment is accepted. The reconciliation logic, the supplier part cross-reference, and the receipt posting stay in your code and your ERP or WMS.
Conclusion
Packing list OCR pays off at the line-item table, and that table is not something a generic field catalog should guess at. Get the page as Markdown, parse the table with code you can test against your own suppliers, and compare it with the open PO before anyone counts. Use typed JSON for the seller and buyer, the GS1 check digit for SSCCs, and a review queue for everything that disagrees. Post the counted quantity, keep the original scan linked to the receipt, and measure correction rates per supplier before you widen the rollout.
To try it, get a key from get-key.json, run last week’s messiest packing list through /ocr, and feed the Markdown and a PO export to the checker above. Product details and pricing are at ocrskill.com.
