Freight Forwarding OCR: Extract Data from Invoices, Bills of Lading, and AWBs
If you searched for freight forwarding OCR, you probably have a shipment file that arrives as a pile of PDFs, phone photos, and email attachments: a commercial invoice from the shipper, a packing list, a house or master bill of lading, an air waybill, and an arrival notice from the carrier or overseas agent. Someone on the ops desk retypes the parties, dates, references, and totals into the forwarding system. Short answer: OCR removes most of that retyping when you split the job in two. Use Markdown OCR to make every page searchable and to catch reference numbers, use typed JSON only for the fields your extraction API actually supports, validate container and air waybill numbers with their check digits in your own code, and send anything that fails to a human review queue before it touches a booking or a customs filing.
This post is about tool choice and pipeline shape for forwarders, NVOCCs, and customs brokerage desks. It is not an archive blueprint. If you need intake folders, document types, and a Paperless-ngx archive for BOLs, PODs, and freight invoices, start with the logistics paperless document workflow and come back here for the extraction layer.
Which documents a forwarder actually needs to read
A single ocean or air shipment can produce a surprising number of documents, and they come from different parties with different layouts:
- Commercial invoice from the seller or shipper: parties, invoice date, goods description, quantities, values, currency, and charges
- Packing list: package counts, marks and numbers, weights, and dimensions
- Bill of lading (house and master), or a sea waybill: shipper, consignee, notify party, ports, vessel and voyage, container and seal numbers
- Air waybill (house and master): the 11-digit AWB number, airports, pieces, weight, and handling information
- Arrival notice from the carrier or destination agent: ETA, free time, charges, and pickup references
- Certificates of origin, permits, and other support pages that a broker asks for during entry
Some of this is already digital. IATA’s e-AWB program lets airlines and forwarders replace the paper air waybill with an electronic shipment record under the Multilateral e-AWB Agreement (Resolution 672). That helps on lanes where it is active, but forwarders still receive scanned house documents, overseas agent PDFs, and shipper paperwork that never went through EDI. That leftover paper is where OCR earns its place.
What “freight forwarding OCR” should mean in practice
Character recognition is the easy part. The real question is what the ops desk gets back and whether it can be trusted. Score any tool against these points:
- Input realism. It should accept multipage PDFs, phone photos of stamped originals, and email attachments without a separate conversion step.
- Two output contracts. Full-page text for search and review, plus named fields for the values your system needs every time.
- Loud failure. A missing required value should return an error you can route, not an empty column you discover during a customs query.
- Original retained. Keep the scan next to the extracted data. Brokers, carriers, and auditors ask for the document, not your JSON.
- Honest schema. Know which fields are supported out of the box and which you have to parse yourself.
That last point matters more in freight than in most industries. A typical commercial invoice fits common invoice fields well. A bill of lading has carrier-specific reference numbers, container lists, and seal numbers that no general field catalog will name for you.
Markdown OCR for every page, typed JSON where the schema fits
OCRskill exposes both contracts on the same Bearer-token auth. POST /ocr returns Markdown for any page, and POST /ocr.json returns typed JSON for the fields you list. Both accept images (PNG, JPEG, WebP, GIF, BMP, TIFF) and documents such as PDF, Word, and Excel files, up to 20 MB per request. A free key from get-key.json is enough for testing.
Start with Markdown for the whole shipment file. It gives your team searchable text for every page and gives your code something to scan for reference numbers:
curl https://api.ocrskill.com/get-key.json
export API_KEY="sk-your-key-here"
curl https://api.ocrskill.com/ocr \
-H "Authorization: Bearer $API_KEY" \
-F "file=@hbl-scan.pdf" > hbl-scan.md
For commercial invoices, the invoice fields in the documented catalog map well to what a forwarder needs first: seller_name, buyer_name, invoice_date, and total_amount. Mark the shaky ones optional with a trailing ? so a single unreadable total does not reject the whole request:
curl "https://api.ocrskill.com/ocr.json?fields=seller_name,buyer_name,invoice_date,total_amount?" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@commercial-invoice.pdf"
An illustrative response shape (values are fictional):
{
"seller_name": "Example Textiles Co., Ltd.",
"buyer_name": "Example Imports GmbH",
"invoice_date": "2026-09-18",
"total_amount": 48250.0
}
Dates come back as YYYY-MM-DD. A required field that the API cannot find returns 400 with the extracted input_text, so your review queue can show the clerk what was read. Unknown field names are also rejected with 400, which keeps you honest: there is no bill_of_lading_number or container_number field in the catalog today, so do not pretend there is. The full list lives in the OCR JSON API reference, and the form data extraction API guide covers how to design field lists for recurring document families.
Validate container and AWB numbers in your own code
Reference numbers are where OCR mistakes hurt most. A misread 0 for O on a container number sends a trace request to the wrong box. The good news is that both container numbers and air waybill numbers carry check digits, so you can catch most single-character errors without a human.
- Container numbers (ISO 6346) have a three-letter owner code, an equipment category letter (
U,J, orZ), a six-digit serial, and a check digit. The BIC check digit calculator page describes the check digit as a way to catch data-entry errors and erroneous OCR readings. - Air waybill numbers are a three-digit airline prefix, a seven-digit serial, and a check digit equal to the serial modulo 7, as described in the Australian Border Force air waybill validation note.
This Python script scans OCR Markdown for both patterns and flags the ones that fail their check digit:
import re
import sys
# ISO 6346 letter values: A=10 upward, skipping multiples of 11
LETTER_VALUES = {}
value = 10
for ch in "ABCDEFGHIJKLMNOPQRSTUVWXYZ":
if value % 11 == 0:
value += 1
LETTER_VALUES[ch] = value
value += 1
CONTAINER_RE = re.compile(r"\b([A-Z]{3}[UJZ])\s?(\d{6})\s?(\d)\b")
AWB_RE = re.compile(r"\b(\d{3})[-\s]?(\d{7})(\d)\b")
def container_ok(owner_cat: str, serial: str, check: str) -> bool:
chars = owner_cat + serial
total = sum(
(LETTER_VALUES[c] if c.isalpha() else int(c)) * (2 ** i)
for i, c in enumerate(chars)
)
return (total % 11) % 10 == int(check)
def awb_ok(serial: str, check: str) -> bool:
return int(serial) % 7 == int(check)
def find_ids(markdown: str) -> dict:
text = markdown.upper()
containers = [
{"value": "".join(m.groups()), "valid": container_ok(*m.groups())}
for m in CONTAINER_RE.finditer(text)
]
awbs = [
{"value": f"{m.group(1)}-{m.group(2)}{m.group(3)}",
"valid": awb_ok(m.group(2), m.group(3))}
for m in AWB_RE.finditer(text)
]
return {"containers": containers, "awbs": awbs}
if __name__ == "__main__":
print(find_ids(open(sys.argv[1], encoding="utf-8").read()))
Run it against the Markdown from the previous step with python3 find_ids.py hbl-scan.md. On a test string, CSQU3054383 (a widely used ISO 6346 example) passes, CSQU3054384 fails, and 016-12345675 passes the AWB check.
Two caveats. A passing check digit proves the number is well formed, not that it belongs to this shipment, so still match it against the booking. And the AWB pattern is just “eleven digits”, which can also match phone or account numbers, so only trust it near an AWB label or inside an air shipment file.
A practical pipeline for a forwarding ops desk
Keep each step small and boring:
- Capture. Pull attachments from the ops mailbox, agent portal downloads, and scanner drops into one staging folder per shipment or job number. Keep originals unchanged and record a content hash.
- Classify. Tag each file as commercial invoice, packing list, HBL, MBL, HAWB, MAWB, arrival notice, or other. A wrong document type causes more downstream damage than a slow OCR call.
- OCR to Markdown. Run
/ocron everything. Store the Markdown beside the original so search works across the whole job file. - Typed fields where they fit. Run
/ocr.jsonon commercial invoices for seller, buyer, date, and total. Leave B/L and AWB layouts on Markdown plus your own parsing. - Validate. Check container and AWB numbers, compare invoice parties against the booking, and flag invoice dates that do not fit the shipment timeline. Anything that fails goes to review.
- Hand off. Push only validated values into the forwarding or customs system, with a link back to the original scan.
For high volumes, two paid-key features help. A document URL lets OCRskill fetch a file that already sits in your object storage, and an OCR callback returns 202 right away and delivers the result to your endpoint later, so a large batch of agent PDFs does not hold open connections. Free keys get 403 for both.
Limitations to plan for before the pilot
Freight paperwork is a hard document class, and a pilot goes better when everyone knows that up front:
- Line items. Commercial invoices can run to hundreds of lines with descriptions, quantities, and unit values. The typed catalog covers header-level invoice fields, not per-line tariff data. Keep line items in Markdown and parse tables in your code, or have a person handle classification.
- Multipage packets. Agents often send one PDF with the invoice, packing list, and B/L stapled together. Split by document before typed extraction so fields from page 4 do not land on the wrong record.
- Stamps, signatures, and carbon copies. Faded originals and stamp-heavy pages raise error rates. Keep the image and expect more review on them.
- Language. Invoices from overseas suppliers may not be in English. For US imports, 19 CFR 141.86 requires the invoice to be in English or to carry an accurate English translation. OCR gives you the text that is on the page. It does not supply that translation.
- Compliance stays with people. That same regulation lists what a US commercial invoice must contain, from the parties and goods description to itemized charges and country of origin. Extraction helps a broker find those values quickly. Whether the invoice is complete, what the tariff classification should be, and what gets filed are decisions for the licensed broker and your customs system.
Do not quote accuracy numbers you have not measured. Pick 50 real shipment files from the last quarter, run them through the pipeline, and track how often a clerk still corrects a value before it is trusted.
Where OCRskill fits and where it does not
OCRskill handles the extraction step: files in, Markdown or typed JSON out, with clear errors when a required field is missing. It does not replace your forwarding system, your customs filing software, carrier EDI, or e-AWB messaging. It also does not know your booking data, so matching extracted values to jobs belongs in your code. If you want the archive around these files, the logistics paperless document workflow covers document types, naming, and Paperless-ngx retrieval for claims. For more detail on how typed extraction behaves, see the structured OCR JSON API post.
Conclusion
Good freight forwarding OCR is mostly about knowing where a generic schema stops. Commercial invoice headers fit typed fields. Container and AWB numbers are best pulled from Markdown and checked against their standard check digits, which catch most single-character OCR errors for free. Everything else, from line items to compliance decisions, belongs in a review step with the original scan one click away. Start with one lane and one document family, measure how often clerks still correct values, and widen only when that number is low enough to staff.
To try the extraction half, get a key from get-key.json, run a recent commercial invoice through /ocr.json, and run the matching bill of lading through /ocr and the validator above. Product details and pricing are at ocrskill.com.
