Commercial Invoice OCR: Extract Parties, Totals, and Line Data for Customs
If you searched for commercial invoice OCR, you probably have a stack of seller invoices that still arrive as PDFs, phone photos, or faxed scans, and someone retypes parties, invoice numbers, totals, and line values into a customs system or an AP screen. Short answer: commercial invoice OCR works when you split the page the same way a broker reads it. Use typed JSON for the header fields your extraction API actually supports (seller, buyer, invoice date, total). Use Markdown for everything else (invoice number, currency, Incoterms, country of origin, itemized charges, and the line-item table). Then validate the extracted values against the packing list, the purchase order, and the checklist your entry team already uses, and send every failure to a human before the file touches a customs filing or a payment run.
This post is about the commercial invoice as its own document class: valuation evidence for customs and the bill that AP pays. It is not a full freight shipment walkthrough. For bills of lading, air waybills, and the multi-document forwarder packet, read freight forwarding OCR. For packing-list line reconciliation at the dock, read packing list OCR. For intake folders and a Paperless-ngx archive around BOLs, PODs, and freight invoices, use the logistics paperless document workflow.
What a commercial invoice has to show
A commercial invoice is the seller’s bill to the buyer for the goods in a shipment. The US International Trade Administration describes it as a legal document between the exporter and the foreign buyer that states what is being sold and what the customer must pay, and notes that customs authorities often use it when assessing duties. It is not a packing list, and a packing list is not a substitute for it.
For merchandise entering the United States, 19 CFR 141.86 sets out what each invoice of imported merchandise must set forth. The regulation is worth reading in full when you design a review checklist. In practice, brokers and importers look for values in these groups:
- Shipment context: port of entry destined for the goods, and when and where the sale (or shipment) happened
- Parties: who sold and who bought (or who shipped and who received, when there is no purchase)
- Goods: a detailed description per item, including how the trade knows the goods, plus package marks and numbers
- Quantities and prices: quantities in usable weights or measures, unit purchase prices or values, and the kind of currency
- Charges: freight, insurance, commission, packing, and other charges itemized by name and amount when they are not already in the price
- Origin: country of origin of the merchandise
- Discounts and assists: every discount that affects the purchase price, and any assists that are not already in the price
The invoice and its attachments must be in English or carry an accurate English translation. OCR returns the characters that are on the page. It does not invent a translation when the seller sent only a foreign-language original.
That list is the reason commercial invoice OCR is different from generic “invoice OCR” for domestic AP. A restaurant receipt or a software SaaS invoice needs a merchant, a date, and a total. A commercial invoice for customs also needs origin, detailed goods descriptions, itemized charges, and often Incoterms and a currency that match the contract. Design the pipeline for the stricter checklist, then reuse the header extraction for AP when the same scan feeds both desks.
Commercial invoice versus packing list and pro forma
Three documents get mixed up on the same email thread:
| Document | What it mainly proves | OCR priority |
|---|---|---|
| Commercial invoice | What was sold, for how much, under which terms | Parties, date, total, currency, origin, charges, line values |
| Packing list | What went into each package | Line quantities, weights, package marks; reconcile to the PO |
| Pro forma invoice | What the seller quotes before shipment | Same shape as a commercial invoice, but not the final customs bill |
The ITA notes that the commercial invoice should reflect what the packing list shows. Your validation code should treat that as a soft check: line descriptions and quantities that disagree are review items, not silent merges. A pro forma that never became a final commercial invoice should not feed an entry or a payment.
Two outputs: typed JSON for the header, Markdown for the rest
OCRskill exposes both contracts on the same Bearer-token auth. POST /ocr returns Markdown for any page. POST /ocr.json returns typed JSON for the fields you list. Both accept images (PNG, JPEG, WebP, GIF, BMP, TIFF) and documents including PDF, Word, Excel, and CSV, up to 20 MB per request. A free key from get-key.json is enough to test.
For commercial invoices, the invoice fields in the documented catalog map cleanly to the header: seller_name, buyer_name, invoice_date, and total_amount. Mark the total optional with a trailing ? when a faded stamp often hides it, so one unreadable figure does not reject the whole request:
curl https://api.ocrskill.com/get-key.json
export API_KEY="sk-your-key-here"
curl "https://api.ocrskill.com/ocr.json?fields=seller_name,buyer_name,invoice_date,total_amount?" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@commercial-invoice.pdf"
An illustrative response shape (values are fictional):
{
"seller_name": "Example Textiles Co., Ltd.",
"buyer_name": "Example Imports GmbH",
"invoice_date": "2026-09-18",
"total_amount": 48250.0
}
Dates come back as YYYY-MM-DD. A required field the API cannot find returns 400 with the extracted input_text, so the review queue can show the clerk what was read. Unknown field names also return 400. There is no invoice_number, currency, incoterm, hs_code, or country_of_origin field in the catalog today, so do not invent them in the fields list. Pull those from Markdown instead. The full catalog lives in the OCR JSON API reference, and the structured OCR JSON API post covers required versus optional behavior and discovery mode.
Always run Markdown on the same file:
curl https://api.ocrskill.com/ocr \
-H "Authorization: Bearer $API_KEY" \
-F "file=@commercial-invoice.pdf" > commercial-invoice.md
On a clean PDF you typically get plain header lines plus a line-item table (sometimes as an HTML <table> block inside the Markdown). That Markdown is where invoice numbers, Incoterms, currency codes, country of origin labels, HS or Schedule B codes, and itemized freight or insurance lines usually live. Keep it next to the original scan so a broker can search the page later without opening the image.
Validate what OCR returns before anyone files or pays
Extraction without checks only moves the typing error into your database. Wire a small validation layer around the JSON and Markdown:
- Party match. Compare
seller_nameandbuyer_nameto the vendor and ship-to on the purchase order or booking. Fuzzy match is fine for legal suffixes; a completely different seller is a hard stop. - Date sanity. Reject invoice dates far outside the shipment window you already know from the booking or ASN.
- Total versus lines. When the Markdown table has unit prices and quantities, sum them in code and compare to
total_amount. A mismatch means a missing line, a wrong total, or a charge that lives outside the table. - Customs checklist. Scan the Markdown for labels your entry team requires (country of origin, currency, Incoterms such as FOB or CIF, and any HS codes the seller printed). Absence is a review flag, not a made-up value.
- Cross-document checks. Compare quantities and descriptions to the packing list and to open PO lines. Over-shipments and unknown SKUs belong in the same review queue the dock already uses.
Here is a compact Python sketch that checks a typed total against a Markdown HTML table and looks for a few common labels. Adapt the synonyms to the languages your suppliers actually print:
import re
from html.parser import HTMLParser
class TableParser(HTMLParser):
def __init__(self):
super().__init__()
self.rows, self._row, self._cell, self._in_td = [], [], [], False
def handle_starttag(self, tag, attrs):
if tag in ("td", "th"):
self._in_td, self._cell = True, []
elif tag == "tr":
self._row = []
def handle_endtag(self, tag):
if tag in ("td", "th") and self._in_td:
self._row.append("".join(self._cell).strip())
self._in_td = False
elif tag == "tr" and self._row:
self.rows.append(self._row)
def handle_data(self, data):
if self._in_td:
self._cell.append(data)
def parse_amount(text: str) -> float | None:
m = re.search(r"[-+]?\d{1,3}(?:[,\s]\d{3})*(?:\.\d+)?|[-+]?\d+(?:\.\d+)?", text.replace(" ", ""))
if not m:
return None
return float(m.group(0).replace(",", ""))
def line_sum_from_markdown(md: str, qty_idx: int, price_idx: int) -> float | None:
p = TableParser()
p.feed(md)
if len(p.rows) < 2:
return None
total = 0.0
for row in p.rows[1:]:
if max(qty_idx, price_idx) >= len(row):
continue
qty, price = parse_amount(row[qty_idx]), parse_amount(row[price_idx])
if qty is None or price is None:
return None
total += qty * price
return total
LABELS = {
"origin": re.compile(r"country\s+of\s+origin|origin\s*:", re.I),
"currency": re.compile(r"\b(USD|EUR|GBP|JPY|CNY)\b|currency\s*:", re.I),
"incoterm": re.compile(r"\b(FOB|CIF|CFR|EXW|DDP|DAP|FCA)\b|incoterm", re.I),
}
def review_flags(md: str, total_amount: float | None, qty_idx=3, price_idx=4) -> list[str]:
flags = []
for name, pattern in LABELS.items():
if not pattern.search(md):
flags.append(f"missing label for {name}")
line_sum = line_sum_from_markdown(md, qty_idx, price_idx)
if total_amount is not None and line_sum is not None:
if abs(line_sum - total_amount) > 0.05 * max(abs(total_amount), 1.0):
flags.append(f"line sum {line_sum:.2f} disagrees with total {total_amount:.2f}")
elif total_amount is None:
flags.append("total_amount missing from typed JSON")
return flags
Column indexes differ by supplier, so discover them from the header row once per vendor template rather than hard-coding forever. The point of the sketch is the workflow: typed total in, Markdown table parsed in your code, and a short list of review reasons out.
A practical pipeline for broker, importer, and AP desks
Keep the steps small enough that the person who owns the review queue can explain them:
- Capture. Save the seller PDF or scan into a staging folder keyed by shipment or PO. Keep the original bytes and a content hash.
- Classify. Confirm the page is a commercial invoice (not a packing list, arrival notice, or credit note) before typed extraction.
- Header JSON. Call
/ocr.jsonforseller_name,buyer_name,invoice_date, and optionaltotal_amount. - Full-page Markdown. Call
/ocrand store the Markdown beside the original. - Validate. Run party, date, total-versus-lines, and customs-label checks. Failures go to review with the image one click away.
- Hand off. Push only validated header values into the customs or AP system. Attach the scan and Markdown. Classification, valuation method, and whether the invoice is complete remain human decisions.
For higher volumes, two paid-key features help. A document URL lets OCRskill fetch a file already in your object storage, and an OCR callback returns 202 and posts the result to your endpoint later. Free keys get 403 for both.
Limitations to plan for before the pilot
Commercial invoices look structured and still break naive automation:
- Line items and tariff data. Typed fields cover header-level invoice data, not per-line HS codes or unit values. Parse tables from Markdown, or keep a person on classification.
- Multipage packets. Sellers often staple the invoice, packing list, and certificate of origin into one PDF. Split by document type before field extraction so page four does not overwrite page one.
- Currency and rounding. OCR may return a total without a currency code. Currency usually lives in Markdown or a symbol next to the amount. Do not assume USD.
- Language and translation. Non-English originals need an English attachment for US entry. OCR will not produce that attachment for you.
- Compliance stays with licensed people. Extraction helps a broker find parties, totals, and printed origin statements quickly. Whether the invoice meets 19 CFR 141.86, how goods are classified, and what gets filed are decisions for the broker and your customs system, not for the OCR call.
Measure before you claim accuracy. Take fifty real commercial invoices from the last quarter across your top suppliers, run them through the pipeline, and track how often a clerk still corrects a value. That number tells you which vendors are ready for light-touch review and which still need a person on every page.
Where OCRskill fits
OCRskill handles the reading step: files in, Markdown or typed JSON out, with a clear 400 when a required field is missing or a field name is unknown. It does not replace your customs brokerage system, your ERP, EDI invoice messages from suppliers who already send them, or the licensed judgment behind an entry. Matching extracted values to purchase orders, computing duties, and deciding whether an invoice is complete stay in your code and your people.
Conclusion
Commercial invoice OCR pays off when you treat the page as customs evidence first and an AP bill second. Pull the few header fields the catalog supports as typed JSON, keep the rest as searchable Markdown, and validate parties, totals, and origin labels before anything reaches an entry or a payment. Cross-check the packing list and the PO so quantity fights surface early. Start with one trade lane and one supplier family, measure correction rates, and widen only when the review queue is quiet enough to staff.
To try it, get a key from get-key.json, run a recent commercial invoice through /ocr.json and /ocr, and feed the Markdown plus the typed total into the checker above. Product details and pricing are at ocrskill.com.
