Paperless Expense Management: From Receipt Photo to Approved, Archived Claim
← All posts
GuideOct 7, 2026· 10 min read

Paperless Expense Management: From Receipt Photo to Approved, Archived Claim

If you searched for paperless expense management, you are probably tired of the month-end envelope: employees hand in crumpled receipts, someone retypes merchant names and totals into a spreadsheet, a manager signs a printout, and finance staples the whole thing into a binder nobody opens again until an audit. The short answer is that paperless expense management replaces each of those paper steps with a data step. The employee captures the receipt once, close to the moment of purchase. OCR turns the image into a few typed fields (merchant, date, total) plus searchable text for everything else. Policy checks and approvals run on those fields. The original image stays attached to the claim and lands in an archive with a retention date. Nobody should ever have to type a total that is already printed on the receipt.

This guide covers employee-paid spend: out-of-pocket receipts, travel, meals, and the reimbursement claim that wraps them. It is not about supplier invoices. If your problem is vendor invoices, remittances, and statements reaching AP and the ERP, read the accounting paperless document workflow instead. That post designs the finance archive. This one designs the path a single receipt takes from an employee’s phone to a paid, defensible reimbursement.

What a receipt has to prove

Before you design screens or pick tools, decide what evidence each claim line must carry. Tax rules give you a useful floor.

In the US, IRS Publication 463 says documentary evidence is ordinarily adequate if it shows the amount, date, place, and essential character of the expense. A restaurant receipt should show the restaurant name and location, the number of people served, and the date and amount. A hotel receipt should show the hotel name and location, the dates of the stay, and separate amounts for lodging, meals, and phone charges. Receipts are generally not required for an expense other than lodging that is under $75, but the employee still has to record the business purpose. For employers running an accountable plan, the same publication treats accounting for expenses within 60 days and returning excess reimbursement within 120 days as a reasonable period.

In the EU, the VAT Directive lets suppliers issue a simplified invoice when the amount is not higher than EUR 100 (Article 220a), and Member States may allow it for some higher amounts. A simplified invoice must carry at least the date of issue, the identity of the supplier, the type of goods or services, and the VAT amount or the information needed to calculate it (Article 226b). Article 233 also requires the legibility of an invoice to be ensured until the end of its storage period, which matters a lot for thermal paper that fades in a drawer.

Turn that into a checklist your workflow can test automatically:

  • Merchant name, and location where the rules ask for it
  • Transaction date
  • Total amount and currency
  • Tax or VAT amount, when you plan to reclaim VAT
  • Business purpose, written by the employee (OCR cannot read intent off a receipt)
  • Attendees, for meals
  • The original image or PDF, legible

Your own policy and your tax advisor set the final list. The point is that most of these values are printed on the receipt, so software should extract them and humans should only add what the paper cannot say.

The paperless expense management workflow in five stages

A workable flow looks like this:

  1. Capture. The employee photographs the receipt or forwards the emailed PDF right after the purchase.
  2. Extract. OCR returns typed fields for the values your checks need, plus full text for the rest.
  3. Check. Code compares the claim line against policy: missing fields, late submission, duplicates, amounts over a limit, and a card transaction to match.
  4. Approve. A manager sees the image next to the extracted values and decides on exceptions, not on arithmetic.
  5. Reimburse and archive. Payroll or AP pays the claim, and the original, the extracted data, and the approval trail are stored together with a retention date.

Each stage has one owner and one output. When a claim stalls, you can see which stage it is stuck in.

Capture: fewer, better images

Most extraction errors start at capture, so make capture easy to do well:

  • Capture early. A receipt photographed at the table survives. A receipt found in a coat pocket three weeks later may be faded, torn, or gone. Pub 463 notes that a record made at or near the time of the expense has more value than one prepared later.
  • One receipt per image. Photos of four receipts laid out on a desk produce one muddled record. Ask for one image per receipt, flat, with all four edges visible.
  • Prefer the itemized receipt to the card slip. A card terminal slip usually shows the merchant and total but not what was bought or how many people ate. For meals, the itemized bill is the evidence.
  • Accept emailed PDFs directly. Ride-hailing, rail, airline, and hotel receipts mostly arrive by email. Let employees forward them to a capture address instead of printing and rephotographing them.
  • Ask for the business purpose at capture time. One short text field while the context is fresh beats a reconstruction at month-end.

Extract: typed fields where the schema fits, text for the rest

OCRskill exposes two endpoints that fit this stage. POST /ocr.json returns typed JSON for the fields you name, and POST /ocr returns Markdown for the whole page. Both take images (PNG, JPEG, WebP, GIF, BMP, TIFF) and documents such as PDF, up to 20 MB per request, with a Bearer token. A free test key comes from curl https://api.ocrskill.com/get-key.json.

For receipts, the documented field catalog includes seller_name, receipt_date, and total_amount. Dates come back as YYYY-MM-DD and the total comes back as a number:

export API_KEY="sk-your-key-here"

curl "https://api.ocrskill.com/ocr.json?fields=receipt_date,total_amount,seller_name?" \
  -H "Authorization: Bearer $API_KEY" \
  -F "file=@dinner-receipt.jpg"

An illustrative response (values are fictional):

{
  "receipt_date": "2026-09-29",
  "total_amount": 186.4,
  "seller_name": "Example Bistro SRL"
}

Fields without a suffix are required. If the API cannot find one, it returns 400 and includes the extracted input_text, which is exactly what a reviewer needs to see. The trailing ? makes seller_name optional, so a receipt with an unreadable logo still returns its date and total.

Be clear about what the catalog does not include. There is no currency field, no VAT or tax field, no line-item field, and no expense category. Unknown field names are rejected with 400, so you cannot quietly ask for them. Take those values from the Markdown text in your own code, or have the employee confirm them. The structured OCR JSON API post explains the required and optional field behavior in more detail.

A working example: one receipt in, one claim line out

This Python script calls both endpoints for a single receipt, pulls a currency hint and any tax or VAT lines out of the Markdown, and flags anything a person should look at. It uses the requests library.

import hashlib
import json
import re
import sys
from datetime import date

import requests

API = "https://api.ocrskill.com"
API_KEY = "sk-your-key-here"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
FIELDS = "receipt_date,total_amount,seller_name?"
SUBMIT_WINDOW_DAYS = 60  # set this from your own expense policy
LARGE_RECEIPT = 75.0     # example threshold for an extra tax-line check

CURRENCY_RE = re.compile(r"\b(EUR|USD|GBP|RON|CHF|PLN)\b|([€$£])")
SYMBOLS = {"€": "EUR", "$": "USD", "£": "GBP"}
TAX_WORDS = re.compile(r"\b(VAT|TVA|MWST|IVA|BTW|TAX)\b", re.IGNORECASE)
AMOUNT_RE = re.compile(r"\d+[.,]\d{2}")


def ocr_fields(path):
    with open(path, "rb") as f:
        r = requests.post(f"{API}/ocr.json?fields={FIELDS}", headers=HEADERS,
                          files={"file": f}, timeout=120)
    if r.status_code == 400:
        try:
            return None, r.json().get("input_text", r.text)
        except ValueError:
            return None, r.text
    r.raise_for_status()
    return r.json(), None


def ocr_markdown(path):
    with open(path, "rb") as f:
        r = requests.post(f"{API}/ocr", headers=HEADERS,
                          files={"file": f}, timeout=120)
    r.raise_for_status()
    return r.text


def currency_hint(text):
    m = CURRENCY_RE.search(text)
    if not m:
        return None
    return m.group(1) or SYMBOLS[m.group(2)]


def tax_lines(text):
    found = []
    for line in text.splitlines():
        if TAX_WORDS.search(line) and AMOUNT_RE.search(line):
            found.append(line.strip())
    return found


def build_claim_line(path, employee, purpose, seen_keys):
    with open(path, "rb") as f:
        digest = hashlib.sha256(f.read()).hexdigest()
    line = {"employee": employee, "purpose": purpose, "file": path,
            "sha256": digest, "flags": []}

    fields, failed_text = ocr_fields(path)
    if fields is None:
        line["flags"].append("date or total not found, needs manual entry")
        line["ocr_text"] = failed_text
        return line
    line.update(fields)

    markdown = ocr_markdown(path)
    line["currency"] = currency_hint(markdown)
    line["tax_lines"] = tax_lines(markdown)

    if "seller_name" not in fields:
        line["flags"].append("merchant name not found")
    if line["currency"] is None:
        line["flags"].append("currency not detected")
    if not purpose.strip():
        line["flags"].append("business purpose missing")

    age = (date.today() - date.fromisoformat(fields["receipt_date"])).days
    if age > SUBMIT_WINDOW_DAYS:
        line["flags"].append(f"submitted {age} days after purchase")

    key = (fields.get("seller_name", "").lower(), fields["receipt_date"],
           fields["total_amount"])
    if digest in seen_keys or key in seen_keys:
        line["flags"].append("possible duplicate")
    seen_keys.update({digest, key})

    if fields["total_amount"] >= LARGE_RECEIPT and not line["tax_lines"]:
        line["flags"].append("no tax or VAT line found on a large receipt")
    return line


if __name__ == "__main__":
    seen = set()
    result = build_claim_line(sys.argv[1], "emp-1042", "Client dinner, Cluj", seen)
    print(json.dumps(result, indent=2, ensure_ascii=False))

Run it with python3 claim_line.py dinner-receipt.jpg. A clean receipt comes back with an empty flags list. A faded one comes back with a flag and the OCR text, ready for a reviewer. In a real system, seen_keys would be a database lookup across all open and paid claims, not a set in memory.

A few design choices are worth copying even if you rewrite the code:

  • The file hash catches the same image submitted twice. The merchant, date, and total key catches the same receipt photographed twice, or once as a photo and once as an emailed PDF.
  • Tax lines stay as raw text. Receipts print VAT in many layouts, sometimes with several rates. Show the lines to the reviewer instead of guessing which number to post.
  • Currency is a hint, not a fact. A $ sign can mean several currencies. When it matters, match the claim line to the card transaction, which carries the settled amount and currency.

Check and approve: send people the exceptions

Once claim lines are data, most of the checking is code. A starting rule set:

Check Source What happens on failure
Date and total present /ocr.json required fields Line goes to review with OCR text
Submitted inside policy window receipt_date vs. submission date Flag for the approver
Not a duplicate File hash, merchant, date, total Hold until finance clears it
Card transaction matched Card feed export Mark as out-of-pocket or ask the employee
Over a per-category limit total_amount plus employee-chosen category Route to a second approver
Business purpose and attendees Employee input Return to the employee

The approver should see three things side by side: the receipt image, the extracted values, and the flags. Their job is judgment on the unusual cases (an expensive dinner, a weekend hotel night, a receipt in a language nobody on the team reads), not checking whether 42.50 plus 9.35 equals 51.85.

Keep the rules visible to employees. If people know that a missing purpose returns the claim to them, they fill it in at capture time.

Reimburse and archive: keep the evidence with the money

After approval, the payment runs through payroll or AP like any other. What makes the process defensible later is what you store next to it:

  • The original image or PDF, unchanged, with its hash
  • The extracted fields and the OCR text
  • The employee’s purpose and attendees
  • Each approval, with who approved it and when
  • A link between the claim and the ledger posting, and a retention date

The IRS accepts electronically stored records when the system meets the conditions in Rev. Proc. 97-22: an accurate and complete transfer of the paper record, controls against unauthorized changes, an indexing and retrieval system, legible reproduction, and an audit trail between the general ledger and the source documents. For how long, the IRS page on how long to keep records gives three years as the general period for records supporting a return, with longer periods in specific cases and at least four years for employment tax records. In the EU, Article 247 of the VAT Directive leaves the storage period to each Member State, so check the rule for every country where you reclaim VAT. Talk to your advisor before you shred paper originals.

If you want a self-hosted archive for the stored receipts, the Paperless-ngx tutorial shows how to stand one up, and the accounting post linked earlier covers document types, tags, and folder paths for finance records.

Limitations to plan for

  • Thermal paper fades. Many till receipts are printed on thermal paper. Capture early and treat the image, not the paper, as the record.
  • Phone photos vary. Shadows, folds, glare, and skew lower OCR quality. Expect a higher review rate on photos than on emailed PDFs, and measure it.
  • Handwritten tips and totals. A tip written on a card slip may not match the printed subtotal. Decide whether the card transaction or the slip is the source of truth.
  • Foreign receipts. Text in other languages and scripts still needs a person to confirm what was bought, and currency conversion should come from the card statement or your policy rate, not from the receipt.
  • Itemized lines. The typed catalog covers header values, not line items. If you need to split alcohol from food or separate hotel extras from the room rate, parse the Markdown table or ask the employee.
  • Policy is yours. OCR reads what is printed. It does not know your per diem rates, your category limits, or whether a trip was business travel under the rules that apply to you.

Do not quote accuracy numbers you have not measured. Take one month of real claims, run them through the extraction and check steps, and count how often a reviewer still had to correct a value.

Where OCRskill fits

OCRskill handles the extraction stage: receipt image or PDF in, typed receipt_date, total_amount, and seller_name out, plus Markdown text for currency, tax lines, and line items, with a clear 400 when a required value is missing. It is not an expense management app. It does not store claims, run approvals, match card feeds, pay employees, or apply your policy. Those stay in your expense tool, payroll, ERP, or the small service you build around the code above. For high volumes, paid keys add document URLs and OCR callbacks so receipts already in object storage can be processed without holding connections open.

Conclusion

Paperless expense management is less about the scanner and more about where data enters. If the receipt is captured once, early, and turned into a date, a total, and a merchant without retyping, every later step gets easier: duplicate checks become a lookup, late submissions become a date comparison, and approvers spend their time on the claims that actually need judgment. Keep the original image beside the extracted values and the approval trail, set retention from the rules in each country where you file, and send unreadable receipts to a person instead of letting them pass with a blank total.

To try the extraction step, get a test key from get-key.json, run a few of last month’s receipts through /ocr.json with receipt_date,total_amount,seller_name?, and see how many come back with no flags. Product details and pricing are at ocrskill.com.