KYC Form Data Extractor: Turn ID Packets into Typed JSON
← All posts
GuideSep 29, 2026· 9 min read

KYC Form Data Extractor: Turn ID Packets into Typed JSON

If you searched for a KYC form data extractor, you are not looking for another wall of OCR text. You need named values from identity cards, onboarding forms, and proof-of-address pages so your compliance or onboarding system can validate and store them. Short answer: treat the extractor as a schema-field OCR step: declare the keys your KYC record requires, upload each document image or PDF, get typed JSON back, then run your own checks and review queue. OCRskill does the extraction half with POST /ocr.json?fields=.... It does not replace biometric match, sanctions screening, or a full KYC vendor.

This post is about the document extraction slice of Know Your Customer onboarding. For general multi-field forms (claims, surveys, partner applications), see the form data extraction API guide. For the product overview of typed JSON from images, see structured OCR JSON.

What a KYC form data extractor actually does

In product language, a KYC form data extractor reads customer-submitted documents and returns a record your backend can insert. Typical packet pieces:

  • Government ID card or similar identity page (name, birthdate, nationality, document number, expiry)
  • Proof of address (utility bill, bank letter, or registration extract with a postal address)
  • Onboarding or account-opening form the customer filled (often overlapping name and address fields)
  • Supporting company papers when the customer is a legal entity (company_name and related keys)

Buyers usually want four properties at once:

  1. Named fields that match the onboarding schema (last_name, birthdate, expiration_date, address parts)
  2. Predictable types so dates land as YYYY-MM-DD and strings arrive trimmed
  3. Required vs optional behavior so a missing critical ID field fails closed instead of writing a half-empty applicant
  4. Upload flexibility for phone photos, scanner PDFs, and the same patterns used for plain OCR

Plain Markdown OCR helps humans and search indexes. KYC automation needs columns. That is why “extractor” in this query almost always means structured output, not a text blob. Format tradeoffs after extraction are covered in How OCR can output structured JSON or XML.

Extraction vs verification (do not mix the jobs)

Search results for KYC tools often blur two products:

Job What it answers Typical owner
Document data extraction What values are printed or filled on this page? OCR / document AI API
Identity verification Is this person real, present, and allowed to open the account? KYC / AML vendor or internal compliance stack

Extraction pulls first_name, id_card_series_number, or town_or_city from pixels. Verification covers liveness, face-to-document match, document authenticity forensics, watchlist and sanctions checks, and policy decisions. A solid onboarding design keeps them separate: extract first, then validate and verify with systems built for those risks.

OCRskill is the extraction layer. Wire it beside your compliance tools; do not treat a successful JSON response as a completed KYC decision.

Schema-field extraction beats per-template zones

KYC packets vary by country, issuer, and capture quality. Template engines that read fixed coordinates break when a new ID layout appears or when someone photographs a card at an angle. Schema-field APIs flip the contract: you declare the logical keys you need; the model maps readable content onto those keys without a zone file per issuer.

That matches how OCRskill’s POST /ocr.json works. You pass a comma-separated fields list. Required names fail with 400 when absent. Optional names use a trailing ? and are omitted when missing. Unknown field names also return 400, so you stay inside the documented catalog.

Use discovery mode (omit fields) only while learning a new document type. Freeze a fixed schema before production traffic so response shape stays stable for your validators.

A practical KYC extraction call

Get a test key, declare identity fields, upload an ID image or PDF:

curl https://api.ocrskill.com/get-key.json

export API_KEY="sk-your-key-here"

curl "https://api.ocrskill.com/ocr.json?fields=last_name,first_name,birthdate,nationality,id_card_series_number,expiration_date,issue_date?,id_card_issuer?" \
 -H "Authorization: Bearer $API_KEY" \
 -F "file=@id-card.jpg"

Example response shape:

{
 "last_name": "MUSTERMANN",
 "first_name": "ERIKA",
 "birthdate": "1964-08-12",
 "nationality": "DEUTSCH",
 "id_card_series_number": "T22000129",
 "expiration_date": "2029-08-12",
 "id_card_issuer": "STADT MUSTERSTADT"
}

issue_date was optional and absent on this sample, so it is omitted. Required fields without ? fail with 400 if the page does not support them. For KYC that is usually what you want: better to reject a blurry phone photo early than to persist an incomplete applicant.

Proof-of-address pages lean on the address and company fields:

curl "https://api.ocrskill.com/ocr.json?fields=company_name?,street_name,street_number?,town_or_city,region?,country,invoice_date?" \
 -H "Authorization: Bearer $API_KEY" \
 -F "file=@utility-bill.pdf"

Romanian national ID flows can include cnp and judet when those values appear on the document. Always confirm each key against the OCR JSON API field reference before you hard-code a schema. Maximum upload size is 20 MB per request. Supported inputs include common images (PNG, JPEG, WebP, GIF, BMP, TIFF) and documents (PDF plus common Office and OpenDocument types listed in that reference).

Pipeline shape for KYC packets

Keep extraction inside a short, auditable path:

  1. Capture. Accept uploads into a staging bucket with virus scanning and retention rules that match your privacy policy.
  2. Classify. Route ID cards, address proofs, and filled onboarding forms to different field lists (rules, a light classifier, or human triage for new countries).
  3. Extract. Call /ocr.json with the schema for that document class and Authorization: Bearer sk-....
  4. Validate. Check required keys, date sanity (birthdate in the past, expiry in the future when policy requires it), ID number format in your own layer, and cross-field consistency (name on the ID vs name on the form).
  5. Review. Send 400 failures, low-quality pages, and policy exceptions to a human queue with the original image attached.
  6. Hand off. Pass the typed record into your verification and case-management systems; store the original file for audit.

For searchable archives of the same packets after onboarding, pair structured fields with a DMS metadata model the way industry paperless guides do for ops teams. Extraction still owns the record; the archive owns long-term retrieval.

Validation patterns that keep KYC extractors honest

  1. One schema per document class. Do not reuse an invoice field list on an ID card. Map each inbound type to its own fields string.
  2. Fail closed on identity-critical keys. Treat missing last_name, birthdate, or document number as review, not a silent retry loop.
  3. Normalize once. Dates already arrive as YYYY-MM-DD. Avoid re-parsing locale strings unless you have a second source of truth.
  4. Domain rules stay in your app. Country codes, ID checksums, age thresholds, and PII retention windows are compliance logic, not OCR responsibilities.
  5. Cross-check across pages. If the ID and the address bill disagree on the surname spelling, escalate. Extraction reports what each page says; it does not reconcile conflicts.
  6. Keep Markdown for free-text attachments. Use POST /ocr when a supporting letter needs full-page text for review or RAG, not a fixed schema.

Limitations to respect

  • Blurry or cropped captures. Phone photos with glare, cut-off edges, or heavy stamps still fail. Ask for a rescan instead of burning retries.
  • Handwriting. Printed ID fields are the strong case. Messy pen on paper onboarding forms needs a review path.
  • Field catalog boundaries. Only documented field names are valid. If a passport layout exposes name and dates under the people and ID-card keys, use those; do not invent undocumented keys in the query string.
  • Multi-page packets. Decide whether one request covers a multipage PDF or you split pages and merge keys in your code. Do not assume every page carries every field.
  • Not a full KYC suite. No liveness, no face match, no sanctions list, no document hologram forensics. Pair OCRskill with tools that own those controls.
  • Privacy. Identity documents are sensitive. Minimize retention, encrypt at rest, restrict who can download originals, and follow the laws that apply to your customers.

Choosing a KYC form data extractor

When you compare options, ask production questions:

  • Can I declare required and optional fields without prompt engineering?
  • Do dates and strings arrive normalized?
  • Is there a field reference and clear 400 behavior when a required value is missing?
  • Do uploads match how applicants already send photos and PDFs?
  • Is discovery available for new country layouts before I freeze schemas?
  • Does the vendor honesty-check what extraction does not cover?

OCRskill’s answer for the extraction slice is https://api.ocrskill.com/ocr.json with Bearer auth, name vs name? syntax, discovery when you omit fields, and the same upload styles as Markdown /ocr. Grab a key from get-key.json, lock an ID schema and an address schema, and keep verification in the systems built for it. Broader form OCR patterns live in the form data extraction API post; buyer-oriented image-to-text shapes are in the image to text API guide.

Conclusion

A KYC form data extractor earns its place when it returns a typed onboarding record from messy uploads, fails closed on missing identity fields, and leaves biometric and AML decisions to the compliance stack. Schema-field OCR is the practical core of that extractor: declare the keys, upload the page, validate the JSON, review the failures. OCRskill’s /ocr.json endpoint covers the pixel-to-field step for ID cards, address proofs, and related onboarding documents without per-issuer templates.

If you are wiring this week, start with a small holdout set of real applicant captures, run them through a fixed identity schema, measure how many pages need human repair, then expand country by country. The win is not a prettier OCR dump. It is a KYC record your application can trust enough to pass into verification.