TL;DR
- OCR outputs raw text plus a per-word confidence score; deciding which number is the invoice total or tax ID is a separate extraction step (IDP).
- Thai OCR is harder than English: no spaces between words, stacked vowels and tone marks, and look-alike letters such as ด/ต and บ/ป.
- Tesseract (open source) and Google Cloud Vision both support Thai, but accuracy depends on image quality, so test with 50 to 100 of your own documents.
- Extracted text from ID copies or forms is personal data under Thailand's Personal Data Protection Act of 2019, so limit fields, access and retention.
OCR (optical character recognition) is technology that reads the characters in an image or scanned file and turns them into text a computer can search, edit and pass to other systems. For a business owner, the practical meaning is that OCR cuts the manual typing of invoices, receipts, contracts and business cards into your systems, but it only pays off when the extracted data is checked and flows into the ERP, CRM or marketing tools your team already uses.
This guide is written for owners, finance managers and marketers who have heard the word OCR from software vendors and want to know what it can and cannot do, where Thai text causes trouble, and which documents to start with. No programming background is needed.
What OCR is, explained for non-programmers
OCR stands for Optical Character Recognition. Think of a PDF from a scanner or a phone photo of a receipt. To a computer, those files are just coloured dots arranged into a picture. The words "Grand total" or the figure "12,500.00" are not in there in any searchable form. If you press Ctrl+F in a scanned file and find nothing, the file has not been through OCR.
After OCR, the file gains a text layer. You can search it and copy from it. More important for a business, that text can be handed to another system, for example to pull the tax invoice number, date and amount into accounting software.
One point to understand early: OCR on its own produces raw text. It knows which characters are on the page, but not which number is the total and which is the vendor's tax ID. Assigning meaning is a separate layer, covered in the IDP section below.
How OCR works, step by step
Each OCR product differs in detail, but the main steps are similar. Knowing them explains why some documents read cleanly and others come out full of errors.
- Image intake from a scanner, a phone camera, a PDF or a photo a customer sends in chat. Image quality at this stage affects every step after it.
- Pre-processing straightens skewed pages, removes noise, adjusts contrast and usually converts to black and white to separate characters from the background. Folds, hand shadows and stamps over text make this step harder.
- Layout analysis separates paragraphs, tables, logos and images, then decides reading order. Multi-column pages and nested tables often go wrong here.
- Character recognition. Most modern systems use machine-learning models that read a line or word at a time, instead of matching one character shape at a time as older systems did, so they cope better with varied fonts.
- Post-correction uses dictionaries or language models to fix misread words, then outputs text with its position on the page and a confidence score for each word.
The confidence score is the number a business should care about, because you can build rules on it. For example, any word below a threshold you set goes to a staff member for review instead of flowing straight into the accounts.
How accurate is Thai OCR, and what are the limits?
There is no single accuracy figure for Thai OCR that holds across cases, because it depends on image quality, fonts, document type and the tool. Clean scans of printed documents usually read well. Skewed photos taken in low light, or handwriting, produce noticeably more errors. The honest approach is to test with your own documents instead of trusting a brochure's accuracy number.
Thai has several features that make OCR harder than English:
- No spaces between words. The system must segment words itself after reading characters. Wrong segmentation breaks search and extraction downstream.
- Stacked vowels and tone marks. Upper vowels, lower vowels and tone marks sit above or below consonants. In a blurry image or where table lines cross the text, these small marks disappear or get swapped.
- Look-alike letters. Pairs such as ด and ต, บ and ป, ถ and ภ, ข and ช differ only by a small tail or notch in some fonts or at low resolution.
- Thai numerals and bilingual documents. Some receipts and government forms use Thai digits or mix Thai and English on one line, so the system must handle both.
- Handwriting. Reading Thai handwriting is still inconsistent. If your key documents are handwritten forms, plan for human review every time.
Tools with Thai support include open-source Tesseract, which has a downloadable Thai model, and cloud services such as Google Cloud Vision, which supports Thai. There are also local providers with models built for specific Thai documents such as ID cards or tax invoices. Choose based on test results with real documents, not on brand names.
How OCR differs from AI document processing (IDP)
This is where buyers get confused most often. OCR answers "what is written on this page?" IDP (Intelligent Document Processing) answers "what is this document, and what does each field mean?" IDP uses OCR as its first step and adds several layers on top:
- Document classification decides whether an incoming file is an invoice, receipt, purchase order or ID copy.
- Field extraction pulls vendor name, tax ID, date, pre-tax amount, VAT and total out as named fields.
- Validation checks, for example, that pre-tax amount plus VAT equals the total, that the tax ID has 13 digits, or that the vendor already exists in the system.
- Human-in-the-loop review lets documents that pass every rule flow in automatically while the rest go to a queue for staff.
More recently, multimodal AI models that read both images and text are also used to extract data from documents directly. This handles documents with completely different layouts more flexibly, but it can also return answers that look plausible and are wrong. Validation and human review are still needed whichever technology you use.
In short: if you only need scanned files to be searchable, OCR is enough. If you want data to reach accounting or a CRM without retyping, think in IDP terms from the start. That is as much process design as software selection. See our approach on the document and unstructured data management service page.
Which business documents can OCR handle?
The table summarises documents Thai businesses often start with, what gets extracted and where things usually go wrong, to help decide where to begin.
| Document type | Data usually extracted | Watch out for |
|---|---|---|
| Vendor invoices and tax invoices | Vendor name, 13-digit tax ID, date, pre-tax amount, VAT, total | Every vendor uses a different layout; totals must be checked against line items |
| Staff expense receipts | Merchant, date, amount, expense type | Thermal paper fades quickly; phone photos are often skewed and shadowed |
| ID card copies and customer onboarding forms | Name, 13-digit ID number, address, expiry date | Personal data under PDPA; restrict access and retention |
| Business cards and event registration forms | Name, job title, company, email, phone | Need recorded consent before marketing use; handwriting on forms is error-prone |
| Contracts and old scanned files | Full text for search, parties, start and end dates | Long multi-page files; stamps and signatures over text |
For marketers, the business card and event form row deserves attention. After a trade show, sales teams often come back with a stack of cards. If they are typed into the CRM by hand, follow-up emails go out days later. Reading cards with OCR and sending them into the system tagged with the event name lets marketing process automation, such as a thank-you email or a follow-up sequence, start sooner, provided marketing consent is clearly recorded.
Hypothetical example: time spent before and after OCR
This is a hypothetical case to show the reasoning, not data from a real business. Say a distributor receives 600 vendor invoices a month. An accounts clerk spends an average of 4 minutes per invoice opening, reading and keying it in: 2,400 minutes, or 40 hours a month.
Suppose that after OCR and automated checks, half the invoices pass straight in and the clerk only confirms them (say 1 minute each), while the other half need a few fields fixed (say 2 minutes each). Total time drops to 300 + 600 = 900 minutes, or 15 hours a month.
Real numbers will differ with document quality and vendor count. What the example shows is that the biggest variable is not OCR speed but the share of documents that pass without edits, which depends on validation design and incoming file quality. A good method is to time 50 to 100 real documents before the project, then measure the same way after a month of use.
Connecting OCR to your ERP and CRM
Text that OCR extracts and then leaves sitting in a folder is barely better than the original scan. The value appears when data flows into the systems your team uses daily. A common path looks like this:
- Intake point, such as a shared mailbox where vendors send invoices, a cloud folder, or uploads through chat.
- OCR and extraction reads the document, splits it into fields and gives each field a confidence score.
- Matching compares vendor name or tax ID with master data, matches against open purchase orders and assigns account codes or cost centres.
- Posting to the target system through the ERP or CRM API to create a payable, record an expense or add a new contact, with the original file attached.
- Review queue and edit log. Items that fail the rules go to a person, and a record of who changed which field is kept for audit and for tuning the rules.
Steps 3 and 4 are where most OCR projects spend the most time, because master data in existing systems is rarely clean. One vendor may be stored under several names, or without a tax ID. That connection work is what ERP and CRM integration and synchronisation covers, and if data is spread across several systems, data integration and consolidation into one standard may need to come first.
Check one more thing before investing: some documents do not need OCR at all. If a vendor sends an electronic tax invoice (e-Tax Invoice) with a machine-readable data file, reading that file directly is more accurate than turning an image back into text. OCR suits documents that exist only as images or paper.
Personal data and PDPA
Once documents become text, personal data that was locked inside an image becomes much easier to search, copy and pass around. That is both the benefit and the risk. Thailand's Personal Data Protection Act B.E. 2562 (2019), known as PDPA, applies to this data whether it sits on paper, in an image or in text.
- State the purpose clearly. Data from ID copies collected to verify identity should not be reused for marketing without a legal basis.
- Extract only the fields you need. If you need name and ID number, do not store everything else on the card as text.
- Know where processing happens. Cloud OCR sends document images to the provider's servers. Read the terms on retention, server location and model training, and put a data processing agreement in place.
- Restrict access and set retention. Decide who can see extracted text, and delete both originals and text once they are no longer needed.
- Treat sensitive data carefully. Medical documents or anything with health data fall under stricter PDPA conditions than general data. Consult whoever handles legal matters before automating them.
Starting an OCR project that pays back
For mid-sized businesses, starting small and expanding tends to work:
- Pick one document type with high volume, fairly repetitive layout and someone losing time to it every week. Invoices from a handful of major vendors are often a good start.
- Build a test set of 50 to 100 real documents, including poor-quality ones, so results are not flattering.
- Test 2 or 3 tools on the same set. Count fully correct fields, not characters, because one wrong digit in an amount makes the whole entry wrong.
- Design validation rules and the review queue before connecting to live systems.
- Measure with the same numbers as before: time per document, share passing without edits, and errors that reached the accounts.
OCR is usually the first step of something bigger: turning paper-bound processes into data other systems can use. Seen that way, it belongs in a company-wide digital transformation plan, so each department does not buy its own tool and end up with data that does not connect.
OCR in the Thai business context
Many Thai businesses still receive documents in several forms at once: paper originals, PDFs by email, and photos that customers or vendors send over LINE. That mix makes intake and pre-processing matter more than in many markets. A system tested only on clean PDFs may struggle with a receipt photo sent in chat.
Thai-English mixed documents and government papers that use Thai numerals or Buddhist-era years are the other issue. The system must convert B.E. years into a date format the target system accepts. If that is not configured, dates in the accounting system can end up 543 years off, a real error that appears when foreign software meets a B.E. year for the first time. Testing with real Thai documents matters more than reading a provider's list of supported languages.
OCR FAQ
What is OCR, in one line?
OCR is technology that converts characters in an image or scanned file into text a computer can search and reuse. It reads the text, but knowing which text is the amount or the vendor name needs an extra extraction step.
Can OCR read Thai accurately?
It reads clean scans of printed Thai documents fairly well, but accuracy drops with skewed, shadowed or low-resolution images and with handwriting. Thai has specific difficulties, such as no spaces between words and stacked tone marks, so test with your own documents before deciding.
What is the difference between OCR and IDP?
OCR turns images into text, while IDP uses OCR as the first step and adds document classification, field extraction, validation and a human review queue. If you want data to go straight into accounting or a CRM, you need the IDP approach.
Does using OCR on ID card copies break PDPA?
Not by itself, if you have a legal basis and a clear purpose such as identity verification under a contract. You still need to extract only necessary data, restrict access, set a retention period and check the terms of any cloud provider processing the images.
Where should a small business start with OCR?
Start with one high-volume, repetitive document type, such as invoices from a few main vendors. Test with 50 to 100 real documents, then connect to accounting or the CRM once the validation rules work.
Summary
OCR is where turning paper into data begins, not where it ends. The businesses that benefit are the ones that plan how extracted data gets checked, which system it flows into, and who looks after personal data at each step. If you are weighing this and want a view on where your document processes should start, the Relevant Audience team is happy to talk about document and data management and connecting that data to the systems you already run.







