FORGOT YOUR DETAILS?

CREATE ACCOUNT

Business OCR software: zone OCR, dynamic OCR, and the engines that read them.

SimpleIndex is Windows OCR software for business. It reads batches of scanned documents and puts the values into index fields, file names and databases, not just text into a file. First you tell it where each value is: in a box you drew (zone OCR, also called zonal OCR), anywhere on the page matching a pattern (dynamic OCR), or by asking the page a question in plain English. Then you choose which engine reads it, from local engines with no per-page cost to cloud and AI engines for the hard fields. Licensed once, with a free trial on your own documents.

01

Five ways to get a value off a page

Each capture method runs on whichever engine you pick in the next section, and most real jobs use two or three together: a barcode to separate the batch, a zone for the values that stay put, a pattern for the ones that move.

Also called zonal OCR or template OCR. Draw a box on a sample page and read whatever falls inside it on every page after. When documents come off one template, this is the fastest and most exact method there is, and there is nothing to configure per document. Large zones with a pattern inside them (Professional) give you room for the skew and shift that scanning introduces.

Uses Fixed layouts · both editions

Find a value by its shape rather than its position. Enter ###-##-#### and SimpleIndex searches the whole page until it finds something matching, which is usually the one social security number on it. Works inside a zone to discard surrounding text, or across the full page when the value could be anywhere.

Uses Structured values · Professional

Give it a list of values that are possible and it searches for each one until it finds a match, for example a vendor list, a department list, or a set of document types. It classifies without training on thousands of samples, because you have pre-defined a set of answers to match.

Uses Known value sets · Professional

New in version 12: Put a question to the page in plain English, "what is the invoice number", and the answer comes back into a named field with no prior zone, pattern, or template. This is what handles the file where two hundred vendors each have their own layout, and none of them warn you when it changes.

Uses Varying layouts · ProfessionalRuns on a cloud engine or a language model, so pages sent this way are billed by the provider.

Read the whole page and write the text back into the PDF as an invisible layer, so the file is searchable in whatever you store it in. It also feeds the other methods: where a page already carries text, zone fields read their values straight from it rather than recognizing the same area twice.

Uses Searchable PDF · both editions
Field 02 which_method

How to choose your own
OCR adventure.

The big difference in zone versus dynamic methods is whether the value lands in the same place on every page. If it does, draw a zone and stop. If your value does not occur in the same place on every page you need dynamic methods to search the page for the desired value.

Many SimpleIndex jobs use both zone and dynamic elements: An invoice batch might read the account number from a fixed header zone, find the invoice number by its pattern anywhere on the page, and match the vendor against a dictionary list. All three methods run within the same job in a single pass.

It's important to build your job for your specific needs: zones tuned for layouts that shift produce a job somebody has to correct by hand every morning, and an AI engine needlessly pointed at a form that has looked the same since 2019 is a per-page bill that can be easily avoided.

Same software, two ways to find the valueZone OCRDynamic OCR
Where the value sitsSame place every pageAnywhere on the page
How you define itDraw a box on a sampleA pattern, a list, or a question
A layout you have not seen beforeNeeds a new zoneUsually already covered
SpeedFastest availableSlower, and slower again on a cloud engine
Cost per pageNoneNone, unless a cloud or AI engine reads it
Right forForms from one templateInvoices, statements, anything that varies

Both columns are SimpleIndex. This is a configuration choice inside one product, not a comparison between two of them.

02

Engines that read the page

We have three OCR engine options with SimpleIndex that run locally on your machine. They all turn pixels into readable text and the recognition is then reported to you. Nothing leaves your network and there is no per-page charge, whichever of the three you use.

The open source engine, included with every edition: accurate enough to make a document searchable and to read a clean, well-scanned zone. Language data now downloads on-demand with version 12, so the installer no longer carries a default gigabyte of dictionaries you may never use.

Uses Included in both editions

The open source engine, included with every edition: accurate enough to make a document searchable and to read a clean, well-scanned zone. Language data now downloads on-demand with version 12, so the installer no longer carries a default gigabyte of dictionaries you may never use.

Uses Included in both editions

The commercial engine: when time is money and it needs to be right the first time, the value it returns is right often enough to trust for document filing. It also reads hand printed characters in defined boxes, with no per-page cost.

Uses Professional
03

Engines that analyze the page

These analysis engines run as a second process, and only on the fields you point them at. That is our purposeful design: the local OCR engines handle the bulk of the needed recognition at no additional cost, and the analyzing engines are pointed only at the fields that need more capability than the local engines can provide.

Reads labeled fields, tables and totals without a template, and reads unconstrained handprint and cursive better than any local engine. Map the labels it returns to your index fields and the values come back wherever they appear on the page.

Uses Plain text · forms · tables · invoices · IDs Professional. Needs your own AWS account, and AWS bills you per page.

New in version 12, and ideal if your organization already runs on Azure, because the account, the billing, and the data residency arrangements are ones you have already approved.

Uses Plain text · layout and tables · forms · invoices · IDsProfessional. Needs your own Azure account, and Microsoft bills you per page.

New in version 12, and the natural choice if your organization already works in Google Cloud. Like the other providers, it has one settings screen with a button that tests your credentials before a batch depends on them, so a wrong key is caught in the wizard rather than three hundred pages into an overnight run.

Uses Plain text · layout and tables · forms · invoices · IDsProfessional. Needs your own Google Cloud account, and Google bills you per page.

Choose a model, paste in an API key, and the job can put questions to a page instead of reading zones off it. Version 12 also lets a model read marks, so checkboxes and filled circles come back as values, and it reports how many tokens each document consumed so the bill is not a surprise. Fields the model has already answered are not read again by the other engines.

Uses Questions · classification · marks · handwritingProfessional. The model provider bills you, by tokens rather than by page.

Point SimpleIndex at a model server you host and the analysis happens only on your hardware, only on your network, and at no per-page cost. It has been tested working on a server with 8 GB of memory, and it runs in an air-gapped system with no internet connection at all.

Uses Questions · classification · marks, with no data leaving your networkProfessional. No per-page charge from anyone.
Field 03 how_the_engines_run

We triage for the lowest cost and
the highest accuracy.

We understand every business has to balance cost against performance, so we triage the recognition: the extra cost of an analyzing engine is only incurred on the fields you send to it.

  1. The included engines go first

    Tesseract, SimpleOCR or FineReader read every page in the batch, on your machine, at no cost per page. Most fields on most documents are finished at this point.

  2. They report what they could not get

    A field that came back empty, or came back failing its format rule, is flagged rather than quietly passed along. Where a page already had a text layer, nothing was re-recognized at all.

  3. You decide what happens to those

    In job settings you say which engine takes which field. The invoice total goes to Textract, the vendor name goes to a dictionary list, the rest is already done. It is a setting, not an automatic escalation.

  4. The answers come back into named fields

    Whatever the second engine returns lands in the same index fields as everything else, and the other engines do not read those fields again. From there it is the same filing, naming and export as any other job.

Field 04 what_the_ai_costs

How am I billed for using analyzing engines?

The SimpleIndex license is bought once and has no per-page cost, whichever method you use and however many pages you run through it. The local engines are included in your SimpleIndex purchase and your only ongoing cost from us is optional support and updates from year 2.

What you pay for the analyzing engines is a bill from AWS, Microsoft or Google on your own account, at their published rates. Because these optional analyzing engines only see the pages the included OCR engines could not finish, the billed page count is usually a fraction of the job, which saves you significant money over using these systems for all data.

List prices for the first million pages a month (AWS US West, Oregon; Azure and Google US pay-as-you-go), checked 30 September 2026. Every provider publishes its own rates, all of them change, and volume discounts apply above a million pages. Check the provider's pricing page before you budget.

Billed by the provider, not by us
SimpleOCR, Tesseract, ABBYY FineReader, or a language model on your own server
No per-page cost
AWS Textract, plain text
$1.50 per 1,000 pages
AWS Textract, tables or questions
$15 per 1,000 pages
AWS Textract, forms
$50 per 1,000 pages
AWS Textract, invoices
$10 per 1,000 pages
Azure Document Intelligence
$1.50 plain text; $10 layout, invoices and IDs (per 1,000 pages)
Google Document AI
$1.50 plain text; $30 forms (per 1,000 pages)
Hosted language model (Claude, GPT, Gemini)
Billed by the provider per token; SimpleIndex reports tokens per document
04

Which edition includes which engine?

Both editions do zone OCR and full page OCR: what separates them is the skill level of the engine reading the zone.

EDITION 01Standard

Manual Processing Focused: SimpleOCR and Tesseract make documents searchable and read clean zones. The right choice for barcode indexing, searchable PDF archives, and for jobs where an operator confirms the values.

  • Tesseract and SimpleOCR engines
  • Zone OCR and full page OCR
  • Searchable PDF output
  • Barcode recognition, DTK and Atalasoft
  • No per-page cost on anything
EDITION 02Professional

Automated Processing Focused: All OCR engines including FineReader for accurate local reading, three cloud provider options, and language models cloud hosted or on your own server. This is the edition where the software reads the values instead of a person typing them.

  • Everything in Standard
  • ABBYY FineReader
  • AWS Textract, Azure Document Intelligence, Google Document AI
  • Language models, hosted or self-hosted
  • Dynamic OCR, template, dictionary and pattern matching
  • Mark recognition and handwriting
ADD-ONServer

Unattended Processing: A watched folder runs the job when files land in it, on any SimpleIndex edition. It changes how the job is initiated, not which engine reads the page.

  • Runs as a Windows service
  • Watches a folder and processes what arrives
  • Attaches to either edition

The engine options are the difference maker between Standard and Professional editions, as every other capability on this page is a configuration choice you make in the same job settings wizard. If you are not sure how much accuracy your documents actually need, describe them to us before you buy anything. It is a short conversation, and then the free trial runs on your own documents.

05

Sounds great, so I just buy your product and 1-2-3 done. Right?

We wish we could say our software is near-magic for all jobs, but there are real limitations to even the most cutting edge technologies and it's important to understand what you can and cannot expect from your software solution.

Cursive handwriting

Hand-printed characters in boxes read well with FineReader at no per-page cost, and it handles neat handwriting outside the boxes better than its reputation suggests. Unconstrained print and cursive read far better with a cloud engine or a language model, though not perfectly, and those charge per page. Flowing cursive prose is still where every method struggles.

Poor originals stay poor

Accuracy is bounded by scan quality and no engine changes that. Faded thermal paper, third-generation photocopies and heavy handwriting over print will need review whether a local engine or a language model reads them. The difference is that the language model charges you per page for the same image-compromised answer. The only fix is to increase the quality of the scans.

The cloud engines need the internet

A page sent to AWS, Azure, Google or any other hosted model leaves your network, and the job stops if your internet connection fails. Only the pages you configure go out to any cloud service, and a processing model on your own server avoids the connectivity question entirely, but somebody has to decide how to handle local versus cloud processing needs.

A per-page bill is a real bill

Your return on investment depends heavily on both the quantity and quality of what you need to capture, so it's critical to work out the estimated page count before you design a job around anything that has a per-page bill, and design it so only the pages that need advanced analysis are sent to that system.

Setup is real work

While a zone OCR job on a fixed form is up and running the same afternoon, a job reading two hundred vendor layouts takes a few sessions to tune, whichever recognition method you use. Anyone promising zero configuration is either over-promising, or describing an entirely different product.

06

How to configure SimpleIndex

Start by checking out our wiki that has screen-by-screen settings help. These are the pages people open most often when they are building a recognition job.

Field 05 watch_one_run

Watch SimpleIndex in action.

Follow along with our short screen captures of real jobs to see setup in action.

Field 06 about_recognition

What people ask about OCR.

What is the difference between zone OCR and dynamic OCR?

Zone OCR reads a box you drew, in the same place on every page. Dynamic OCR finds the value by what it looks like or what sits next to it, anywhere on the page. If every document comes off one template, a zone is faster and more exact. If the layout moves, a pattern, a dictionary list, or a question put to an AI engine handles every variant with one rule.

Standard lists zone OCR. Why would I need Professional for it?

Because zone OCR is only as good as the engine reading the zone. Standard includes Tesseract and SimpleOCR, which are accurate enough to make a document searchable and not accurate enough to trust with a value you are going to file by or post to a database. Professional adds ABBYY FineReader and the cloud and AI engines, and that allows automated indexing with accuracy. If you are indexing by barcode, Standard is the right answer and will stay the right answer.

Do my documents leave my network?

Only the ones you send. The included engines run on your own machine and nothing goes out. The cloud engines and hosted language models do transmit the page, and only the pages a rule you wrote sends them. If that is not acceptable, run a language model on your own server instead: it has been tested working on a machine with 8 GB of memory, and it runs in an air-gapped system with no internet connection at all.

What does the AI cost per page?

The short answer: it depends on your documents. The SimpleIndex license is bought once and has no per-page fee. The cloud providers bill your own account at their published rates, and a model you host yourself costs whatever your hardware costs. Because the analyzing engines only see what the included OCR engines could not finish, the billed page count is usually a fraction of the batch.

Can it read handwriting?

Hand-printed characters in boxes, yes, with FineReader and no extra per-page cost. That covers tax forms, credit applications and anything with a letter box per character, and FineReader handles neat handwriting outside the boxes too. Unconstrained handprint and cursive read far better with a cloud engine or a language model, though not perfectly, and those charge per page. More on handwriting recognition.

Which engine should I choose?

Start with an included one, because it costs nothing per page and it will finish most of the work. Then add a second engine for the fields it cannot get right, and say in job settings which engine takes which field. That is the whole design: one engine reads everything, another handles the exceptions, and you decide which is which rather than the software escalating on its own.

Do I need an AWS, Azure or Google account?

Only if you use that provider's engine. You create the account, generate credentials, and enter them on the OCR settings screen, which has a button that tests them before a batch depends on them. The account is yours and the bill goes to you, which also means the data handling terms you agreed with that provider are the ones that apply.

The only way to know which engine you need is to run it on
your documents.

Send a few samples and we will configure a working job on a call, free. That includes telling you when the included engines are enough, which is more often than you would expect.

You only pay our hourly rate if you buy. Never for the demo itself.

TOP