Data Entry Outsourcing
Dedicated teams with the verification method agreed per field: validation rules everywhere, double entry on critical fields, sampling on the rest.

Page guide
On this page
Accuracy is a process problem, not a typing problem
Data entry is the most commonly outsourced back-office task and the most commonly done badly, because it is bought on the wrong measure. Suppliers compete on speed and price, buyers compare hourly rates, and nobody specifies the accuracy control - so the work arrives fast, cheap and wrong in ways that surface months later.
Fast, careful people still make mistakes. Any team, anywhere, at any rate, has an error rate. The question is not whether errors happen but whether your process catches them before they reach the system of record. That is a design decision, and it is the only thing that actually determines the accuracy you get.
RHI provides dedicated data entry teams with the verification method agreed and measured rather than assumed.
What the team handles
- Document to system entry - forms, invoices, applications, orders, contracts.
- Data migration and cleansing - moving records between systems and correcting them on the way.
- Database maintenance - updating, deduplicating and enriching existing records.
- Product and catalogue data - attributes, descriptions and specifications, often at volume.
- Transcription and digitisation - handwritten or scanned material into structured data.
- Data validation - checking existing records against a source of truth.
- Research and list building - compiling information from defined, permitted sources.
- Exception handling - the records a rule-based process could not resolve, which is where humans genuinely add value.
How verification actually works
Three methods, with different costs and different accuracy. We agree which applies to which field, because applying the most expensive method to every field is waste and applying the weakest one to critical fields is negligence.
Validation rules. The system rejects impossible entries - a date that cannot exist, a postcode in the wrong format, a total that does not reconcile. This is the least expensive control and it should always be exhausted first. Where your system lacks these, adding them usually beats adding checking effort.
Sample-based review. A defined proportion of records checked by a second person against the source. Suitable where an individual error is recoverable, and it gives you a measured error rate rather than a hoped-for one.
Double entry. Two people enter the same record independently and the system compares. Roughly doubles the cost and catches nearly everything, which makes it right for critical fields - financial amounts, identifiers, clinical or legal data - and wasteful for descriptive text.
Most engagements use all three: validation rules everywhere, double entry on critical fields, sampling on the rest.
The source documents are usually the real constraint
Clients tend to assume errors originate with the operator. Frequently they originate upstream: illegible handwriting, ambiguous forms, inconsistent formats, fields that could reasonably be read two ways, scanned pages with parts missing.
Our rule is that an ambiguous source is queried rather than guessed. That means we will return records to you rather than invent a plausible value, and it means the queue includes an exception path. A supplier reporting a very low error rate and never querying anything is guessing on ambiguous records, and you will not find out until the data is used.
We report source-quality exceptions as a category in their own right, because a recurring ambiguity is a form design problem worth fixing at source. Fixing the form removes the error permanently; checking harder does not.
What accuracy target to set
We agree a target field by field rather than one figure for everything, because a name misspelled and a payment amount mistyped are not equivalent failures.
We report the measured rate against it, including when we miss. What we will not do is state an accuracy figure before seeing your source material and agreeing the verification method - the number is a function of both, so quoting it in advance is meaningless. Any supplier offering a headline accuracy figure without asking to see your documents first is describing an aspiration.
Where automation belongs
Optical character recognition and rule-based extraction handle clean, structured documents well and should be used where they fit - paying people to type what a machine can read reliably is waste.
They handle poor scans, handwriting, unusual layouts and genuine ambiguity badly, and they fail with confidence, which is worse than failing visibly. The sensible arrangement is automation for the structured majority with human handling of the exceptions and the confidence-flagged output. That is usually cheaper and more accurate than either approach alone, and we will tell you where the line falls for your material.
Security
Data entry work often involves complete records rather than the fragments a support agent sees, which raises the exposure. Access is least-privilege and role-based, and where the data is sensitive we configure restricted environments - no local storage, no export, no removable media - during scoping rather than afterwards.
Where volume can be processed without identifying information, we will suggest splitting or masking it, because the least risky arrangement is one where sensitive fields never leave your systems. See security and data protection.
Coverage, systems and quality
Data entry is asynchronous, which makes it well suited to overnight processing: work handed over at the end of your day is complete before your next one. Coverage is built from shift units.
Agents work in your systems wherever possible. Reporting covers volume processed, measured error rate by field group against target, rework volume, source-quality exception rate, and throughput per person. Interactions and output are sampled and scored - see quality assurance. Figures published elsewhere on this site are historical or representative and campaign-dependent.
Getting started, and what it costs
Priced per dedicated agent per month with a five-agent minimum - published rates and what moves them. Double entry on critical fields costs more because it is two people doing one job, and that is the correct trade for fields where an error is expensive.
Bring a representative sample of your source documents - including the messy ones, which are the informative ones - your volume, and which fields matter most. Related: office administration, finance and accounting. Talk to us.
Frequently asked questions
We will not state a figure before seeing your source material and agreeing the verification method, because accuracy is a function of both. Any supplier offering a headline accuracy number without asking to see your documents first is describing an aspiration.
Three methods, agreed per field. Validation rules reject impossible entries and are the least expensive control, so they are exhausted first. Double entry - two people entering independently, system comparing - catches nearly everything and roughly doubles the cost, so it is applied to critical fields. Sampling covers the rest and gives you a measured rate.
They are queried, not guessed. We will return records rather than invent a plausible value. A supplier reporting a very low error rate and never querying anything is guessing, and you find out when the data gets used.
For clean, structured documents, yes - paying people to type what a machine reads reliably is waste. OCR handles poor scans, handwriting and unusual layouts badly and fails with confidence, which is worse than failing visibly. Automation for the structured majority plus human handling of exceptions is usually cheaper and more accurate than either alone.
Often not. It frequently originates upstream in illegible handwriting, ambiguous forms or fields that can reasonably be read two ways. We report source-quality exceptions as their own category, because fixing the form removes the error permanently while checking harder does not.
Often yes, and we will suggest it. Where volume can be processed with sensitive fields split out or masked, the least risky arrangement is one where those fields never leave your systems.