Deep Dive · Vendor evaluation

Looking Past Coverage: What Really Drives Your Match Rate

Learn why region-specific match rate is more important than coverage claims when it comes to evaluating data vendors for identity verification workflows.

Data Zoo9 min readUpdated September 2026
On this page

Key takeaways

  • Coverage describes how much of a country’s population a vendor’s sources can reach. Match rate depends on the segments within that population you actually serve.
  • Identity infrastructure is local: Mexico spreads identity across separate population, electoral and tax registries, while Colombia centralizes it in one civil registry. Regional knowledge is part of what you are buying.
  • A quoted match rate is only as meaningful as the match rules behind it — ask which population, what counted as a match, and what was in the denominator.
  • Source order, fallback rules and what counts as a pass should be your decision, because your organization is the one defending it. A good data partner configures that logic to your requirements and returns results you can interrogate.

If you’re building identity verification workflows, you’ll find that most data vendors lead with coverage, like how many sources they connect to, how many countries they reach, and how many consumers they can verify. And coverage is a reasonable place to start. It tells you whether a source exists in a market and whether the vendor can reach it.

What a coverage figure describes, though, is how much of a country’s population a vendor’s sources can reach. Your match rate depends on the segments within that population you serve, and those are two different questions. That’s why a vendor with strong coverage in a given country can still return thin results for the customers you’re onboarding.

The distinction changes how you evaluate a global identity data partner.

Where coverage and match rates diverge#

A vendor might support dozens of countries, connect to hundreds of sources, and cover billions of consumers, and your match rate can still come in lower than you expected. Two things account for most of that difference.

Identity is local#

Every country has its own registries, credit bureaus, telcos, government sources and document types, so coverage in one market can mean something structurally different from coverage in the next.

Mexico and Colombia, both major markets in Latin America, illustrate the range.

  • In Mexico, identity is spread across separate government registries, each tied to its own identifier:
    • The Registro Nacional de Población (National Population Registry, or RENAPO) assigns the Clave Única de Registro de Población (Unique Population Registry Code, or CURP).
    • The Instituto Nacional Electoral (National Electoral Institute) issues the voter credential.
    • The Servicio de Administración Tributaria (Tax Administration Service) issues the Registro Federal de Contribuyentes (Federal Taxpayer Registry, or RFC).
    Which registry can verify a customer depends on which identifier they present.
  • Colombia centralizes citizen identity instead. A cédula de ciudadanía (citizenship card) is checked against a single register, the Registraduría Nacional del Estado Civil (National Civil Registry).

In practice, this shapes how you design your onboarding flow. The identifiers you collect determine which sources you can verify against, and the right choice differs from one market to the next.

This is why regional knowledge is part of what you’re buying from a data vendor. If you’re entering a market your team hasn’t operated in before, a vendor who already understands its source landscape shortens that learning curve considerably.

You’re targeting a segment within a population, not the whole country#

A vendor’s country-level coverage can be entirely credible and still produce a match rate below what you expected if the sources behind it hold thin records for a segment that matters to you.

For example, migrants often have thin records in the most commonly used sources, simply because they haven’t yet built much history in that market. And younger customers — common in crypto and buy-now-pay-later use cases — frequently have little to no credit footprint at all.

Neither group is well served by the sources a verification flow tends to reach for first. The solution is usually a different source altogether. Going back to our example, superannuation and payroll data in Australia often reaches the migrants that standard sources miss, and telco data in the United States verifies younger users who might otherwise be declined.

Data Zoo has seen this play out directly. A global payments company handling cross-border remittances (a service popular with underbanked and unbanked immigrants sending money home to their families) was running a traditional, single-source verification flow built on datasets like credit bureaus. Its five priority markets were the United States, France, Germany, Spain, and Italy — all countries where credit bureau coverage is strong on paper. But the company’s target customers had little to no financial history in those files, so verifications failed at a high rate and showed up as a weak pass rate at onboarding.

By the company’s own account, every 1% drop in pass rate was costing it millions in annual revenue.

The vendor’s coverage figures for those markets were accurate, but they represented a population the company wasn’t onboarding. So, after evaluating several alternative identity data providers, the company partnered with Data Zoo.

Our solution delivery team worked with them to map their customer demographics in each market, identify where the existing flow was losing customers, and build a strategy for lifting pass rates in each one. The payments company’s own team then configured the flow it wanted, sequencing across several complementary sources instead of depending on credit bureau data alone, and has continued to fine-tune it since.

The result was an average uplift of roughly 5% in approval rate per market, more than $30 million in new annual revenue, and 166,000 additional customers. The company put the value of a single percentage point of new-user conversion at around $1.1 million per year.

The same logic holds no matter which region you’re building for. If your customer base skews toward a segment a given source wasn’t built to verify, your match rate will lag behind a vendor’s headline coverage number. But if you can pinpoint the segments your customer base actually skews toward, you can evaluate vendors on the match rate they’re able to demonstrate for your target segments in the markets you’re entering.

What to ask when a vendor quotes a match rate#

Asking for a demonstrated match rate is the first step to a more informed vendor evaluation. But it’s incomplete on its own, because a match rate is only as meaningful as the rules used to calculate it.

Match rules define what counts as a match in the first place, and different rules can produce a different match rate even when the underlying data is the same. In effect, two vendors quoting figures on the same data can land a long way apart without either of them being dishonest.

That’s why useful follow-up questions are all about methodology:

  • Which population does the figure describe? A general-population rate tells you little about a customer base weighted toward a particular segment or market.
  • What rule defines a match? A rate built on a composite score is not comparable to one built on individual response elements.
  • What was in the denominator? Records that failed pre-validation are sometimes excluded, which lifts the figure without any source having performed better.

Better still, agree to the rules before the test. Running a proof of concept against your own data, with success criteria set in advance, produces a figure that means something for your build and avoids comparing two numbers that were never calculated the same way.

With Data Zoo, those rules are yours to set. Each source response is scored on which identity attributes agreed. You define which combinations of those scores clear the bar for a match, at either the organization or the user account level.

Go deeper: how Data Zoo structures match results and scores.

Why you should stay in control of verification logic#

The right combination of data sources changes with every market you enter and every customer segment you serve, and neither of those factors stands still. That’s why your team should decide which sources you call, in what order, and what counts as a pass.

There’s a second reason to stay in control. Your organization, not your data vendor, is responsible for defending a verification decision to a regulator, an auditor, or a risk committee. That’s easier when the logic behind it reflects decisions your team made and can explain.

That control should take a few concrete forms:

  • Data source sequencing should reflect your requirements. If one source cannot produce a match, the verification should automatically retry in real time against the next best source, following rules set to your requirements. This is the key to lifting the overall match rate.
  • Sequencing should account for how sources perform in practice. If a source is slow to respond, the flow should move to the next source after a set wait time. A legitimate customer shouldn’t be held up by one source’s delay.
  • Additional checks should be applied selectively. Structured, attribute-level results give your orchestration layer what it needs to decide when a customer needs a further step, such as a document or biometric check. Most legitimate customers can then move through the workflow with no added friction.

Thresholds are the other half of the picture. A threshold defines what counts as a match, and it works at two levels. At the attribute level, it sets how closely a name or address must agree with the source record. At the verification level, it sets which combination of matched attributes is enough to verify an identity. An exact match on a government ID number and date of birth is stronger evidence than a close match on a street name, and your rules should reflect that.

What matters is that the level reflects your risk appetite. Set it too low, and you risk waving through a bad actor. Set it too high, and you risk rejecting customers who should have passed. Your compliance team should decide where that line sits, and your data partner should apply it consistently to every check. That’s only workable if what comes back is a structured result you can interrogate — an outcome for each attribute checked and which source produced it — rather than a single pass or fail with no way to see how it was reached.

The advantages of controlling your own verification logic#

The first advantage of controlling your own verification logic is a verification decision you can easily defend. When a regulator or auditor challenges a verification, they want to see which sources were called, in what order, and what each one returned. If the only evidence of a decision sits in several sources’ separate logs, you’ll be reconstructing it under pressure. That record belongs in your system rather than your vendor’s, since you’re the one who has to produce it — sometimes years after the fact.

What a data layer owes you is the raw material to build that defense, including a structured, attribute-level result on every check, with the source behind each result, in a form you can log and retrieve on demand. What it shouldn’t be doing is holding your customers’ personal data any longer than the check itself requires.

The second advantage is agility. Markets change, your product changes, and the mix of people signing up changes with them. A source that performed well for last year’s customer base may be a poor fit for this year’s. When you control which sources you call and in what order, responding to that is a configuration change rather than an engineering project. You can add a source for a new segment, reorder the sequence for a market that has shifted, and leave the rest of the logic untouched.

But while control should sit with your team, you shouldn’t have to work alone. Shifts in your customer mix don’t announce themselves, and the source landscape moves as well, with new options coming online in markets where they used to be thin. The right vendor not only knows that configuration should be revisited regularly, they help you do it by watching match rate performance market by market, raising a flag when results start drifting, and bringing the local knowledge of which source or sources would close the difference.

Own your decisioning#

High-performing identity verification workflows need a data layer built to be authoritative, transparent, and configurable from the ground up.

Data Zoo provides authoritative identity data that helps organizations verify individuals accurately in regulated markets. Through one API, we match identity attributes against independent and reliable data sources and return structured results that customers and partners can apply within their own decisioning logic.

Evaluate us on your own data

Set the match rules before the test and run a proof of concept against the segments you actually onboard.

Frequently asked questions#

Coverage describes whether a data source exists and is reachable in a given market. Match rate describes the percentage of your customers who verify successfully against it once your flow is live. Coverage is a useful starting point, but it describes a population on average, while match rate depends on where your specific customers sit within that population. A source can have full national coverage and still return a weak match rate for a particular segment.

A match rate can differ between vendors because of the match rules used to calculate it, as those rules vary. A rate based on a composite score will differ from one based on individual response elements. A rate that includes records which failed pre-validation will look lower than one that excludes them. Neither figure is necessarily wrong, but they aren’t comparable. When you compare vendors, establish what counted as a match, what was included in the denominator, and whether the figure reflects your customer profile or the general population.

Identity verification needs to be handled differently by country because the documents people hold, the data sources that are authoritative, and the demographic makeup of a population all vary from market to market. A verification approach tuned for one country’s identity infrastructure — its ID document types, government registries, and credit bureau depth — will not perform the same way in another. Treating verification as local rather than global is what keeps match rates high as you expand.

Data Zoo differs from other identity data vendors by focusing on authoritative identity data. It provides the data layer beneath verification, fraud and decisioning workflows, and many identity platforms build on it to extend their reach into new markets. That focus supports deeper knowledge of each market’s sources, which helps customers achieve more reliable verification outcomes across jurisdictions. Every check returns a structured, attribute-level result that your team can apply within its own decisioning logic.

With Data Zoo, decisioning stays with the team building on it. Data Zoo brings the authoritative data sources and the infrastructure to route and sequence checks against them. Your team defines which sources are invoked, in what order, and under what conditions, and Data Zoo executes that logic automatically and consistently across every identity check.

Data Zoo is not a replacement for your orchestration platform, document verification tool, biometrics, or fraud tooling. It’s the authoritative data layer underneath them. It connects your existing stack to 150+ authoritative data sources across 40+ countries through a single API, so the tools you already use have better material to work with.