Key takeaways

I have sat in more provider data cleanup meetings than I can count, and they all start the same way. A member complains that the doctor's office at the listed address does not exist anymore. Someone pulls up the directory. The room spends the next hour arguing about whose fault it is. It is never really anyone's fault. It is a structural gap that healthcare has never gotten around to closing.

Ask a health plan or a large provider group where their data quality budget goes, and you will hear the same three answers almost every time: member data, encounter data, and reporting. That makes sense. Member eligibility errors cause denied claims immediately. Encounter data gaps distort risk scores that finance teams watch closely. Reporting mistakes get noticed by regulators and boards. Provider data does not fail the same way. It fails quietly, in the background, and the damage usually gets blamed on something else, adjudication, the directory vendor, the onboarding team, the new AI model, when the real defect is sitting further upstream in a provider record nobody has looked at in years.

After close to two decades building and untangling the systems underneath value-based care, I would put provider data on the same level of importance as the other three. I rarely see anyone treat it that way until it has already cost them months of rework, a regulator's letter, or a member who could not find a doctor that would actually see them.

A provider record is really a dozen copies drifting apart

The moment a provider gets credentialed, their information does not live in one place. It gets copied into the credentialing system, the NPPES registry, the payer's provider master file, the claims adjudication engine, the member-facing directory, the attribution and network engines used for value-based contracts, and the quality reporting system used for HEDIS and Star measures. Each of those systems reads its own copy, at whatever moment it happened to pull the data, and from that point forward the copies drift independently. None of them talk to each other in real time.

Where one bad provider record actually does damage
1
Claims adjudication. Specialty and taxonomy drive whether a service is priced correctly, whether a place-of-service edit applies, and whether the claim pays as in-network. A stale specialty code produces a wrong answer on all three, quietly, on every claim that provider submits.
2
Provider directory and search. A member searching for a cardiologist finds someone who stopped taking that specialty two years ago, or worse, finds an address the provider left before the pandemic. This is exactly the 'ghost network' problem regulators have spent the last few years chasing.
3
VBC attribution. Panel and attribution logic often leans on specialty to decide whether a provider counts as a PCP, a specialist, or gets excluded from attribution altogether. Get the specialty wrong and a whole panel of members gets attributed, or not attributed, incorrectly.
4
Quality measure eligibility. HEDIS and Star measures frequently use provider specialty to decide which denominator a visit counts toward. A generic or incorrect specialty code can quietly exclude a provider's patients from a measure they should count toward, or include them in one they should not.
5
Network adequacy reporting. Plans certify to state and federal regulators that they have enough of the right specialists within a given distance. That certification is only as good as the specialty data sitting behind it.

NPPES is the closest thing to a national provider directory and it was never built to stay current

The National Plan and Provider Enumeration System, NPPES, is where every provider in the country gets their NPI. It is public, it is free, and it is the closest thing healthcare has to a single source of truth for who a provider is. It is also, by design, self-reported. A provider or their office staff enters the information once at enrollment and updates it only when something forces them to: a revalidation cycle, a credentialing audit, a directory complaint. Nothing in the system prompts a provider to log in and update their phone number just because it changed.

That is exactly why the biggest quality problems in NPPES are not the ten-digit identifier itself, that part is reliable, but everything wrapped around it: mailing addresses, phone numbers, and taxonomy codes that were accurate the day they were entered and have not been touched since. A provider changes practices, retires a location, or picks up a new subspecialty certification, and none of that reliably makes its way back into the record every downstream system is quietly trusting.

Real-world scenario: the phone number that outlived the provider
Enrollment
A provider enrolls with a solo practice address and a taxonomy code for internal medicine. The NPPES record is created. The payer loads it into their master file within the quarter.
18 months later
The provider joins a multi-specialty group across town, picks up a new subspecialty certification, and the old office closes. None of this gets reported back to NPPES. The old address and taxonomy remain untouched.
Downstream
The payer directory still lists the old address. A member calls the old number and reaches a disconnected line, or a different practice entirely. The claims system still prices the provider's services under the old taxonomy, generating incorrect specialty-based pricing edits on every claim.

This is not a hypothetical edge case. It is the default behavior of a self-reported system with no ongoing incentive to stay current. The No Surprises Act tried to force the issue by requiring plans to verify directory listings at least every 90 days and to remove or flag a provider within two business days of being notified their information changed. That is a reasonable rule for the plan's own directory. It does nothing to fix the taxonomy and specialty data sitting in the source system every plan is copying from in the first place.

90 days
Maximum interval plans must verify provider directory listings under the No Surprises Act
2 days
Business days a plan has to update a listing once notified a provider's information changed
Self-reported
How NPPES data is entered and maintained, with nothing forcing it to stay current between enrollment events

A specialty code is a small field with an outsized job. It decides which doctor a member is allowed to find, which price a claim gets, and which measure a visit counts toward. Get it wrong once at enrollment, and it stays wrong on every transaction that touches that provider until someone happens to notice.

There is no golden record, because there is no single source of truth

Master data management has a well-known concept called the golden record: one authoritative version of a piece of data that every system defers to. Provider data does not have one. Credentialing has its own version, built from primary source verification. NPPES has its own version, self-reported by the provider. The payer's claims system has its own version, loaded at enrollment and rarely refreshed. The directory team may maintain yet another version, updated manually after member complaints. PECOS, the CMS system providers use to enroll in Medicare, has its own version too, and it frequently disagrees with NPPES on something as basic as which specialty is primary.

None of these systems is wrong, exactly. Each is accurate to its own source and its own update cycle. The problem is that nobody has designated one of them as authoritative for everything, and even if someone did, nothing forces the others to reconcile against it. Matching a provider across systems mostly depends on the NPI as the join key, which works fine when everyone actually uses it consistently. In practice, legacy internal provider IDs, tax IDs, and old paper-based enrollment records still show up, and matching a provider across four systems by name and address alone is exactly as unreliable as it sounds, especially in a large multi-specialty group where six providers share the same suite number.

Why PECOS matters more if your book of business is Medicare Advantage

This is not an abstract data governance problem for most VBC organizations. If your book of business runs through Medicare Advantage, and most value-based arrangements do, PECOS is not just one more system with its own copy of the truth. It is where CMS keeps its own record of who is actually enrolled in Medicare, on what terms, and under what specialty.

PECOS enrollment data includes whether a provider accepts Medicare assignment in full, on a limited case-by-case basis, or not at all. A provider who only assigns claims on a limited basis looks identical to a fully participating provider in most internal systems, right up until a claim processes differently than the network file assumed. PECOS also carries its own specialty taxonomy, sourced separately from the taxonomy code sitting in NPPES, and the two do not always agree on which specialty is primary, alongside enrollment details like graduation year, medical school, and the group practice name the provider is actually billing under.

Real-world scenario: the assignment status nobody checked
Directory
A health plan's Medicare Advantage directory lists a provider as an in-network primary care physician, fully participating, same as every other PCP on the panel.
PECOS
PECOS shows that provider's Medicare assignment status as limited, case-by-case, not full participation. Nobody on the plan's provider ops team has ever pulled this field.
Claim
An MA member sees the provider expecting standard in-network cost sharing. The claim does not process the way the directory implied, and the member ends up disputing a bill nobody at the plan can immediately explain.

None of this is theoretical, and none of it requires guessing. CMS publishes PECOS enrollment data through its Provider Data Catalog, the same Medicare Physician Compare dataset behind the public directory tools, as a free public API. It is the exact data CMS itself references when it audits network adequacy and directory accuracy for Medicare Advantage plans. A provider record that has drifted from PECOS has drifted from the same source the regulator is checking against, which makes reconciling against it a compliance question, not just a data hygiene one.

Legal name versus DBA: the mismatch nobody resolves

Every provider organization has at least two names: the legal business name that shows up on tax documents and contracts, and the DBA, doing business as, name that shows up on the sign outside the building and in patient-facing materials. NPPES has fields for both. Most downstream systems do not consistently pick the same one.

A claim gets submitted under the legal entity name because that is what the billing system was configured to use. The payer's adjudication system is expecting the DBA name on file from the group's contract. The names do not match cleanly, and depending on how strict the matching logic is, the claim either pends for manual review or denies outright. Meanwhile, a member gets an explanation of benefits listing a legal entity name they have never heard of, because they only ever knew the practice by its DBA. They call the plan to dispute a charge from a doctor they insist they never saw. The doctor they saw is the right doctor. The name on the paperwork is just the wrong version of a name that was never wrong to begin with.

The provider did not do anything incorrectly. Two systems just decided, independently and reasonably, to trust a different name field for the same NPI. Nobody designed that conflict on purpose. Nobody designed a way out of it either.

No two clients structure their provider hierarchy the same way, and onboarding pays for it

This is the one that costs the most real time. Every provider organization has some version of a hierarchy: individual providers roll up into locations, locations roll up into a clinic or a group, groups roll up into an IPA or an MSO, and the whole thing may roll up into a health system. What that hierarchy looks like, how many levels it has, which entity actually bills, which entity gets credentialed, which entity the provider is contractually attributed to, is different at literally every client I have onboarded. There is no standard schema for it the way there is at least a defined 837 loop structure for a claim.

So every new client integration starts the same way: weeks, sometimes months, of manually mapping their specific provider-to-entity structure before a single record can load correctly. One client attributes a provider directly to a billing NPI. Another attributes through three levels, provider to location to group, with different rules at each level for which fields are inherited and which are overridden. A third has providers who bill under one entity but are credentialed under another because of a merger nobody fully unwound in the data. None of this is documented anywhere consistent. It lives in whoever built the original integration's head, the same companion guide problem we have written about before, except this time the deviation document does not exist at all. It is tribal knowledge, and it has to be rediscovered from scratch with every new client.

What actually helps

Treat the NPI as the only reliable join key, and stop trusting name-and-address matching

Every process that tries to reconcile a provider across systems should anchor on the NPI first. Name and address matching should be a fallback for the rare cases where the NPI is missing or wrong, not the primary strategy. It sounds obvious. Most legacy integrations still do it backwards because the NPI was bolted onto their schema years after the name and address fields already existed.

Pull specialty and taxonomy from a live source, not a one-time enrollment snapshot

Whatever specialty and taxonomy value got captured at enrollment should be treated as a starting point, not a permanent fact. Cross-checking it against NPPES and PECOS on a recurring basis, and flagging the mismatches instead of silently trusting whichever value loaded first, catches the drift before it turns into a denied claim or a misrouted member. This is worth doing at the API level, not on a quarterly spreadsheet cadence. CMS's Provider Data Catalog is a free public API, so there is no excuse for querying it once at onboarding and never again. The moment a provider is added to a roster is the moment to check it, alongside an OIG exclusion check, and again on a standing schedule after that.

Separate legal name and display name everywhere, and pick deliberately which one each audience sees

Claims and contracts should consistently use the legal name. Directories, patient communications, and anything member-facing should consistently use the DBA. The fix is not complicated. It just requires treating them as two distinct fields with two distinct purposes everywhere in the system, instead of collapsing them into one 'provider name' column and hoping the right value happened to get entered.

Document the hierarchy before onboarding starts, not while it is happening

Before a new client's data ever gets mapped, get their provider-to-entity hierarchy in writing: how many levels, what each level represents, which entity bills, which entity gets credentialed, and which rules govern attribution. It takes a few focused conversations up front. It saves months of reverse engineering later.

Revalidate on a schedule, not in response to a complaint

Directory accuracy and taxonomy accuracy should be checked on a standing cadence, independent of whether a member has complained yet. By the time a complaint surfaces, the bad data has usually already caused a misrouted claim, a frustrated member, or both.

If PHI review is what is slowing your data roadmap, start here

There is also a practical reason provider data is worth prioritizing even ahead of some higher-visibility initiatives. NPI, taxonomy, specialty, addresses, and organizational hierarchy are not member data. None of it is PHI. A provider data quality initiative can typically move through procurement, legal, and security review far faster than anything touching claims or clinical records, while still fixing a defect that quietly degrades claims adjudication, directory accuracy, and attribution at the same time. It is one of the few places in healthcare data where the fix is both high-impact and genuinely fast to start.

It is time provider data got the same standardization push claims data is getting

We have argued before that healthcare's data standards are not really standards once every payer's companion guide is allowed to override them. Provider data has the same disease, without even the companion guide as a starting point. There is no equivalent implementation guide that tells every system exactly which fields describe a provider, in exactly what format, sourced from exactly which authoritative registry. NPPES, PECOS, state licensing boards, and every payer's own master file all describe the same providers, in slightly different ways, and nothing forces them to agree.

FHIR's Practitioner, PractitionerRole, and Organization resources, and CMS's provider directory API requirements under the interoperability rules, are a real start. They give systems a common shape to exchange provider data in. They do not yet solve the upstream problem: the data going into those APIs is still only as fresh and as accurate as whatever self-reported, rarely revisited source fed it. A common format around stale data produces stale data faster, in more places.

A real fix looks less like a new API and more like a real obligation: a single, federally maintained registry that providers, or their credentialing organizations, are required to keep current on a fixed cadence, with taxonomy, specialty, and contact details treated as seriously as the NPI itself already is. Until that exists, every plan, every clearinghouse, and every provider group will keep independently reinventing the same reconciliation work, badly, at their own expense.


Provider data problems do not make headlines the way a denied claim or a botched risk score does. Nobody escalates a call because a taxonomy code was three years stale. But add it up across a health plan's entire network, and it is sitting underneath a meaningful share of the claims denials, member complaints, and integration delays that do get escalated, wearing someone else's name.

The fix is not glamorous. It is treating the NPI as the only trustworthy key, refreshing specialty and taxonomy against live sources instead of enrollment-day snapshots, keeping legal and display names as separate fields with separate purposes, and documenting the provider hierarchy before onboarding starts instead of during. None of it requires new technology. It requires deciding that provider data deserves the same attention member data, encounter data, and reporting already get.

Until the industry gets a real standard for it, that attention is the only thing standing between a clean provider record and the next member who calls a disconnected number looking for a doctor who moved two years ago.

Is your provider data quietly breaking things downstream?

ProviderPulse checks every provider in your roster against NPPES, PECOS, and OIG/SAM.gov exclusion data the moment they are added, flags specialty, taxonomy, and Medicare assignment mismatches before they hit a claim or a directory, and gives your credentialing team a single queue to fix what is actually wrong.

Book a Call Send a Message
Ajay Chaudhary
Founder & Principal Consultant, Vavion Health

Many years working inside healthcare's data and technology infrastructure. Value-based care contract operations, claims data pipelines, EDI and FHIR integrations, and the analytics workflows that determine what gets measured, what gets acted on, and what gets missed.