Uncovering missing and mislabeled data values in your employee, historical, and applicant data is an essential part of your workforce compliance and ongoing metrics monitoring. In this article, we look at three major types of common data error—the 3 “I”s.
The 3 “I”s Defined: How to Recognize Affected Data
The Affirmity team handles tens of thousands of datasets every year. This work has shown us that it’s actually extremely rare for a dataset to be perfect and ready to use when freshly extracted from an organization’s HRIS. A thorough review will almost always reveal data issues, which can be characterized under the three “I”s:
- Incomplete
- Inaccurate
- Inconsistent
Let’s take a closer look at each category of error.
1) Incomplete Data
Incomplete data is data that is missing from your record—it may have been lost at some point in your process, or never captured in the first place. Below, we’ve considered the kinds of data you will encounter across three key file submissions—your employee, history, and applicant file.
Employee File
The following fields should typically be included in your employee file submissions:
- Employee ID
- Employee demographics (race, gender, disability, and veteran status—remember these are voluntary fields. For state and federal reports requiring this information, you will need to establish a business process to gather the necessary information).
- Job titles/profiles
- Grouping variables (typically by job title or job profile—this is how you will review and compare certain groups of people)
- EEO codes
- Location
- Compensation
- Unique plan variables (example: business line, department, executive leadership line, etc. )
- State reporting variables (example: remote work state, state job group—some states, such as Minnesota, have specific job-specific groupings they utilize)
- Union status (if applicable—because of factors such as collective bargaining agreements that will require particular populations to be analyzed differently)
Keep in mind that federal contractors are also required to solicit their current workforce for disability status every five years. A self-service portal function will help with this process.
History File
As with the employee file above, a typical history file has a number of expected fields that should be checked for early in the process of data reconciliation:
- Employee ID
- Employee demographics
- Transaction code (hire, transfer, promotion, or termination)
- Transaction reason (Most critically, was it a transfer which resulted in a change in compensation, or was it just a title change or miscellaneous data correction only?)
- “From” and “To” variables for promotions and transfers
- Job title/profiles
- Grouping variables
- Locations
- Unique plan variables (e.g. business line, product, or service offering—this information helps with historical tracking)
- Union status (if applicable—as above, union members follow different business processes and need to be analyzed separately)
- State reporting variables (if applicable)
Of particular note in this list are the “From” and “To” variables. When someone starts in a particular job, with a particular EEO code and other associated attributes, there needs to be a way of easily linking them to later records when they were transferred or promoted and those variables changed.
Applicant File
This file will be composed of data pulled from your Applicant Tracking System, and can be relatively time-consuming to work with simply because of the large volume of information that flows into your ATS. A typical output will need:
- Requisition ID (helps differentiate different job postings for the same job in the reporting period)
- Requisition title
- Application date
- Applicant/candidate ID (needed in order to understand whether an individual applied to multiple requisitions, or potentially even applied to the same requisition more than once)
- Name
- Applicant demographics
- Final stage
- Final status
- Final reason
- Recruitment source (should be collected to inform a sourcing effectiveness analysis)
- Hired employee ID (because people can sometimes apply with a different name to the one they’re hired with)
- State reporting variables (example: test date, interview date, etc.)
BEST PRACTICES FOR CURRENT COMPLIANCE | ‘How to Navigate the New Patchwork of Employment Law’
2) Inaccurate Data
Moving on to the data types that aren’t straightforwardly missing, “inaccurate data” describes data that is present but untrue due to a mistake made during input or collation. Here are seven common examples:
i) Employee Demographics
Some systems have taken steps to allow employees to better reflect their heritage-they allow for the input of multiple races instead of simply using the “two or more” races option. This can cause confusion about the accurate race data to use for the analysis.
ii) Attributes of Job Titles/Profiles
Each job title should have only one EEO code and grouping variable associated with it. Similarly, salary grade requirements should be the same within your job categories. Checking consistency throughout ensures your analytics are run appropriately based on the intended categories and groupings.
iii) Locations and Office Addresses
Every employee needs a correct work location assignment—something that has become more complex given the rise in remote working arrangements. More often than not, you are required to associate an employee with a specific establishment (i.e. they cannot simply be “working from home”).
iv) Transaction Reasons and Codes
You must ensure that the most accurate transaction reasons have been selected in your system. Some systems have generic-sounding reasons such as “changes” that, in our experience, tend to get selected even though a more specific and technically accurate designation would be better (e.g. a promotion or transfer).
v) Job Transaction Continuity
In some cases, the current snapshot may include an individual who didn’t appear in the last snapshot, but there is no hire record to explain their appearance. Or conversely, someone appears in the older snapshot but not in the more recent one, and there’s no termination record. This naturally requires further investigation to add or correct records appropriately.
vi) Transaction Timelines and Backdated Transactions
It’s important to understand and properly reflect when transactions actually occurred, rather than when they were changed. For example, your last edit date may be within the scope of the timeline you’re considering, but the transaction itself could be from before that range.
vii) Transfer Codes as Hires and Terminations
When people transfer or get promoted into different departments, they can sometimes be erroneously recorded as a hire in the new department (or potentially, as a termination in their original department). You cannot hire someone who already works for your organization, so this transaction will need to be moved into the correct category.
LEARN ABOUT THE LEGAL BASIS FOR USING YOUR WORKFORCE DATA | ‘What the EEOC’s “Record-Breaking” 2025 Results Mean for Your Organization’
3) Inconsistent Data
Inconsistent data is perhaps less apparent at first glance, and only becomes obvious as you compare records captured at different times or on different systems. If you’re getting contradictory information from different sources, you have to work to identify the reasons and remedies for those inconsistencies.
Inconsistency Between Employee and History Files
The following are the most common inconsistencies we find ourselves going back to clients with questions about:
- Employee demographics: An employee has a different race in the employee and history files. The data was entered incorrectly at one stage, but you will need to go back and determine when.
- Location/address changes: These may legitimately differ because of a change in reporting lines, a transfer, a promotion, or an all-office move to new premises.
- Job title/profile attributes: As above, differences may be the result of transfers or promotions that have no transaction to help explain them.
- Job transaction flow: Sometimes, there will be “From” and “To” mismatches that need further investigation
- Duplicate transactions: Was a duplicate made in error, or is something else going on?
- Same-day transactions: If someone accepts an offer but is a no-show, they may have a hire and termination record on the same day.
- Freeform data fields (non-uniform output): Abbreviated job titles (e.g. Sr. versus Senior), and variant names (e.g. Kaylee, Kayleigh, Kaleigh, etc.) require standardization.
- HRIS/ATS transitions during plan year (multiple systems): If you change systems sometime between the creation of your employee and history files, there may be some additional work to draw the line between different records.
- Data impacted by mergers, acquisitions, divestments, etc.: Employees gained through these activities are often recorded as “Hires”, but they should ideally be differentiated (e.g. an “Acquisition Hire”) and not factored into any analysis about your selection decisions.
Inconsistency Between Applicant and History Files
When Affirmity receives a history file, the team will try to hunt down hires from the corresponding hires file and match them in the applicant tracking system—looking for a requisition with a similar title to what is shown in the hires file. However, we may then run into the following problems:
- The employee is flat-out missing from the applicant tracking system—they cannot be found via name or hired employee ID.
- We can see the employee is hired in the ATS, but it’s for a completely different job than in the HRIS data.
- The employee appears to be dispositioned incorrectly in the ATS—they’re still listed at the offer stage, or they’re recorded as not hired in some other disposition.
In these situations, we will typically have to return to the client with questions to ensure that we have the correct requisition, identify contributing factors, and correct any omissions.
We may also encounter applicants marked as “hired” in the applicant tracking system that are not in the hires file. These often turn out to be internal hires.
Finally, we also question misaligned application and hire dates. If the dates are nearly a year apart, that may indicate an error or that we’re looking at the wrong records.
Inconsistency Within Your Applicant File
We commonly run into yet more inconsistencies within the applicant data itself.
- Duplicate records or multiple applications for the same person in the same requisition: Perhaps the candidate applied using two different profiles, giving them two different candidate IDs. We would typically need the talent acquisition team’s help in identifying these duplicates.
- No applicants corresponding to a hire: Having a single applicant and a single hire, raises a red flag. There may be an evergreen or sourcing requisition from which you’ve spun off child requisitions leading to more hires. You need a way of recording the relationship between these requisitions.
- Requisitions with no hires: This can result from a requisition still being active at the time the data is pulled, an internal placement, or the requisition simply not resulting in a hire (either because it’s cancelled or otherwise unsuccessful)
- Hires with no requisitions: Similar to the above, your analysis period may include hires that resulted from requisitions before the analysis period started. Providing requisition data for three months before the analysis period can provide useful extra context.
- Inconsistent disposition codes: This can often be the result of late dispositioning or “mass dispositioning”, where teams retroactively disposition candidate pools rather than recording the exact stage, step, and reason shortly after decisions are made.
ANOTHER ASPECT OF DATA RISK TO CONSIDER | ‘A 4-Step Approach to Performing an AI Bias Audit’
Keep Learning About the Poor Data Problem
This article is an abridged extract from our full ebook, “The Poor Data Problem: How to Prevent and Correct Incomplete, Inaccurate, and Inconsistent Data”. In the complete guide you’ll additionally find a comprehensive look at:
- The risks associated with poor data
- Why data errors happen
- Where to look for errors and start your remedial efforts
- The best practices for data review
Ensure your data avoids liability—and take your data collection practices beyond simple compliance. Download the full ebook, then contact us to find out how we can help.
About the Author
Aly Ferguson is a Senior Business Consultant for Affirmity and has been with the organization for over ten years. She consults with clients in a variety of industries concerning workforce compliance and non-discrimination best practices along with inclusion planning, implementation, and measurement.