Data cleaning and deduplication:
one record per customer, not five
A customer list, a product catalog or a CRM export almost always has the same entity spelled three different ways, phone numbers in four formats and rows nobody dares delete. An agent finds the duplicates, standardises the formatting, and brings the genuinely ambiguous merges to a person instead of guessing.
The process today
A dataset that has been fed by more than one source for more than a year almost never stays clean on its own. A customer signs up once through a web form and once through a sales call logged by a different manager, and now the same person exists as “J. Smith,” “John Smith” and “john.smith+promo@gmail.com” in three separate rows, each with slightly different order history attached. A product catalog imported from two suppliers has the same SKU under two different internal codes with different unit labels. None of this is anyone’s fault exactly; it is what happens when data arrives from different places over time without anyone’s job being to reconcile it.
The cost is not abstract. Marketing sends the same promotion to the same person three times under three record variants, which looks sloppy and burns send quota. A support agent pulls up one version of a customer’s record and misses the order history sitting under the duplicate, and gives an answer that contradicts what the customer was told last time. A finance report counts revenue against a product twice because two SKUs that are really one product were never merged, and the discrepancy takes an afternoon to track down and explain.
Manual deduplication does not scale past a few hundred rows before people start making the same judgment calls over and over, inconsistently, because fatigue sets in and the twentieth “is this the same person” decision gets less care than the first. And nobody wants to run a bulk merge by hand on live customer or financial data, because one wrong merge is a genuinely bad day, so the cleanup gets postponed until the dataset is too painful to work with, at which point the project has grown to a size that makes the postponement worse.
What the agent does
The agent starts with an audit of the dataset as it actually exists: duplicate rate by match type, the specific formatting inconsistencies present (date formats, phone number styles, inconsistent units or currencies), and which fields are missing often enough to matter. This runs on a copy of your data, never the live system, so the audit itself carries zero risk.
Matching runs on fuzzy logic across names, emails, phone numbers and addresses rather than exact string matching, which catches the “John Smith” versus “J. Smith” cases that simple deduplication tools miss. Every match comes with a confidence score. Above a threshold you set, matches merge automatically, keeping the most complete and most recent values from each duplicate and logging exactly what was merged from where. Below that threshold, the match goes to a review queue with both candidate records shown side by side, so a person makes the genuinely ambiguous call instead of the agent guessing on a coin flip.
Formatting gets standardised across the board at the same time: dates into one format, phone numbers into E.164 or whatever your system expects, currency and units normalised so a report does not silently mix pounds and kilograms. The whole pass, including every automatic merge and every queued decision, is written to an audit log that lets you see and reverse any single change, not just restore from a full backup.
For data that keeps accumulating duplicates from an ongoing source, like a form that is not enforcing uniqueness, the same matching logic runs on a schedule, catching new duplicates as they arrive instead of letting the dataset drift dirty again a month after the first cleanup.
What stays with humans
Every merge below the confidence threshold, any decision to delete rather than merge a record, and the final choice of which source record’s values win when two duplicates disagree on something important (a different address, a different price) stay with a person who knows the business context the agent does not have.
Guards
Every automatic merge is logged with both original records attached and can be reversed individually, not just restored from a full backup. The first pass on any new dataset runs as a dry run against a copy, and nothing touches the live system until that dry run has been reviewed and the confidence threshold has been tuned against real examples from your own data. Access during the project is scoped to the fields needed for matching, not the full dataset.
Price and timeline
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| Single automation | from $500 | One dataset cleaned and deduplicated, up to ~50,000 rows, full audit log | 3 to 10 days |
| Department package | from $2,500 | Cleanup plus spreadsheet workflow automation and a recurring monthly clean for the team that owns the data | 2 to 4 weeks |
Running cost for a recurring monthly clean is usually $10 to $50 in model usage depending on dataset size, with a budget cap set before launch.
Related
This pairs directly with spreadsheet workflows automation once the underlying data is clean enough to automate on top of, and with weekly reports in plain language and dashboard commentary, both of which are only as trustworthy as the data feeding them. See the automation-everything overview and the AI agents service page for the broader catalogue. We did a version of this at real scale in the self-hosted ERP recovery for a supplements factory, exporting and rebuilding 51 tables and 12,039 rows, and the two-brand analytics warehouse depends on the same kind of clean, joined data underneath its dashboards.
If a dataset has reached the point where nobody fully trusts a report built on it, get in touch and we will send back an audit of what is actually wrong with it before touching anything.
Tired of doing this by hand? We can take the whole routine off your team, not just this step: Routine takeover, from $400 →
FAQ
How much does data cleaning and deduplication cost?
From $500 for a one-time cleanup of a dataset up to around 50,000 rows, delivered with a full audit log. Larger datasets, multiple source systems, or a recurring monthly clean are $1,500 and up.
How long does a cleanup take?
3 to 10 days depending on dataset size and how many fields need standardising. Most of that time goes into setting the matching rules correctly for your specific data, not running the match itself.
Which tools and formats does it connect to?
CSV and Excel exports, direct database connections (PostgreSQL, MySQL), and API pulls from CRMs like HubSpot, amoCRM or Salesforce. Output goes back into the same system or to a clean file, whichever you need.
What if the AI merges two records that should stay separate?
Every merge below a confidence threshold you set goes to a review queue instead of happening automatically, and every merge that does happen automatically is logged with both original records attached, so it can be reversed. We never delete the pre-cleanup data.
Is our customer data safe during the process?
The first pass always runs on a copy, never your live system, and access is limited to the fields needed for matching. Your data stays in your own database or spreadsheet; we do not retain a copy after the project closes unless you ask us to keep maintaining it.