Stijn AI
Data

Data Cleaner

Normalises messy records, finds duplicates and explains every change.

Data Cleaner takes a batch of records and returns them standardised: names cased properly, addresses normalised, phone numbers in E.164, companies deduplicated across spelling variants, and a change log explaining every edit it made.

Every change is explained

The output includes a diff. For each field the agent changed, you get the original value, the new value and a one-line reason. Data cleaning that you cannot audit is data corruption you have not noticed yet.

Fuzzy matching that understands context

Deduplication is where naive tools fail hardest. String similarity says "Acme Corp" and "Acme Corporation" are different, and that "Acme Ltd" and "Acme Limited" are the same as each other but not as the first two. The agent reasons about them as company names, and it knows that "Apple Inc" and "Apple Records" are genuinely different organisations.

Non-destructive by default

It returns a proposed cleaned set alongside the original. Applying the changes is your step, not its. For a first run against a production database, apply the diff to a copy and read it.

What it can do

  • Standardise names, addresses and phone numbers
  • Deduplicate across spelling and legal-suffix variants
  • Infer and fill obviously missing fields
  • Flag records that look fabricated or test data
  • Produce a full auditable change log
  • Normalise country, currency and date formats
  • Detect and separate merged fields
  • Return records it declined to touch, with reasons

Inputs and outputs

It takes

  • Record batch (JSON or CSV)
  • Target schema
  • Cleaning rules (optional)

It returns

{ "cleaned[]": ... "changes[]": ... "duplicates[]": ... "flagged[]": ... "untouched[]": ... }

Fit

Good for

  • CRM hygiene
  • Pre-migration cleanup
  • Merging acquired customer lists
  • Mailing list normalisation

Not for

  • Real-time validation at form submit
  • Authoritative address verification
  • Financial reconciliation