On this page
The problem data processing solves
Data rarely arrives ready to use. Customer names are spelled three ways, dates come in different formats, the same product has different codes in the store and the warehouse system, and someone’s export has a blank row halfway down. Before anyone can report on it or import it anywhere, it has to be cleaned and combined.
Doing this by hand once is tedious. Doing it every month is a real cost — and each manual pass risks slightly different results. A data processing pipeline does the same job the same way every time.
Where it helps
- A D2C brand merges orders from its website, marketplaces and courier reports to see true delivered revenue per channel.
- A real estate agency combines lead lists from portals, ads and walk-ins, removing duplicates before they go into the CRM.
- A logistics firm reconciles its own delivery records against client invoices and carrier statements.
- A manufacturer consolidates production logs from several machines and shifts into one daily dataset.
- A SaaS startup prepares usage and billing data for analysis or migration into a new system.
What a pipeline does, step by step
- Ingest files or pull data from databases and APIs — CSV, Excel, JSON, XML or database tables.
- Parse and normalise formats: dates, currencies, phone numbers, units, text case and whitespace.
- Validate each record against rules you agree — required fields, allowed values, sensible ranges.
- Deduplicate using exact and fuzzy matching, with clear rules for which version of a record wins.
- Merge and enrich, joining sources on shared keys and mapping codes between systems.
- Output to the destination you need: a clean spreadsheet, a database table, a CRM import file or an API.
- Report what happened — records processed, merged, rejected and why.
Tools and technology
We choose tools according to data size and where the pipeline will run:
| Purpose | Tools |
|---|---|
| Tabular processing | pandas, or Polars for larger files |
| Validation | Pydantic or pandera schemas |
| Fuzzy matching | RapidFuzz |
| Databases | PostgreSQL, SQLite, or your existing database via SQLAlchemy |
| Scheduling | cron or a job scheduler on a server |
| Excel input and output | openpyxl |
What affects timeline and cost
- The number of sources and how different their formats are.
- How messy the data is, and how many business rules are needed to decide what’s correct.
- Data volume, which influences the tools and hosting needed.
- Whether the pipeline runs once (a migration) or repeatedly on a schedule.
- Where the output needs to go, and whether that destination has an import API.
Common mistakes
- Cleaning in the spreadsheet by hand, so the steps are lost and next month starts from zero.
- Deduplicating too aggressively and merging two genuinely different customers with similar names. Matching rules need review with someone who knows the data.
- Overwriting source files. Raw inputs should be kept untouched so any step can be re-run.
- No record of rejections, which hides data quality problems at the source.
One-off cleanup or recurring pipeline
It helps to be clear early about which kind of job this is, because it changes how the work is built.
- One-off cleanup — typically before moving to a new CRM, ERP or accounting system. The focus is getting the historical data right once, with a clear report of what was changed, merged or left out.
- Recurring pipeline — daily, weekly or monthly data that arrives in the same shape each time. The focus is reliability: it runs on a schedule, handles the usual exceptions on its own and alerts someone when something genuinely new turns up.
A one-off job often becomes recurring once people see the result. We write even one-off scripts cleanly enough that turning them into a scheduled pipeline later is a small step, not a rewrite.
Frequently asked questions
Can the cleaned data go straight into our CRM or database?
Yes. Output can be written directly to a database, pushed through a CRM’s import API, or produced as an import file in exactly the format the destination system expects.
How large a dataset can you handle?
From a few hundred rows to many millions. For larger volumes we use tools such as Polars or a database for processing, rather than loading everything into a spreadsheet-style workflow.
Can you work with our confidential data?
Yes. We can work on sample or masked data, run the pipeline on your own server, and agree confidentiality terms before any data is shared.
Is this a one-off job or ongoing?
Either. A one-off cleanup is common before a system migration; a scheduled pipeline suits recurring monthly or daily data.
What do you need from us to start?
Representative sample files, including awkward ones, and a description of what the final output should look like and who uses it.
Talk to us about data processing
Cleaning, transforming and merging large datasets into something you can actually use.