On this page
What data extraction covers
Where web scraping is about collecting pages on a schedule, data extraction is about reading specific fields accurately out of whatever source you have. That source might be a set of public web pages, a folder of PDF invoices, supplier catalogues, rate cards, government notices or exported reports.
The output is a dataset mapped to your schema — the column names, types and formats your systems already expect — rather than whatever structure the source happened to use.
Where it helps
- A logistics firm that receives rate sheets from carriers as PDFs in different layouts and needs them in one comparable table.
- An e-commerce seller that gets supplier catalogues as PDFs or spreadsheets and wants product specs loaded into its store.
- A clinic or diagnostic centre that needs fields pulled from its own standard lab or billing documents into a reporting sheet.
- An accounts team that wants invoice numbers, dates, totals and tax lines extracted from incoming vendor invoices.
- A research team collecting published figures from public reports and notices into a single dataset.
Where documents contain personal data — patient or customer details, for example — we work only on documents you are entitled to process, and handle them under the data-protection obligations that apply to you.
Deliverables
- A field map agreeing each output column, its type, its allowed values and where it comes from in the source.
- Extraction code that handles layout variations, multi-page documents and tables that break across pages.
- Validation rules that check types, ranges and totals — for example, that line items add up to the invoice total.
- An exceptions report listing every record that failed validation and why, instead of silently guessing a value.
- Clean output as CSV, Excel, JSON or direct database inserts.
- A repeatable run so next month’s batch goes through the same process without manual effort.
How we approach an extraction job
- Collect samples. You share a representative set of documents or pages, including the messy ones.
- Map the fields. We agree the target schema and how each field should be read and normalised.
- Choose the method. Text-based PDFs and HTML can be parsed directly; scanned documents need OCR, which we use only where it is reliable enough for the purpose.
- Build and test. We run extraction on the samples and compare the output against values checked by hand.
- Handle exceptions. Records that fail validation go into a review queue or report, with the reason attached.
- Automate. The process is scheduled or triggered when new files arrive, and documented for your team.
Tools we use
Extraction is written in Python. HTML pages are parsed with standard parsing libraries, and pages that render with JavaScript are read through Playwright. For PDFs we use established PDF text and table libraries, adding OCR for scanned files. pandas handles reshaping, type conversion and validation before export.
Where a document type is too varied for fixed rules, AI-assisted extraction can help, but we always pair it with validation so that a confident-looking wrong answer is caught rather than trusted.
What drives effort
- Layout variety. Ten suppliers with ten formats take more work than one consistent template.
- Scanned versus digital files. Digital text is faster and more reliable to extract than scans.
- Number of fields. Pulling five fields is simpler than pulling fifty with nested line items.
- Validation depth. Cross-checks such as totals and date logic add build time but prevent bad data.
- Volume and frequency. A one-time backlog differs from a daily inflow that must be processed automatically.
We give a firm estimate after reviewing samples, because the documents themselves decide most of the effort.
Frequently asked questions
Can you extract data from scanned PDFs?
Often, yes, using OCR. Accuracy depends on scan quality, so we test on your real samples first and route uncertain values to an exceptions report rather than guessing.
What if every document has a different layout?
That is common. We build handling for each layout family, or use AI-assisted extraction with strict validation where layouts vary too much for fixed rules.
How do I know the extracted data is correct?
Every field is validated against types, ranges and cross-checks such as totals. Records that fail are listed with reasons, and we compare output against hand-checked samples before handover.
Can this run automatically when new files arrive?
Yes. Extraction can watch a folder, an email inbox or a cloud storage location and process new files as they land, with results sent to your database or sheet.
Do you keep copies of our documents?
Only for as long as the project needs them. We agree retention and deletion up front, work on documents you are entitled to process, and can run the extraction on your own server if you prefer the files never leave your infrastructure.
Talk to us about data extraction
Pull the specific fields you need out of pages, PDFs or documents into a clean dataset.