On this page
When a custom scraper is needed
A custom scraper is a crawler designed around one specific source and its quirks. It is the right choice when a site renders content only in the browser, hides data behind filters and multi-step navigation, loads results as you scroll, or when the data sits inside a portal that you are entitled to access — your own supplier portal, marketplace seller dashboard or internal system without an export feature.
The key word is entitled. Where a login is involved, we work only with your own authorised account, on data you have the right to retrieve, and within the platform’s terms. We do not bypass logins, paywalls or CAPTCHAs, share or buy credentials, or access accounts that are not yours.
Typical situations
- An e-commerce seller that needs its own order or payout reports from a seller portal that offers no suitable export or API.
- A logistics firm retrieving its own shipment statuses from a carrier portal on a schedule, where the carrier permits it.
- A distributor pulling its own price and stock sheets from a supplier’s dealer portal into its inventory system.
- A research team collecting public data from a site built as a single-page application, with filters and infinite scroll.
- A clinic or diagnostic centre exporting its own records from an old web-based system being replaced, as part of a migration.
Where the platform offers an official API or export, we will recommend that first. A custom scraper is the fallback when no reasonable official route exists.
What we build
- A crawler designed for the source, handling rendering, navigation, filters, scrolling and pagination.
- Secure credential handling for your own accounts — stored encrypted, never in code, and revocable by you at any time.
- Loud failure — when something breaks, the scraper stops and alerts rather than returning empty results that look like real data.
- Clean recovery — runs can resume from where they stopped without duplicating records.
- Validation on every run, comparing record counts and field formats with expected ranges.
- Deployment on your server or ours, on a schedule, with logs and documentation.
How we build it
- Suitability check. We confirm the data is public or yours to retrieve, review the terms and robots directives, and look for an official API first.
- Source analysis. We study how the site loads data — rendered HTML, background data requests, or both — and pick the most stable approach.
- Prototype. A small working version proves the approach on real pages.
- Hardening. Retries, timeouts, rate limiting, resumable runs and validation are added.
- Deploy and monitor. The scraper runs on schedule, with alerts on failure or suspicious output.
- Maintain. When the source changes, we update the crawler under a maintenance arrangement.
Tools and technology
Custom scrapers are built in Python. Playwright drives a real browser for JavaScript-heavy pages and multi-step flows. Where a page loads its data from structured background requests, reading those directly is often faster and more stable than parsing the rendered page. Scrapy handles larger crawls, and pandas validates and shapes the output.
Runs are scheduled with standard job schedulers or containers, logged, and wired to alerts by email, Telegram or Slack. Every scraper is rate-limited to stay well within what a normal user of the site would generate.
What affects timeline and cost
- Source complexity — rendering, navigation depth and how often the site changes.
- Authentication — working with your own account adds secure session handling and careful testing.
- Volume and frequency — browser-based scraping is heavier to run at scale than plain requests.
- Reliability requirements — a daily business-critical feed needs more monitoring than an occasional export.
- Output destination — files versus direct integration with your systems.
We scope after a short source analysis, because the site itself decides most of the effort.
Frequently asked questions
Can you scrape data from behind a login?
Only using your own authorised account, for data you are entitled to retrieve, and within the platform’s terms. We do not bypass logins, paywalls or CAPTCHAs, or access accounts that are not yours.
Do you solve or bypass CAPTCHAs?
No. A CAPTCHA is a clear signal that a site does not want automated access at that point. We respect it and look for an official API or another legitimate route instead.
How do you keep my account credentials safe?
Credentials are stored encrypted, kept out of source code and logs, and used only by the scraper. You can revoke or change them at any time, and we recommend a dedicated account where the platform allows it.
What happens when the site changes?
The scraper detects unexpected output and alerts instead of passing bad data along. Under a maintenance arrangement we update it and resume collection.
Talk to us about custom scrapers
Purpose-built crawlers for sites with logins, pagination or heavy JavaScript rendering.