All projectsWeb scraping · Data engineering

Real project

20+ public data sources, collected automatically.

More than 20 source-specific collectors that transformed statistical and customs websites into repeatable, multi-year ingestion workflows.

20+sources automated
The problem
Customs and statistics websites each publish data differently. Historical imports took days of manual downloads and couldn't be repeated reliably.
What I built
20+ source-specific scrapers feeding one standard schema, with checks that stop bad data instead of passing it on.
My role
Investigated each source, wrote the extraction logic and packaged runs with Docker. Junior Data Analyst.
Result
Up to 80% less processing time on customs and statistical data. Multi-year imports now repeatable.
Why it was hard
  • Four formats: HTML, Excel, CSV and PDF
  • Websites that change structure without warning
  • Large multi-year data pulls
Tools
Python · BeautifulSoup · Requests · Pandas · Docker
Laptop displaying source code in a developer workspace
Illustrative photo.

THE PROBLEM

Why the system needed to exist.

Useful public data was trapped behind inconsistent tables, downloads, page flows and formats. Manual retrieval made historical imports take days and was difficult to reproduce.

MY ROLE

Where I created leverage.

I investigated each source, designed resilient extraction logic, normalized the output and packaged the scripts for consistent local and cloud execution.

SANITIZED WORK SAMPLE

Source quirks stay outside the canonical dataset.

The pattern below captures how independent collectors feed one reliable analytical contract.

01 / SOURCE ADAPTERHTML / XLSX / CSV / PDF

Each collector handles navigation, pagination and file-specific behavior.

02 / CANONICAL CONTRACTONE SCHEMA

Names, dates, units and identifiers are normalized before delivery.

03 / QUALITY GATESCHEMA + VOLUME + FRESHNESS

Failures become explicit review states rather than silent partial exports.

Architecture-level evidence only. Selectors, access patterns, schedules and destinations are omitted.

Technical details

SYSTEM FLOW

From friction
to a repeatable flow.

  1. Source discovery
  2. Request strategy
  3. Extraction
  4. Schema detection
  5. Normalization
  6. Validation
  7. Incremental export
  8. Deployment

CONSTRAINTS

The difficult parts.

  • HTML, Excel, CSV and PDF source formats
  • Changing website structures
  • Large multi-year pulls
  • Local and cloud runtime consistency

DESIGN DECISIONS

How I approached them.

  • Kept extraction adapters separate from the canonical output schema.
  • Added schema checks and clear failures instead of silently producing partial data.
  • Used streaming or incremental output for large files to control memory use.
  • Standardized runtime behavior with Docker and development containers.

OUTCOMES

Project outcomes

20+

automated scrapers

Across customs and statistical sources.

80%

faster processing

Up to 80% less processing time in key workflows.

Repeatable

delivery

Consistent local and cloud execution.

TECHNICAL SURFACE

PythonBeautifulSoupRequestsPandasJupyterDockerSchema validation
NEXT SYSTEM / 06

Eight product catalogs merged into one online store.