Remote
Scraping Engineer
About this role
ABOUT THE ROLE You'll be an embedded member of a fast-moving AI startup team, owning the reliability and quality of a web extraction infrastructure that processes millions of pages daily through hundreds of scripts. Reporting to a Forward Deployed Engineer, this role is critical to keeping a high-volume data pipeline running smoothly and accurately. WHAT YOU'LL DO - Maintain and monitor a triage queue, resolving broken scrapers and data quality alerts to keep operations running smoothly.
- Build and deploy new web scraping scripts for websites requiring data extraction. - Validate scraped data for accuracy and investigate discrepancies. - Create dashboards to visualize and monitor scraped data in real time. - Collaborate with AI agents to fill capability gaps, fix issues, and improve existing scripts. WHAT WE'RE LOOKING FOR - 1+ years of hands-on, professional experience building or maintaining web scraping solutions.
- Proficiency with TypeScript and Node.js for building and debugging scraping scripts. - Proficiency with SQL for data querying and validation. - Experience with Puppeteer or similar browser automation libraries. - Experience debugging and fixing broken scraping scripts in production environments. - Familiarity with message queues (e.g. RabbitMQ) and asynchronous job processing systems. - Experience with Redis or similar in-memory caching systems.
- Experience with Google Cloud Platform or equivalent cloud infrastructure. - Comfort with bash scripting, git, and gRPC. - Exposure to BigQuery or similar data warehousing solutions is a plus. - Experience building dashboards or data visualization tools is a plus. - Strong async communicator who flags blockers proactively and thrives with ambiguity in a fast-paced environment. - Availability for 8+ hours per day with overlap during US business hours.
LOCATION Fully remote — all time zones welcome, provided you can maintain overlap with US business hours.
Source listing: ashby_clera