Build and own custom web crawling tools, services, and workflows to improve scrape analysis and reporting. Develop scripts to automate data processing, transforming complex raw data into structured formats. Manage data ingestion pipelines and ensure quality through AI-driven processes. Design and implement anomaly detection models for geospatial and domain-specific requirements using libraries like Prophet. Communicate specific data requirements for web scraping with third-party vendors.
Responsibilities
- Build and maintain custom web crawling tools and services.
- Automate data processing to transform raw data into structured formats.
- Manage end-to-end data ingestion pipelines and ensure data quality.
- Design anomaly detection models for geospatial requirements.
- Coordinate data requirements with third-party vendors.
Required Skills
- 5+ years of experience with Python, PySpark, SQL, Scala, and Shell scripting.
- Proven experience running large-scale web scrapes and analyzing scraping requirements.
- Strong proficiency in Python, SQL, C#, and Git.
- Experience working as a Data Engineer in a production environment.
- Hands-on experience with SQL and cloud platforms including Azure, AWS, or GCP.
- Deep understanding of data modeling concepts for efficient storage and retrieval systems.
- Background in building anomaly detection models and understanding machine learning algorithms.
- Experience with Spark and large-scale data processing.
Preferred Skills
- Experience with MySQL and distributed notebook environments like Databricks or Azure Synapse.