
In Progress
Posted
Paid on delivery
I'm in need of a robust Python-based web crawler that can efficiently scrape and aggregate data from multiple news sites. The crawler will need to extract various types of data including article content, author information, publication dates, tags, and cover thumbnails. Key Features: - Proxy Support: The crawler should be equipped with proxy support to bypass geographic and IP restrictions, ensuring uninterrupted access to these news sites. - Configurable: It should be configurable via environment variables, dynamically adjusting proxy usage and frequency to comply with site terms and avoid bans. - Resume Feature: The crawler should include a resume feature, allowing it to track and stop at the last crawled position, making subsequent crawls faster and avoiding redundant data collection. - Data Storage: The extracted data needs to be saved in JSON format to MongoDB Ideal Skills: - Extensive experience with Python and web scraping - Familiarity with JSON and data aggregation - Understanding of proxy usage and its configuration - Ability to implement a resume feature in a web crawler.
Project ID: 38733237
110 proposals
Remote project
Active 2 yrs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs