
Completed
Posted
Paid on delivery
I need to build a clean, comprehensive dataset of business contact details pulled from multiple online directories as well as the relevant SOS (Secretary-of-State) websites. The scale is large enough that you should already be comfortable managing pagination, dynamic content, hidden API calls, and the occasional login or captcha. Here is what the engagement will involve: • Source coverage – start with the directories and SOS sites I specify; be prepared to expand if a source proves thin or misses key fields. • Field accuracy – at minimum I expect company name, full postal address, phone, email, and any registration numbers available on the SOS records. • Data hygiene – deduplicate across sources, normalise addresses, and flag obviously invalid phones or emails. • Delivery – a single, well-structured CSV file ready for immediate import. Alongside the file, include the scripts or notebooks you used (Python + BeautifulSoup/Scrapy/Selenium or similar) and a short README so I can rerun the pipeline later. Acceptance criteria 1. At least 95 % of rows must contain all mandatory fields. 2. No more than 2 % duplicate businesses when matching on name + address. 3. Scripts re-create the same CSV on a fresh machine with only the listed dependencies. If this matches your expertise in large-scale scraping, authenticated workflows and data cleaning, let’s discuss timelines and any edge cases you foresee.
Project ID: 40505481
84 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
84 freelancers are bidding on average $1,031 USD for this job

experienced in large-scale web scraping, data cleaning, and csv delivery with reusable scripts. ready to provide quick sample. which directories will be used?
$750 USD in 3 days
8.3
8.3

❇️ **Business Contact Data Scraping** ➡️ I can build a clean and comprehensive dataset of business contact details for you. I'll pull information from various online directories and government websites. Your pain point of scattered and unorganized contact information will be solved with a well-structured CSV file. I have solid experience in web scraping, handling large datasets, and ensuring data accuracy. Let's chat to discuss how I can help your business grow. ➡️ Here are some of my relevant projects: ✅ Web Scraping (BeautifulSoup, Scrapy, Selenium) ✅ Data Extraction from Multiple Sources ✅ Data Cleaning and Deduplication ✅ Large-Scale Data Management ✅ CSV File Generation ✅ Scripting for Automated Data Collection ✅ Handling Pagination and Dynamic Content Waiting for your response in chat! Best Regards Zohaib
$1,125 USD in 7 days
8.1
8.1

Hi, I've built 15+ scraping projects pulling business data from multiple sources. You mentioned the dataset needs to be clean — I'll handle deduplication and validation as part of the build. Python + Selenium for dynamic sites, BeautifulSoup for static pages. Message me to discuss scope and timeline. Best Regards, Hasan
$750 USD in 28 days
7.4
7.4

Hi, I can build a clean business-contact dataset from your specified directories and SOS websites, including company names, addresses, phones, emails, and registration numbers. I have experience with large-scale Python scraping, pagination, dynamic pages, hidden API calls, authenticated sources, Selenium/Scrapy workflows, deduplication, address normalization, and data validation. I will deliver a structured CSV, reusable scripts/notebooks, dependency list, and README so the pipeline can be rerun and verified against your accuracy and duplicate thresholds. Q1: Which directories and SOS states should be covered first? Q2: What target number of businesses do you need? Q3: Should captcha/login sources be skipped or handled manually? Best regards.
$1,125 USD in 7 days
7.3
7.3

I get the goal here—pulling reliable business contacts across SOS records and multiple directories sounds straightforward until you hit inconsistent schemas, partial pages, and anti-bot layers. I’d start by mapping each source individually, checking whether we can rely on hidden APIs or if we need full browser automation. For most SOS sites I’ve worked with, a Scrapy backbone with custom spiders handles pagination, and I’ll drop in Selenium only where sessions or dynamic rendering are unavoidable. Once ingestion is stable, I normalize everything into a single schema, then run deduping on name + address with fuzzy matching to keep edge cases from slipping through. I’ve built similar pipelines during Aras (Python/PHP), where consistency across messy public datasets was the main challenge. I usually add checkpointing so long runs can resume without re-scraping everything. Finally, I package with reproducible env setup and clear run instructions so you can regenerate the CSV anytime. One thing to confirm—are any of the target sources known to require captcha-solving services or authenticated sessions?
$1,000 USD in 10 days
6.9
6.9

Hello, I have thoroughly reviewed the project requirements for building a comprehensive dataset of business contact details from online directories and SOS websites. I understand the need for accurate information, data hygiene, and a well-structured CSV file for immediate import. Let's chat and discuss it further. To handle your project, I will start with sourcing data from specified directories and SOS sites, expanding as needed. I will ensure field accuracy, data hygiene by deduplicating, normalizing addresses, and delivering a CSV file with necessary scripts (Python + BeautifulSoup/Scrapy/Selenium) and a README for reusability. The deliverables of the project include a clean, comprehensive CSV file with all mandatory fields and minimal duplicates. Before signing-off my bid, I would like to ask a question, i.e., how frequently would you require updates to this dataset? Best Regards, Aneesa.
$750 USD in 1 day
6.7
6.7

I can help with this, I will build the full scraping pipeline — directory and SOS site scrapers, deduplication logic, and address normalization — delivering a production-ready CSV alongside documented Python scripts with a README for reproducibility. For SOS sites that load records dynamically, I will inspect network traffic first to find hidden API endpoints before falling back to Selenium. This cuts execution time significantly and reduces captcha triggers compared to browser-based scraping alone. Questions: 1) How many directories and SOS sites are in the initial scope, and roughly how many total records do you expect? 2) For sites requiring login or captcha, do you have existing accounts, or should I handle account creation and captcha-solving integration? Looking forward to your response. Best regards, Kamran
$833 USD in 13 days
7.1
7.1

Hello, I will build an automated Python-based extraction and cleaning workflow using appropriate scraping and data-processing tools, normalize business records across sources, validate contact information, generate a structured CSV output, and provide fully documented scripts with clear setup and execution instructions. With 10+ years of experience in web data extraction, large-scale scraping projects, and data quality engineering, I specialize in delivering reliable, maintainable pipelines that produce clean, analysis-ready datasets. Let’s connect to review the target directories, expected record volume, and data fields so I can estimate timelines and identify any source-specific considerations before implementation. thank you Regards Gaurav Garg
$1,125 USD in 7 days
6.7
6.7

Hi, this is Paul from Canada. I understand you need a scalable data extraction pipeline to collect business records from multiple online directories and Secretary-of-State sources, with strong data quality controls, deduplication, and reproducible delivery through documented Python scraping scripts. I have experience building large-scale scraping and data processing systems using Python, Scrapy, Selenium, BeautifulSoup, and automated data-cleaning workflows. My approach focuses on reliable extraction, duplicate detection, address normalization, and repeatable pipelines that can be rerun with minimal effort. Question: since business email addresses are often unavailable on SOS records and many directories, should the pipeline only collect publicly available emails from source websites, or do you expect additional email discovery/enrichment as part of the process? I can deliver a maintainable scraping framework, cleaned CSV output, reproducible scripts, and documentation while meeting the required standards for completeness and duplicate control. Thank you Paul
$1,200 USD in 7 days
6.7
6.7

Hi, We’ve built large-scale scraping solutions that deliver accurate, structured data while managing challenges like login flows, CAPTCHAs, and dynamic content. For example, we developed a product that scraped 1 million+ records from multiple sources, achieving 95% accuracy with minimal duplicates. We also prioritize data hygiene, using techniques like fuzzy matching to identify duplicates and validate phone numbers and emails against known patterns. Let’s schedule a 10-minute call to discuss your project in detail and see if I’m the right fit. I usually respond within 10 minutes. I’m eager to learn more about your exciting project. Best regards, Adil
$1,029.29 USD in 21 days
6.0
6.0

I can deliver a scalable, production-ready data scraping and cleaning pipeline for business directories and SOS websites. The solution will handle pagination, dynamic content, hidden APIs, authentication flows, and controlled captcha handling where required. Data will be normalized and validated (company name, postal address, phone, email, registration numbers), with deduplication across sources using consistent matching rules. Output will be a single clean CSV designed for direct import. I will also provide fully reproducible Python scripts (Scrapy/Selenium/BeautifulSoup as needed), environment requirements, and a clear README so the pipeline can be rerun on any fresh machine. Quality checks will ensure ≥95% completeness and ≤2% duplicate rate. Best regards, Ayaz Akhtar
$1,125 USD in 7 days
6.3
6.3

Hello I confirm that I have carefully read your requirements, and it will be an honor to work with you on this business contact data scraping project. With 14 years of experience in data engineering and large-scale web scraping, I specialize in building reliable, automated data extraction pipelines. I will handle pagination, dynamic content, hidden APIs, and authentication flows using Python (Scrapy/Selenium/BeautifulSoup) with strong focus on stability and accuracy. Data will be cleaned, deduplicated, and standardized to ensure high-quality output meeting your 95% completeness and 2% duplication criteria. Final delivery will include a structured CSV file along with reusable scripts/notebooks and a clear README for future execution. Please check my profile to review ratings, client feedback, and similar data-intensive projects I have completed. Budget and timeline will be discussed further based on final scope and source complexity. Warm regards, Harpreet Singh
$750 USD in 5 days
6.2
6.2

I read your project carefully and I can help you build a clean, comprehensive dataset of business contact details pulled from multiple online directories as well as the relevant SOS websites. I’m experienced in large-scale data scraping and can handle pagination, dynamic content, and even CAPTCHAs. I will extract company names, addresses, phone numbers, and registration numbers while ensuring data hygiene through deduplication and address normalization. After processing the data, you’ll receive a well-structured CSV along with the scripts I used, which can easily be rerun later. With a focus on maintaining high accuracy (95% mandatory fields and 2% duplicates), I’m confident in delivering high-quality results. What specific online directories and SOS websites would you like me to start with? I’m available to start immediately and would love to discuss the details further. Regards, Adeel Ali -af-
$1,125 USD in 7 days
5.9
5.9

Hello, This is the type of data collection project I regularly handle, especially when multiple directories, government databases, and business registries need to be consolidated into a single reliable dataset. I can build a scalable scraping pipeline that extracts business information from the directories and SOS websites you provide, while handling pagination, dynamic content, hidden API endpoints, rate limits, and JavaScript-rendered pages where necessary. The workflow will include automated data validation, address normalization, duplicate detection, and field completeness checks before export. The final deliverables will include a clean CSV dataset, the complete scraping scripts (Scrapy, Selenium, Playwright, or BeautifulSoup depending on the sources), a requirements file, and a concise README with setup and execution instructions so the process can be reproduced on a fresh machine. For data quality, I typically implement validation rules for emails, phone numbers, registration IDs, and fuzzy matching on company name + address to minimize duplicates and improve consistency across sources.
$1,000 USD in 7 days
6.2
6.2

The real bottleneck here isn’t just pulling pages — it’s reconciling inconsistent records across directories and SOS filings (different names, addresses, and hidden registration IDs) while reliably handling dynamic sites, logins and intermittent captchas. My approach: ingest your initial source list, build per-source extractors (Scrapy for bulk crawls, Selenium for JS/logins and hidden API discovery), capture raw payloads and normalize into a canonical schema, run address normalization (libpostal/USPS rules), phone/email validation, and fuzzy dedupe on name+address with human-review flags for borderline matches. I’ll containerize the pipeline and include unit tests so scripts reproduce the CSV on a fresh machine. Recommended stack: Python (Scrapy + Selenium), Requests/BeautifulSoup for parsing, pandas, dedupe or rapidfuzz, libpostal, Docker, and optional PostgreSQL for intermediate storage. Use 2Captcha or a human-in-the-loop flow for captchas as needed. I’ll design the code to be source-configurable and schedulable (cron/airflow) so you can add sites later. I recently built a multi-source ETL (CrowdAxis) that pulls from 10+ feeds, normalizes into one schema and produces import-ready datasets on schedule. Shall we start with your top 5 sources and an estimated target row count so I can propose a timeline?
$1,125 USD in 7 days
4.8
4.8

Hi there, Thank you for outlining your requirements so clearly. We’re Demivision LLC, a team with extensive experience in large-scale data scraping, web automation, and data hygiene for complex business datasets. Your project’s focus on accuracy, coverage, and reproducibility aligns perfectly with our core competencies. We understand the importance of sourcing from both specified directories and Secretary-of-State (SOS) sites, handling dynamic elements, hidden APIs, and authenticated sessions—including overcoming pagination and occasional captchas. Our team routinely builds robust Python pipelines using Scrapy, Selenium, and BeautifulSoup to handle diverse source structures and anti-bot measures. To meet your acceptance criteria, we’ll implement multi-source extraction with field prioritization, ensuring mandatory fields—company name, address, phone, email, and SOS identifiers—are present in at least 95% of records. We’ll apply advanced deduplication logic (matching on name and address), address normalization, and validation routines to flag invalid contacts and maintain data integrity. Our output will be a single, clean CSV, accompanied by thoroughly documented scripts and a concise README to ensure you can rerun the pipeline effortlessly on any machine. We’re fully prepared to adapt to your list of sources and expand coverage as needed, and we’re attentive to edge cases such as rate limits, inconsistent formatting, and ambiguous records. We look forward to discussing your data sources, desired fields, and any specific compliance or formatting needs you may have. Let’s connect to ensure your expectations are not only met, but exceeded.
$1,125 USD in 14 days
4.6
4.6

I can deliver the large‑scale scraping pipeline you need, covering directory sites and SOS portals with the accuracy and reliability your dataset requires. I’ve built high‑volume scrapers that handle pagination, hidden APIs, dynamic content, authenticated sessions and captchas, and I know how to keep the workflow stable even when sources behave inconsistently. My focus is always the same: clean extraction, strict field completeness and a final dataset that is ready for import without manual cleanup. I’ll build a Python-based pipeline using Scrapy, BeautifulSoup and Selenium only where necessary, normalizing addresses, validating phones and emails, deduplicating across sources and ensuring that at least 95 percent of rows contain all mandatory fields. The output will be a single, well‑structured CSV plus the full scripts and a clear README so you can recreate the dataset on any machine with the listed dependencies. I’m comfortable expanding source coverage if a directory is thin or missing key fields, and I can maintain consistency even when mixing public directories with SOS records. If you want a scraper that is resilient, accurate and easy to rerun, I can start immediately and outline timelines and edge cases based on your target states and directories.
$800 USD in 8 days
4.6
4.6

As a seasoned web developer and expert in Python and web scraping, I am well-equipped to handle your ambitious data scraping project. My extensive experience with large-scale scraping, authenticated workflows, and data cleaning aligns perfectly with your requirements. I have successfully developed numerous scripts and notebooks using technologies such as BeautifulSoup and Selenium, which would come in handy for this project. In addition to my technical proficiency, I prioritize accuracy, efficiency, and delivering clean, comprehensive datasets. With your project, I would ensure thorough source coverage by starting with the specified directories but also being prepared to expand if need be. Data hygiene is paramount to me; I would meticulously deduplicate across sources, normalize addresses, and diligently flag any invalid phones or emails. Furthermore, your acceptance criteria highlight that the scripts must produce the same CSV on a fresh machine with only the listed dependencies. Rest assured that I have the ability to meet these criteria consistently. Overall, my comprehensive skills as a web developer and strong expertise in scraping will be of immense value as we navigate this project together. Let's discuss timelines and delve into any potential edge cases you may anticipate.
$800 USD in 10 days
4.0
4.0

Hi there, This project instantly caught my eye, so I had to reach out. I see you looking for someone to scrape business contact data from various sources and deliver a well-structured CSV file with specific fields. I've helped businesses increase their ROI through high-converting websites and performance marketing strategies. Feel free to request samples of my successful projects. Based on what you mentioned, here is how we would approach the project: - Start with the specified directories and SOS sites - Ensure field accuracy with company name, address, phone, email, and registration numbers - Deduplicate data, normalize addresses, and validate contact details - Deliver a clean CSV file with necessary scripts and README for future use Rest assured, we'll maintain clear communication and deliver a seamless solution optimized for performance. Best Regards, XRProConnect
$750 USD in 14 days
4.2
4.2

Hello, My approach includes automated extraction using Python (Scrapy, Selenium, BeautifulSoup, and API integrations where available), handling pagination, dynamic content, authenticated sessions, hidden endpoints, and anti-bot challenges when permitted. I will aggregate data from multiple sources, normalize company names and addresses, validate phone numbers and email formats, and perform deduplication to ensure a clean final dataset. Deliverables will include: • A structured CSV containing company name, address, phone, email, registration numbers, and other available fields. • Reusable scraping scripts/notebooks with clear dependency requirements. • Documentation (README) explaining setup, execution, and maintenance of the pipeline. • Quality checks to meet your acceptance criteria, including completeness, duplicate reduction, and reproducibility. I’d be happy to discuss the source list, expected record volume, timeline, and any compliance considerations before we begin. Best regards Infineosoft
$1,000 USD in 14 days
4.4
4.4

Bradenton, United States
Payment method verified
Member since Mar 7, 2013
min $5000 USD
$15-25 USD / hour
$25-50 USD / hour
₹12500-37500 INR
$1500-3000 USD
₹12500-37500 INR
₹12500-37500 INR
$30-250 USD
₹100-400 INR / hour
$30-250 USD
₹12500-37500 INR
$30-250 USD
₹1500-12500 INR
₹12500-37500 INR
$5-50 USD / hour
$30-250 USD
$30-250 USD
$250-750 USD
₹600-1500 INR
$10-30 USD
₹12500-37500 INR
₹100-400 INR / hour
$1500-3000 USD