
In Progress
Posted
I am assembling a corpus of real-world documents that contain publicly available PII across multiple domains and need a consultant who can combine smart desk-research with targeted web-scraping to locate, capture, and catalogue them for later analysis. Scope of work • Source material: real-world data sets, reports, filings, or any document that actually exposes PII (names, addresses, SSNs, medical record numbers, account details, etc.). • Focus domains: Healthcare, Finance and Education. • Geographic emphasis: North-American publications at this stage (government portals, open-data sites, public court documents, regulatory disclosures, etc.); other regions may follow once this tranche is complete. • Methods: a mix of automated scraping (Python, BeautifulSoup/Scrapy/Selenium or similar) and classic desk research to reach sources that resist automation. • Output: an organised folder structure plus a spreadsheet/JSON catalog listing document title, source URL, date accessed, domain tag, and a short note of the specific PII fields present. Acceptance criteria 1. Minimum 250 unique documents, balanced across the three domains and all the documents ya must be single page with maximum 300 words 2. Each entry must include working source links and clear evidence of at least one PII field. 3. No paywalled or illegally obtained content—everything must be freely reachable on the open web. 4. Scripts (if used) are handed over, well-commented, and runnable on a vanilla Python environment. If you have experience harvesting open data, navigating government portals, and keeping an eye on compliance while still finding those hard-to-spot files, I’d like to hear how you’d tackle this and how quickly you could deliver the first batch.
Project ID: 40520675
6 proposals
Remote project
Active 7 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

I am highly interested in your project to assemble a corpus of real-world documents containing publicly available PII across Healthcare, Finance, and Education domains. With strong expertise in data scraping, open-data research, and Python automation, I can efficiently locate, extract, and catalogue compliant datasets from North American sources. My approach combines automated scraping using Python, BeautifulSoup, Scrapy, and Selenium with manual desk research to identify government portals, public filings, and regulatory disclosures. I ensure full compliance with open-access standards, avoiding paywalled or restricted content. The deliverables will include a structured folder system and a spreadsheet/JSON catalog listing document titles, URLs, access dates, domain tags, and identified PII fields. All scripts will be clean, well-commented, and runnable in a vanilla Python environment. I will deliver a minimum of 250 unique documents balanced across the three domains, each verified with source links and clear evidence of at least one PII field. With experience in data management and information security, I guarantee accuracy, compliance, and reproducibility. The first batch can be delivered within 10–12 days, ensuring quality and efficiency. I look forward to collaborating on building a robust, compliant, and scalable PII corpus that meets your analytical goals.
₹110 INR in 24 days
0.0
0.0
6 freelancers are bidding on average ₹218 INR/hour for this job

My name is Hamza and I specialize in providing top-notch data services, blending my expertise in Data Science, Statistical Analysis, and Web Scraping using Python to deliver actionable insights. My broad knowledge of Python libraries, automation techniques, as well as data cleaning and preprocessing makes me an ideal candidate for this task. With over three years' experience in leveraging data analytics, I am comfortable navigating various web sources -- including government portals -- while still ensuring compliance with regulations. I have previously harvested open data from multiple online resources while maintaining a strict adherence to ethical retrieval and usage.
₹250 INR in 40 days
3.1
3.1

We can deliver this engagement through a structured, compliance-conscious data acquisition and document intelligence workflow focused on accuracy, traceability. Our team has experience building web-scraping systems, large-scale data processing pipelines, search and indexing engines, and AI-powered information extraction solutions. For this project, we will combine targeted desk research with automated collection pipelines to identify, validate, and catalogue publicly accessible documents containing real-world PII across Healthcare, Finance, and Education domains. Our methodology includes: • Discovery of legitimate public sources including government repositories, regulatory disclosures, court records, educational archives, and open-data portals • Automated collection using Python, Scrapy, BeautifulSoup, Selenium. • Classification by domain, source type, document category, and exposed PII fields • Delivery of reusable, documented collection scripts for future expansion Deliverables: • 250+ unique documents meeting all stated requirements • Organized folder hierarchy with consistent naming conventions • Spreadsheet and JSON catalog containing source URL, access date, domain tags, document metadata, and identified PII indicators • Collection methodology and validation report The expected timeline for this project could vary from 30 hrs to 50hrs due to authenticity. AlphaFusion Corporation AI • Data Engineering • Automation • Enterprise Intelligence Solutions
₹200 INR in 80 days
0.0
0.0

Hello, I have strong experience in web scraping, OSINT, and large-scale data collection using Python (BeautifulSoup, Scrapy, Selenium, and custom automation tools). For this project, I can combine automated scraping with manual research to identify and catalog publicly accessible documents across Healthcare, Finance, and Education domains. I will deliver: • 250+ unique documents meeting your criteria • Organized folder structure • Excel/CSV/JSON catalog with source URL, access date, domain tag, and identified PII fields • Clean, well-documented Python scripts for reproducibility • Link validation, deduplication, and quality checks I can provide an initial sample batch quickly for review before scaling to the full dataset. Looking forward to discussing the timeline and requirements. Best regards
₹100 INR in 40 days
0.0
0.0

Assembling a corpus of real-world documents containing publicly available PII is no small task, especially when it requires both smart desk research and targeted web scraping. Your need for a reliable consultant who can navigate diverse domains like Healthcare, Finance, and Education is clear. With over 12 years of experience in data sourcing and web scraping using Python libraries such as BeautifulSoup, Scrapy, and Selenium, I can efficiently locate the required documents while ensuring compliance with legal standards. My background also includes working with government portals and public databases to harvest open data responsibly. I will provide a structured output including source links and well-documented scripts that are easily runnable in a vanilla Python environment. To ensure we meet your expectations, I would love to clarify: how do you envision balancing the document types across the three specified domains? Looking forward to collaborating on this important project!
₹400 INR in 7 days
0.0
0.0

Hi Client's, My name is Anggita Ardiyani, and I’m interested in supporting your document research and data collection project. With 8 years of experience in documentation management, data organization, and administrative support, I have strong attention to detail and experience working with structured datasets, spreadsheets, and large volumes of information. I have experience conducting web research, organizing databases, verifying information accuracy, and maintaining well-structured records. Recently, I organized a database of 100+ professional contacts with 95% data accuracy and improved search efficiency through consistent formatting and verification. For this project, I would follow a structured process to identify publicly available documents, verify source accessibility, categorize findings by domain, and maintain a detailed catalog including URLs, access dates, and relevant document notes. I am proficient with Excel, Google Sheets, Google Drive, and data organization workflows, and I am comfortable learning additional tools required for the project. I would be happy to discuss your timeline and the expected delivery schedule for the first batch. Best regards, Anggita Ardiyani P.S. Just curious — do you already have preferred source portals for Healthcare, Finance, and Education, or would you like me to identify and prioritize them as part of the research process?
₹250 INR in 40 days
0.0
0.0

Bangalore, India
Payment method verified
Member since Mar 29, 2015
₹12500-37500 INR
$2-8 USD / hour
₹37500-75000 INR
$10-30 CAD
$250-750 USD
min $50 USD / hour
₹1500-12500 INR
$200-600 USD
$1500-3000 SGD
$30-250 USD
₹12500-37500 INR
₹12500-37500 INR
$30-250 USD
$250-750 USD
min £36 GBP / hour
$100-101 USD
$250-750 USD
₹1500-12500 INR
₹100-400 INR / hour
$15-25 USD / hour
$2-8 USD / hour
₹1500-12500 INR
$15-25 USD / hour