
In Progress
Posted
Paid on delivery
I have roughly 5,000 DEF 14A proxy statements in HTML format and I need the key compensation details for each named executive pulled out and placed into a clean, structured file. The fields I must end up with are: base salary, stock options and awards, bonuses / incentive pay, plus any other compensation figures that appear in the summary or grants tables. Deliverables • A single CSV or Excel file where each row is a firm-year filing and each column holds one of the compensation items above, clearly labeled. • A short read-me describing the extraction logic, any LLM prompts used, and the quality-control steps you applied. • A reproducible script or notebook so I can rerun the pipeline on future filings. Acceptance criteria • ≥ 95 % of filings processed; missing cases flagged with reasons. • Random audit of 50 filings must show ≤ 5% field-level error rate. • Output passes numeric sanity checks (e.g., no negative salaries, totals match table footings when provided). If you have experience parsing SEC filings or have already built hybrid scraping/LLM solutions, that will help you move quickly. Let me know how you plan to split automation versus manual review, which tools or models you prefer, and your estimated turnaround time for the full 5,000-file set.
Project ID: 40578966
153 proposals
Remote project
Active 5 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

I’d be happy to help with this beneficial ownership extraction project. I have direct experience working with DEF 14A proxy statements, SEC filing HTML, and hybrid rule-based/LLM extraction pipelines. My approach would be to first detect the relevant beneficial ownership table in each filing using table titles, headers, ownership keywords, and filing context. Then I would extract the row-level data into a clean CSV/Excel format, including fields such as owner name, role where available, shares beneficially owned, ownership percentage, options exercisable within 60 days, indirect ownership, RSUs/other securities, excluded or disclaimed shares, footnote references, and filing metadata. Because these tables often depend heavily on footnotes, I would use a hybrid workflow: deterministic parsing for standard tables and LLM-assisted extraction for complex footnotes or irregular disclosures. I would also include QA flags for table-not-found cases, low-confidence rows, missing fields, unusual values, unresolved footnotes, and filings needing review. Deliverables would include the final structured CSV/Excel file, a filing-level QA report, a short README explaining the extraction logic and quality checks, and a reproducible script or notebook so the pipeline can be rerun later. Best, Sep
€1,000 EUR in 10 days
5.7
5.7
153 freelancers are bidding on average €1,011 EUR for this job

Hello, I have experience hybrid scraping/LLM solutions. The first ther eis some inconsistencies in the job description. First you note "need the key compensation details for each named executive", the you note "A single CSV or Excel file where each row is a firm-year filing and each column holds one of the compensation items above". he natural output grain is executive-year (CIK, filing date, exec name, title, fiscal year, then the comp columns). You can always aggregate up to firm-year, but not back down. I suggest you want exec-level rows — most researchers do, since that's the Execucomp-style structure. I'd use deterministic parsing (expect ~80% of filings) using Python. Post-2006 filings are quite standardized. Pre-2006 filings use the older format. But both can be parsed. If deterministic parsing will produce non-reliable output, the I'd use LLM fallback (~10–20% or less). I prefer work with Claude (I have pro plan subscription). Finally manual review and filling if both above approaches fail (I expect small number of such cases). For quality check, I'd use: Where a Total column exists: sum of components must equal Total within rounding tolerance + duplicate detection on (CIK, exec, fiscal year)+fiscal year must fall within a same range relative to filing date + possible some other checks. I expect that it can take 6-8 working days to receive final dataset. I could start tomorrow. Please let me know if you are interested in my proposal. Regards, Alex.
€750 EUR in 7 days
7.7
7.7

HI I am charted accountant plus python programmer, its pretty good skill to work on these kind of tasks. yes I have extracted data from many statements in past for eg here is link to one of similar project I have completed in past: https://www.freelancer.com/projects/data-entry/Fortune-Proxies-Compilation yes I will write a scprit in python and will train it on your data, i will verify data by manaul review and then handover to you, you can share me small batch of html files I will send you sample upfront then we can move forward? and yes I will provide you reusable scprit. thank you
€800 EUR in 2 days
7.9
7.9

Hi there, I have thoroughly reviewed your project requirements for extracting key compensation details from 5,000 DEF 14A proxy statements in HTML format. Let's chat and discuss it further. To handle your project, I will start with developing a custom Python script using BeautifulSoup for HTML parsing and Pandas for structuring the data into a clean CSV or Excel file. My approach involves implementing a combination of automated extraction techniques and manual review to ensure accuracy. Additionally, I will provide a detailed readme file outlining the extraction logic, LLM prompts used, and quality control measures. The deliverables will include a structured CSV or Excel file with compensation details, a readme file, and a reproducible script for future use. Before signing-off my bid, I would like to ask a question, i.e., have you considered any specific data privacy regulations that may impact the extraction process? Warm Regards, Aneesa.
€750 EUR in 1 day
7.1
7.1

Hello I have experience with both Python and parsing SEC filings, and I have even completed related project, therefore I am sure I can help you. I have read acceptance criteria and I am able to meet Let's start!
€765 EUR in 2 days
7.2
7.2

• A single CSV or Excel file where each row is a firm-year filing and each column holds one of the compensation items above, clearly labeled. • A short read-me describing the extraction logic, any LLM prompts used, and the quality-control steps you applied. • A reproducible script or notebook so I can rerun the pipeline on future filings.
€750 EUR in 7 days
6.9
6.9

Warm Hello! Extracting executive compensation from 5,000 DEF 14A filings requires a reliable, reproducible pipeline that balances automation with rigorous quality control. I have over 9 years of experience in this field. I’ve built data extraction workflows for complex HTML documents and can deliver a scalable solution with high accuracy, structured outputs, and complete documentation. Here's how I can help: Build a hybrid HTML parsing + LLM extraction pipeline Extract base salary, stock awards/options, bonuses, incentive pay, and other compensation fields Generate a clean CSV/Excel with clearly labeled columns Flag unprocessed filings with detailed reasons Perform numeric validation and random QA to meet your accuracy targets Provide a reproducible Python script/notebook and a read-me with extraction logic, prompts, and QC process Are the 5,000 HTML files already downloaded locally with consistent naming, or should the pipeline also handle file organization? Do you need compensation for every Named Executive Officer in each filing, or only the principal executive (e.g., CEO)?
€1,125 EUR in 7 days
6.6
6.6

Hi There, Ready right now I'm ready to key compensation details for each named executive pulled out and placed into a clean, structured file. I will show you sample for your satisfaction and project accuracy then we will go to start, so please contact me and share more details thanks. Check My Profile: https://www.freelancer.pk/u/WelcomeClient I would like to work on this project and can complete with 100% accuracy within the time frame. https://www.freelancer.pk/projects/excel/business-profit-loss-reporting-excel/reviews https://www.freelancer.pk/projects/data-entry/copy-listings-from-website-another/reviews Thanks, Umer
€750 EUR in 2 days
6.9
6.9

Hi, I have good working experience with Python,Data processing,Data Extraction,Scraping,Visualization. I can provide you the demo of my previous similar work which I have done earlier. Message me here. I am looking forward to an early and positive response. Regards, Shalu
€750 EUR in 7 days
6.8
6.8

Hello, I can build a reliable pipeline to process all 5,000 DEF 14A filings and extract executive compensation into a clean, structured CSV/Excel. My approach is to combine HTML parsing with an LLM only where needed, giving better accuracy while keeping the process fast and reproducible. Every filing will be validated with numeric checks, missing values will be flagged with reasons, and the final script/notebook will let you rerun the extraction on future filings without extra work. One question: are all 5,000 filings already downloaded locally in HTML format, or should the pipeline also handle different SEC filing layouts and table variations? Looking forward to discussing the best extraction strategy. Hopefully the only thing harder than parsing DEF 14A will be choosing a coffee break. ? Best regards, Dev S.
€1,500 EUR in 10 days
6.7
6.7

Good to see this project, I will build a Python pipeline that parses your 5,000 DEF 14A HTML filings, extracts the Summary Compensation Table and Grants of Plan Based Awards, and outputs a single CSV with base salary, stock options/awards, bonus/incentive pay, and other compensation per firm, year, and named executive. On a similar SEC parsing project, adding table structure validation against reported totals caught misaligned columns early, dropping field errors well below your 5% threshold. Questions: 1) Are the 5,000 HTML files already downloaded locally, or do I need to pull them from EDGAR using CIK/accession numbers? 2) Should each named executive get a separate row, or do you want one row per filing with columns grouped by executive? Looking forward to discussing further. Best regards, Kamran
€836 EUR in 25 days
6.4
6.4

I propose using NLP and custom data extraction scripts to accurately extract compensation details from 5,000 DEF 14A proxy statements. I will utilize NLP models, regular expressions, BeautifulSoup, NLTK, and Pandas for parsing and data manipulation. I estimate a X-week turnaround time and will provide detailed documentation and a reproducible script. Your feedback on specific requirements is welcomed for fine-tuning. I aim to exceed expectations and establish a long-term partnership for future projects.
€1,350 EUR in 5 days
6.3
6.3

Hey! We’re a team of 62 professionals specializing in data extraction and SEC filing analysis with 9+ years of experience building automated parsing and validation workflows. Here's how we can help: - Extract executive compensation from thousands of DEF 14A filings - Build reproducible scripts with structured CSV or Excel outputs - Apply quality checks with flagged exceptions and validation reports - Balance automation and manual review for high data accuracy Could you clarify whether the HTML filings follow a consistent SEC format, and do you have any preferred Python libraries or LLM models for the extraction pipeline?
€1,125 EUR in 7 days
5.4
5.4

Hey, 5,000 DEF 14A filings is a solid batch but the summary comp table structure is consistent enough to automate most of it. I’d parse the HTML tables directly for salary, bonus, and stock awards, then use an LLM pass only for the messy footnote cases where numbers get buried in text. The real headache is always executives who show up under different name spellings across years, so I’d build a matching step for that before anything else. Want me to send a quick plan on the automation-to-manual-review split before you commit?
€750 EUR in 14 days
5.4
5.4

With my expertise in web development and data analysis, I offer a highly capable solution to your SEC filing data extraction project. I have a keen understanding of data processing and my experience with Python will be invaluable in effectively parsing through the 5000 DEF 14A proxy statements. My core skill set notably includes Data Analysis, Data Processing and Natural Language Processing which are directly transferrable to this project. To ensure quality control, I plan to implement a two-step approach -- a hybrid of automation and manual review. Utilizing powerful tools like LLM and applying my understanding in Natural Language Processing, I can effectively extract the compensation details you require. My focus will also be on creating a reproducible script or notebook for future use ensuring your project’s longevity. Having built scalable web applications and API integrations before, I have the knowledge to create a clean and structured file, meeting your deliverables. In alignment with your acceptance criteria, I'll run numeric sanity checks to verify all output total figures, ensuring no negative salaries or mismatched totals. Let me handle these filings for you with dedication and efficiency. First-time right.
€1,200 EUR in 7 days
4.9
4.9

Hi i have expertise in html file conversation to CSV file with the help of python.I can help you to write your all assignments with diagrams and tables plus Matlab calculation.
€1,125 EUR in 7 days
5.2
5.2

Hi there yeah I've understood the project and I'm sure that I can do this ASAP Kindly send me a message we'll discuss further Really looking forward to hearing from you Thank you
€1,125 EUR in 3 days
5.3
5.3

Your 95% success threshold will fail if you rely purely on regex or purely on an LLM. DEF 14A tables use inconsistent HTML structures across issuers, nested spans for footnotes, and merged cells that break naive parsers. The real risk is silent data corruption where your script extracts the wrong column and you don't catch it until after delivery. Quick questions - are you planning to validate against EDGAR's structured XML feeds where available, or is this purely HTML-only? And do you need the data tied back to CIK identifiers so you can merge with other SEC datasets? Here is the architectural approach: - PYTHON + WEB SCRAPING: Build a two-pass extraction pipeline where BeautifulSoup isolates Summary Compensation and Grants tables by section headers, then a fine-tuned GPT-4 model extracts numeric fields with confidence scores. - DATA PROCESSING + NLP: Use named entity recognition to map executive names across tables, handle footnote references that inflate raw numbers, and flag filings where table structure deviates from the 90th percentile pattern. - DATA VALIDATION: Implement cross-table reconciliation checks, flag negative values or missing required fields, and generate an audit log showing which 50 filings to manually review based on confidence thresholds. I've built similar hybrid pipelines for a hedge fund that processed 12K 10-Ks with a 97% first-pass accuracy rate. Let's schedule a 20-minute call to walk through your sample HTML files and confirm the edge cases before I scope the timeline.
€1,020 EUR in 30 days
5.4
5.4

Hola, soy Leo Sarmiento, especialista en extracción de datos y procesamiento de documentos con Python con más de 10 años de experiencia en parsing de estructuras complejas y pipelines de datos. El verdadero desafío de este proyecto no es recorrer 5.000 archivos, sino lidiar con la variabilidad en la presentación de las tablas de compensación dentro de los DEF 14A, donde los nombres de columnas y el orden de los conceptos cambian entre emisores y años; por eso, planteo un enfoque híbrido que combine BeautifulSoup para la extracción estructural de tablas con un modelo LLM para interpretar las cabeceras y clasificar los valores en los campos solicitados, aplicando además reglas de validación numérica para identificar valores atípicos o faltantes antes de la revisión manual de una muestra.
€1,000 EUR in 7 days
5.1
5.1

Hi, I will extract Table III compensation fields from your 5,000 DEF 14A HTML filings and deliver a validated CSV or Excel, a short readme with the extraction logic and prompts, and a reproducible script/notebook to rerun the pipeline. I built an extraction pipeline that processed 4,200 DEF 14A and 10 K filings to CSV. My approach: normalize HTML with BeautifulSoup and pandas.read_html, canonicalize headers with regex and a mapping table, use rule based extraction for well formed tables and OpenAI GPT 4 to resolve ambiguous labels and map grant rows to named executives. Automation will handle the majority of files and a focused manual review of heuristically flagged filings will cover edge cases so the final dataset meets your acceptance criteria. Quality control includes numeric sanity checks, totals reconciliation against table footings, automated unit tests, and a 50 filing random audit. Estimated turnaround for all 5,000 files is four weeks from file access. Are the HTML files available in a single cloud location I can access? Happy to jump on a quick chat. Ali Zain
€1,125 EUR in 7 days
4.8
4.8

Hi there, I understand you need to extract key compensation details from approximately 5,000 DEF 14A proxy statements in HTML format. This is crucial for ensuring structured and easily accessible executive compensation data. I propose using Python for efficient web scraping and data extraction. Utilizing libraries like BeautifulSoup and Pandas, I will parse the HTML documents and accurately extract the required details, organizing them into a clean and structured format such as CSV or Excel. This approach will ensure data accuracy and integrity, and the output can be easily manipulated for further analysis or visualization as needed. Best Regards, Khorshed Alam, RS Software
€750 EUR in 6 days
5.0
5.0

SEGOVIA, Spain
Payment method verified
Member since Mar 13, 2026
₹12500-37500 INR
₹75000-150000 INR
₹750-1250 INR / hour
$10-30 USD
$100-250 AUD
$30-250 NZD
₹600-1500 INR
₹600-1500 INR
$250-750 USD
$30-250 USD
₹12500-37500 INR
$30-250 USD
$30-250 USD
₹12500-37500 INR
$30-250 USD
₹1500-12500 INR
€30-250 EUR
$75-100 USD
₹12500-37500 INR
$30-250 USD