
Closed
Posted
Paid on delivery
Project Description We are looking for experienced freelancers or teams to provide or build a high-quality call center conversation dataset for AI training. We are open to both: * Custom data collection and annotation * Ready-made (pre-built) datasets that meet requirements -Languages Required * English – 500 hours * Hindi – 500 hours * Spanish – 500 hours Total: 1500 hours Data Requirements * Call center style agent–customer conversations only * Clean audio (no background noise) * Minimal silence and natural flow * WAV format * 16 kHz sampling rate (16-bit or higher) * Single channel preferred Transcription Requirements * Full transcription required * Speaker labels (Agent / Customer) * Timestamped alignment required Metadata Requirements (Must Confirm) * Source type (recorded or pre-built dataset) * Transcription type (AI or human-annotated) * Ability to refine AI transcripts if applicable * Speaker diarization accuracy * Timestamp alignment accuracy * Estimated WER (Word Error Rate) * Audio sampling rate details Privacy & Compliance * All personal data must be anonymized or removed * Method of de-identification must be clearly explained * Data must be legally collected with proper consent * Only for internal AI model training use Deliverables * WAV audio files * Transcript files (TXT or JSON) * Metadata (speaker labels, timestamps, language info) Requirements from Freelancer Please include: * Experience with speech datasets or ASR projects * Whether you can provide ready-made datasets * Tools and workflow used * Production capacity * Estimated cost per hour * Delivery timeline Important Note We are open to **ready-made datasets as long as they fully meet the above requirements and are legally authorized for AI training use**.
Project ID: 40470137
9 proposals
Remote project
Active 22 secs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs