
Closed
Posted
Paid on delivery
# Speech AI / Multilingual Cognitive AI Engineer (Contract – 1 Month) ## About Us We are building an AI-powered multilingual cognitive decline screening platform that analyzes short speech recordings to identify early cognitive impairment using speech and language biomarkers. Our current platform already includes: * FastAPI backend * Whisper-based ASR pipeline * Acoustic and linguistic feature extraction * Cognitive scoring pipeline * Web interface * RunPod deployment * Firebase integration * English MVP We are now looking for a Speech AI engineer to improve the core multilingual intelligence of the platform. --- ## Responsibilities * Evaluate the existing ASR and cognitive screening pipeline * Improve transcription robustness for English, Hindi, and Hinglish * Improve handling of code-switched speech * Optimize Whisper or evaluate alternative multilingual ASR approaches * Build confidence scoring and error handling for low-confidence transcripts * Improve acoustic and linguistic feature extraction * Benchmark model performance using objective evaluation metrics * Document findings and recommendations * Work closely with the founding team on iterative experiments --- ## Required Skills * Python * PyTorch * Speech AI / ASR * Whisper or similar speech models * Hugging Face * Audio processing (Librosa, Torchaudio, etc.) * FastAPI * Git --- ## Preferred * Experience with multilingual ASR * Hindi/Hinglish speech processing * Code-switching research * Healthcare AI or cognitive assessment * Model evaluation and benchmarking --- ## Deliverables By the end of the engagement, the engineer should deliver: * Improved multilingual ASR pipeline * Better English, Hindi, and Hinglish performance * Benchmark report with accuracy metrics and failure analysis * Production-ready inference improvements * Fully documented code integrated into the existing VoiceMind repository --- ## Duration 1 Month (Remote) ## Compensation ₹15,000 (Fixed Contract) Outstanding performance may lead to a longer-term research collaboration.
Project ID: 40578650
31 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
31 freelancers are bidding on average ₹24,243 INR for this job

Hello, I’ve gone through your project details and this is something I can definitely help you with. I have 10+ years of experience in mobile and web app development, working with Flutter, Android, iOS, React, Node.js, and APIs. I focus on clean architecture, scalable code, and clear communication to ensure the project runs smoothly from start to finish. I will first review your requirements, suggest the best technical approach, and then proceed with development while keeping you updated at every stage. Here is my portfolio: https://www.freelancer.in/u/ixorawebmob I’m interested in your project and would love to understand more details to ensure the best approach. Could you clarify: 1. Do you need this for mobile, web, or both? 2. Do you already have UI/UX designs or should we create them? 3. Will there be any third-party API or payment gateway integration? 4. What is your expected timeline for completion? 5. Are there any reference apps or websites you like? Let’s discuss over chat! Regards, Arpit Jaiswal
₹27,750 INR in 30 days
7.4
7.4

Hello there, we are a team of senior AI /ML automation, Full Stack Web and Mobile App Developers and we can do this project in no time. Thanks Ashish Kumar.
₹25,000 INR in 7 days
5.9
5.9

Hi, I reviewed your project: Multilingual Speech AI Engineer. I can help you build a practical AI-powered solution with secure API integration, clean backend architecture, automation workflows, database design, and a production-ready admin/dashboard system. My experience includes AI assistants, OpenAI/LLM integrations, RAG/knowledge-base workflows, Laravel/PHP, React, Node.js, APIs, databases, and deployment. Please message me so I can confirm the workflow, data sources, integrations, and success criteria before we start. Portfolio: https://www.freelancer.com/u/irfanui Regards, Mohammad 4th Dimension Partners
₹37,500 INR in 31 days
5.5
5.5

I’ve built ASR-enhanced platforms for cognitive and healthcare AI, with direct experience tuning Whisper and Wav2Vec for multilingual Indian datasets and dealing with code-switched Hindi-English data. For this, transcription robustness means custom decoding strategies and potentially using Whisper-large + fine-tuning or benchmarking alternatives from HuggingFace. I’d dig into your FastAPI backend, tune Whisper for better Hindi/Hinglish, add confidence scoring and fallback logic for low-certainty outputs, then document and benchmark all changes. I use Librosa, Torchaudio, and run full evals with established metrics on a fixed set. Is your ASR pipeline already using diarization, or just straight chunk-processing? Reply if you want fast iteration and clear code. Pradeep
₹25,000 INR in 7 days
5.2
5.2

Hi There , Good afternoon! I’ve carefully checked your requirements and really interested in this job. I’m full stack node.js developer working at large-scale apps as a lead developer with U.S. and European teams. I’m offering best quality and highest performance at lowest price. I can complete your project on time and your will experience great satisfaction with me. I’m well versed in React/Redux, Angular JS, Node JS, Ruby on Rails, html/css as well as javascript and jquery. I have rich experienced in Git, Android, Python, Mathematics, Hugging Face, FastAPI, PHP and Documentation. For more information about me, please refer to my portfolios. I’m ready to discuss your project and start immediately. Looking forward to hearing you back and discussing all details.. I await your immediate response
₹27,750 INR in 3 days
4.5
4.5

Hi, I can evaluate and improve your existing multilingual speech pipeline for English, Hindi, and Hinglish, with focus on code-switched transcription, confidence scoring, feature quality, benchmarking, and production-ready FastAPI integration. The best solution is to first audit the current Whisper pipeline, preprocessing, decoding settings, feature extraction, scoring flow, and RunPod deployment. I’ll then benchmark the baseline using WER/CER and language-specific failure analysis, test Whisper optimizations or suitable multilingual alternatives, and improve handling of accents, mixed-language speech, noise, and low-confidence transcripts. I’m comfortable with Python, PyTorch, Whisper, Hugging Face, Librosa, Torchaudio, FastAPI, audio preprocessing, multilingual ASR evaluation, confidence thresholds, feature extraction, Git, and production inference optimization. Deliverables will include: * Existing pipeline audit * English, Hindi, and Hinglish ASR improvements * Code-switching optimization * Confidence scoring and fallback handling * Acoustic and linguistic feature improvements * WER/CER benchmarking * Failure and edge-case analysis * Production-ready repository integration * Documented code and recommendations I’ll focus on measurable improvements, reproducible experiments, and clean integration into the existing VoiceMind platform. Best regards Ankit
₹12,500 INR in 2 days
3.7
3.7

Hi, Your project aligns closely with my experience in Speech AI, multilingual NLP, and production-grade AI systems. I have worked extensively with Python, PyTorch, Whisper, Hugging Face, FastAPI, Librosa, and Torchaudio to build and optimize speech processing pipelines. I can evaluate your existing VoiceMind pipeline, improve transcription quality for English, Hindi, and Hinglish, optimize code-switched speech handling, implement confidence scoring for low-quality transcripts, and enhance both acoustic and linguistic feature extraction. I'll also benchmark the improved models using WER, CER, latency, and other relevant metrics, providing a detailed performance report with failure analysis and recommendations. My approach focuses on maintaining production readiness while improving accuracy and robustness. The final deliverables will include clean, well-documented code integrated into your existing repository, production-ready inference improvements, and reproducible evaluation scripts. I have experience building AI applications involving speech recognition, multilingual NLP, healthcare-related AI, and model optimization, and I'm comfortable collaborating in an iterative research environment.
₹15,000 INR in 7 days
3.7
3.7

Your current stack already covers the essential components for a strong multilingual screening platform, so the main opportunity here is improving robustness and evaluation quality across noisy multilingual and code-switched speech. I can help by reviewing the existing Whisper-based pipeline end-to-end, identifying transcription failure patterns for English/Hindi/Hinglish, and implementing improvements focused on inference stability, confidence scoring, and benchmark-driven optimization. This includes evaluating decoding strategies, preprocessing pipelines, segmentation behavior, and alternative multilingual ASR approaches when Whisper limitations become evident. From the engineering side, I can work directly on the FastAPI integration layer, production-ready inference improvements, structured evaluation workflows, and reproducible benchmarking using objective metrics such as WER/CER alongside failure analysis for low-confidence outputs. I also have experience building scalable backend systems and ML-oriented APIs, which is important when moving experimental AI workflows into reliable production services. The deliverables would include documented improvements integrated into your repository, clearer observability around transcription quality, and a cleaner experimentation workflow for future iterations. For a 1-month engagement, I would structure the work in short iterative milestones so improvements can be validated continuously instead of waiting for a final delivery.
₹36,791.11 INR in 30 days
3.8
3.8

Hi, We can help with multilingual speech engineering tasks, including speech processing, AI model integration, and optimization for accurate voice solutions. Our team has experience with Python, AI/ML, and language technologies. Ready to discuss your requirements and start immediately. **Regards,** **Acute Tech Solutions**
₹25,000 INR in 7 days
1.9
1.9

As a seasoned Speech AI engineer, my proficiency spans across the critical skills outlined for your project. My extensive experience in Python, PyTorch, Hugging Face, and FastAPI are a perfect match for your existing technology stack. I have successfully improved ASR pipelines using similar Whisper / multilingual models before. Moreover, communication barriers may arise due to code-switched speech in linguistic analysis, which I've tackled effectively in past projects. Not only am I skilled in developing impactful AI/ML models but also proficient in benchmarking their performance to assure top-notch results and objective evaluation metrics - exactly what you're seeking. My prior work on healthcare AI and cognitive assessment gives me an edge as your project concerns cognitive decline screening. Utilizing audio processing libraries like Torchaudio and Librosa to optimize acoustic and linguistic feature extraction falls directly under my skillset. My proven track record of driving real business outcomes through advanced systems complements the ROI-driven approach your company seeks. Let's leverage my expertise for a highly efficient, optimized voiceMind platform.
₹26,000 INR in 10 days
2.7
2.7

The code-switch detection often collapses when Whisper's language token gets confused by rapid Hindi-English alternation, a scenario where Hugging Face language models could help. I’ll isolate the language-id step, run a short-window classifier on the audio, and feed the proper language hint back into Whisper, ready to start immediately. That way the pipeline keeps high confidence scores, uses PyTorch with torchaudio, and falls back to a simple phoneme-based model for low-confidence segments. A common mistake is to trust Whisper's raw confidence without normalizing it across languages, which can hide systematic bias. We'll add a calibrated confidence layer and log failure cases, so the benchmark report will show clear strengths and weaknesses. You can expect a production-ready inference script, updated FastAPI endpoints, and documentation that fits directly into the VoiceMind repo.
₹20,000 INR in 3 days
1.4
1.4

Hello there, I read your project carefully and understand that you need to enhance a multilingual Speech AI pipeline for English, Hindi, and Hinglish by improving ASR accuracy, code-switching support, confidence scoring, and production readiness. I will evaluate the existing pipeline, optimize Whisper or suitable multilingual ASR models, improve feature extraction, benchmark performance, and deliver clean, well-documented code integrated into your FastAPI-based system. I am available for a quick call. One question: Which Whisper model and evaluation datasets are you currently using for English, Hindi, and Hinglish benchmarking? Regards, Rohit
₹20,000 INR in 20 days
1.3
1.3

Hi, see my work first then award your valuable project. I am experienced in **Python, FastAPI, Speech AI, Whisper, Hugging Face, and multilingual ASR pipelines**. I can evaluate your existing VoiceMind pipeline, improve English, Hindi, and Hinglish transcription quality, and enhance cognitive screening accuracy with production-ready improvements. I will optimize the ASR pipeline for multilingual and code-switched speech, implement confidence scoring and robust error handling, improve acoustic and linguistic feature extraction, benchmark the models with objective accuracy metrics, document the results, and integrate clean, well-documented code directly into your existing repository while collaborating closely with your team throughout the project. Let's discuss further on chat and I can help strengthen VoiceMind's multilingual Speech AI pipeline for reliable cognitive screening. Tymofii
₹12,500 INR in 2 days
1.0
1.0

With your project's focus on multilingual cognitive decline screening platform, my team at Naina Technologies brings a unique blend of skills well-suited for the task. Our specializations align almost perfectly with your requirements. Python? Check. PyTorch? Check. Speech AI / ASR? Absolutely. Whisper or similar speech models? No problemo. This goes on to include Audio processing, FastAPI, and more - all areas we're well versed in. Now for the cherry on top: apart from delivering solutions with proficiency, we also have a cognizance of our responsibility towards the eventual end-users. Be it improving transcription robustness for languages like English, Hindi, and Hinglish, optimizing code-switching nuances, benchmarking accuracy metrics, or ensuring low-confidence transcripts are handled gracefully - our end goal is always to offer value and reliability to the target audience. Our solid record in managing fixed-length contracts within scheduled time frames without compromising quality should reassure you that hiring us would be to your advantage. We are keenly prepared to dedicate this next month in contributing towards your project's success. Moreover, we don't just bring ‘out-of-the-box’ solutions; we bring clear communication, documented work ethic and an affinity to iteratively experiment as well as learn which makes us prudent for long-term prospects too.
₹30,000 INR in 7 days
0.0
0.0

I've built exactly this kind of system: an ML model that classifies fluent vs disfluent speech from audio recordings (team project, GitHub: Nimitt07/Stuttering-Detection-AI-ML), plus a machine-translation NLP project. I work in Python with audio feature extraction and model evaluation, and I'm comfortable with Whisper-based pipelines, Hugging Face and FastAPI. As a CSE undergrad at IIT Dharwad I can dedicate the full month to improving your English/Hindi/Hinglish ASR robustness and delivering a benchmark report with documented code integrated into your VoiceMind repo. Happy to start with a quick evaluation of your current pipeline.
₹15,000 INR in 30 days
0.0
0.0

Hi. I understand you need to improve your existing Speech AI platform with better multilingual ASR performance, especially for English, Hindi, and Hinglish speech. The main goal is to enhance transcription accuracy, handle code-switching, improve confidence scoring, and make the pipeline production-ready. I’m fully capable of handling this. I have experience with Python, FastAPI, PyTorch, Hugging Face, Whisper, and audio processing workflows. I can optimize the ASR pipeline, evaluate model performance, improve feature extraction, and deliver documented, tested improvements. I’d like to review your current pipeline and challenges so we can plan the best approach. Looking forward to working with you.
₹25,000 INR in 5 days
0.0
0.0

As an experienced and passionate Data Scientist and AI/ML Engineer, I'm confident in my ability to deliver improved multilingual ASR pipelines that meet your needs for this project. Python, PyTorch, and FastAPI are just some of the tools I'm well-adapted with, ensuring I can optimize Whisper or evaluate alternate intelligible ASR approaches for your multilingual needs. My expertise extends beyond just speech processing and ASR. I've successfully built AI applications that automate business processes and improve user experiences; such as conversational AI chatbots and predictive ML models. That experience will prove invaluable in not only improving the existing pipelines but also in building confidence scoring and error handling mechanisms for low-confidence transcripts. Throughout my career, I have consistently aimed for scalable and maintainable solutions while focusing on delivering actionable insights from data. In line with this approach, you can expect robust, fully-documented code that seamlessly integrates into VoiceMind's repository upon completion. Rather than just ticking off boxes on a to-do list, I'm truly committed to finding the best multilingual ASR solution through research and iterative experimentation. With me on board, you can rest assured of great value for your money and potentially start a longer-term research collaboration. Let's connect soon!
₹12,500 INR in 5 days
0.0
0.0

Dear Client, I am excited to apply for your Speech AI / Multilingual Cognitive AI Engineer contract. My background is in AI, Machine Learning, NLP, and Python development, with hands-on experience building FastAPI applications, integrating Hugging Face models, and developing AI-powered solutions. I have experience working with Python, PyTorch, FastAPI, Hugging Face Transformers, audio processing workflows, and model deployment. I am comfortable evaluating existing AI pipelines, improving model performance, implementing robust error handling, and documenting experiments. I also have experience integrating AI models into production-ready web applications and working with Git-based collaborative development. For your project, I can analyze the current ASR pipeline, identify bottlenecks, improve transcription quality, optimize inference performance, enhance feature extraction, and provide benchmarking with clear evaluation metrics. I write clean, maintainable, and well-documented code and communicate progress regularly throughout the engagement. I am committed to delivering a reliable, production-ready solution within the project timeline and would welcome the opportunity to contribute to the VoiceMind platform and its mission of advancing multilingual cognitive assessment. Thank you for your consideration. I look forward to discussing how I can contribute to your project.
₹25,000 INR in 8 days
0.0
0.0

Hello, Your platform addresses a meaningful real-world challenge, and I'd be excited to contribute to strengthening its multilingual Speech AI capabilities. I can evaluate your existing Whisper-based pipeline, improve transcription accuracy for English, Hindi, and Hinglish, optimize code-switched speech handling, and enhance acoustic and linguistic feature extraction for more reliable cognitive screening. My approach includes benchmarking ASR performance with objective metrics, implementing confidence scoring and robust error handling, and delivering production-ready improvements that integrate seamlessly with your FastAPI architecture. I prioritize clean, well-documented code, collaborative development, and iterative experimentation to ensure the platform is both accurate and scalable for future research and clinical applications. Best regards, Rameen
₹12,500 INR in 1 day
0.0
0.0

Hi, I am an AI Systems Engineer specializing in FastAPI backends and Machine Learning pipelines, with a strong background in multilingual NLP and audio processing systems. I can optimize your VoiceMind screening platform's Whisper pipeline to robustly handle English, Hindi, and Hinglish code-switched speech. How I will deliver this: • Multilingual Whisper Optimization: I will fine-tune the decoding heuristics or configure specialized Hinglish/multilingual Hugging Face models to drastically reduce word error rates in code-switched audios. • Reliability & Error Handling: I'll integrate local Whisper timestamping with confidence-scoring layers to flag low-confidence outputs before they reach your scoring pipeline. • FastAPI & Clean Code: Having built robust, production-ready async FastAPI APIs, I will cleanly integrate these feature-extraction and scoring enhancements into your repository without introducing latency bottlenecks. Let's discuss how I can help your founding team benchmark and productionize these upgrades immediately.
₹25,000 INR in 7 days
0.0
0.0

Gurugram, India
Member since Jan 4, 2025
₹12500-37500 INR
₹37500-75000 INR
₹12500-37500 INR
€18-36 EUR / hour
$1500-3000 USD
₹750-1250 INR / hour
₹12500-37500 INR
€30-250 EUR
$10-20 USD
₹100-400 INR / hour
$10-60 USD
$250-750 USD
₹100-400 INR / hour
$15-25 USD / hour
$10-60 USD
₹12500-37500 INR
₹75000-150000 INR
€250-750 EUR
₹12500-37500 INR
₹12500-37500 INR
₹1000-2000 INR
$30-250 USD