
Closed
Posted
I have worked on multiple AI data generation and model evaluation projects involving both Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF). My work included designing high-quality prompts, analyzing model responses, identifying reasoning and factual errors, and creating golden responses that served as the expected reference outputs. I worked extensively with repository-based tasks using Cursor, where I understood large codebases, executed run scripts, analyzed parse scripts, and validated both Pass-to-Pass (P2P) and Fail-to-Pass (F2P) test cases to ensure the generated solutions were functionally correct. I also reviewed code quality, verified outputs against expected behavior, and ensured compatibility with existing repositories. For tool-use and agentic AI projects, I created prompts to evaluate model capabilities, intentionally designed scenarios to expose model weaknesses, and assessed how effectively models selected and invoked the available tools. I wrote detailed evaluation rubrics defining the expected reasoning process, required tool schemas, correct tool calls, expected outputs, and grading criteria. I compared multiple model responses for correctness, instruction following, reasoning quality, completeness, and safety, while identifying issues such as hallucinations, incorrect tool usage, missing steps, and formatting problems. Throughout these projects, I followed strict quality guidelines, produced consistent annotations, documented evaluation decisions, and collaborated through review workflows to improve dataset quality and model performance. This experience has given me strong expertise in AI data annotation, prompt engineering, rubric writing, code evaluation, tool-use assessment, and LLM response evaluation.
Project ID: 40591662
28 proposals
Remote project
Active 11 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
28 freelancers are bidding on average ₹2,604 INR/hour for this job

Greetings, Thank you for considering my application for this project. As an AI Engineer and Python Developer with over 8+ years of experience, I bring a wealth of knowledge and expertise in the field of Python, Deep Learning. I have carefully reviewed the project description and am eager to discuss your specific needs and requirements in more detail. My commitment is to provide dedicated support and consistent follow-up throughout the project's lifecycle. Please feel free to reach out to me to further discuss how I can contribute to the success of your project. Looking forward to the opportunity of working together. Best regards, KuroKien
₹2,500 INR in 15 days
6.7
6.7

Hi, As per my understanding: You are looking for an AI specialist with hands-on experience in SFT, RLHF, prompt engineering, LLM evaluation, repository-based code validation, and agentic AI workflows. The ideal candidate should be capable of producing high-quality annotations, evaluating model responses, identifying reasoning and factual issues, creating detailed evaluation rubrics, and maintaining consistent quality standards throughout the project. Implementation approach: I have practical experience working on AI data generation and model evaluation projects involving SFT and RLHF. My work includes designing high-quality prompts, creating golden responses, evaluating reasoning quality, and identifying hallucinations and factual inconsistencies. I have also handled repository-based tasks using Cursor, validating P2P and F2P test cases, reviewing code quality, and ensuring compatibility with existing codebases. Additionally, I have developed evaluation rubrics for tool-use and agentic AI projects, assessing model reasoning, tool selection, instruction adherence, and overall response quality while following strict annotation guidelines and review workflows. A few quick questions: 1. Which LLMs and evaluation platform will be used for this project? 2. Will the work primarily involve prompt engineering, code evaluation, or RLHF annotation? 3. Are there any project-specific quality guidelines or annotation standards that I should follow from the beginning?
₹2,500 INR in 40 days
5.5
5.5

I have hands-on experience with SFT and RLHF projects, including prompt engineering, AI response evaluation, rubric creation, and golden response generation. I've worked extensively with repository-based tasks using Cursor, validating P2P/F2P tests, debugging code, and reviewing solution quality. I also evaluated tool-use and agentic AI workflows, assessing reasoning, tool selection, instruction following, and overall correctness. I'm experienced in AI data annotation, code evaluation, prompt design, and maintaining high-quality datasets under strict review guidelines.
₹2,500 INR in 40 days
4.7
4.7

Re: AI Data Generation and Evaluation My initial assessment for creating high-quality AI training data and evaluation frameworks suggests a stack using Python for scripting/analysis and a structured database like PostgreSQL for managing annotations and results. I have direct experience implementing detailed evaluation frameworks for agentic AI. For instance, on a past project, I designed a comprehensive rubric to evaluate a model's tool-use capabilities. This identified critical failure modes in its reasoning and function-calling logic, leading to a 40% improvement in task completion after retraining. I propose the following key steps: 1. Deep dive into your existing SFT/RLHF workflows and quality guidelines to align on evaluation criteria. 2. Develop initial prompt sets and evaluation rubrics for a pilot task to calibrate on quality and expected outputs. Happy to elaborate on my approach. Regards, Anton K.
₹2,500 INR in 7 days
4.2
4.2

Hello there, we are a team of AI/Ml automation, Python, Web and Mobile App developers. Please, send me a message to discuss the work and finish in no time. Thanks Ashish.
₹2,500 INR in 40 days
4.4
4.4

Hi, I can support your AI data generation and model evaluation work, including SFT, RLHF, prompt creation, golden response writing, rubric development, codebase-based evaluation, and tool-use assessment. The best solution is to first review your task guidelines, annotation standards, expected output format, quality benchmarks, and review workflow. I’ll then work on prompt design, response comparison, reasoning/factual error detection, tool-call validation, and detailed grading based on your criteria. I’m comfortable with Python, C/C++, repository-based tasks, Cursor workflows, running scripts, checking parse scripts, validating P2P/F2P test cases, reviewing generated code, identifying hallucinations, assessing instruction following, and writing clear evaluation rubrics. Deliverables will include: * High-quality prompts * Golden/reference responses * SFT/RLHF data support * Model response evaluation * Reasoning and factual error analysis * Tool-use and agentic workflow assessment * Rubrics with grading criteria * Codebase task validation * P2P and F2P test verification * Clear annotation notes * Consistent review documentation I’ll focus on accuracy, consistency, clear reasoning, and strict guideline adherence so the final dataset and evaluations are reliable for improving model performance. Best regards Ankit
₹2,500 INR in 40 days
3.3
3.3

A hidden risk in your RLHF pipeline is inconsistent prompt formatting that can skew evaluation metrics, a pitfall often missed in prompt engineering. I’d start by building a validation layer in Python and a tiny C helper to normalize prompts before they hit the model. That way the generated golden responses stay comparable across SFT and RLHF runs, and any C++ parser can reuse the same data. One common mistake I see is skipping the step that verifies tool schemas against the model’s output format. You’ll see a script that runs after each generation to compare the called tool signature with the schema, using mining checks on RL logs. Result is a clean dataset where errors are caught early and model scores reflect true capability.
₹2,500 INR in 40 days
2.3
2.3

Hi, I hope you're doing well. I understand you need an AI data generation and evaluation specialist with experience in SFT, RLHF, prompt engineering, and LLM response assessment. The goal is to improve model quality through accurate data creation, detailed evaluation, and consistent feedback based on technical and behavioral criteria. I will support AI training workflows by designing high-quality prompts, generating and validating datasets, evaluating model outputs, creating clear rubrics, and identifying issues such as hallucinations, reasoning errors, incorrect tool usage, and instruction-following failures. I can also assist with repository-based code evaluation, test validation, and detailed annotations to ensure reliable model performance improvements. My focus is on delivering accurate AI evaluation work with strong attention to quality, consistency, documentation, and measurable improvements in model reliability. Best regards, Heorhii
₹2,500 INR in 40 days
1.4
1.4

Hello! Your AI project, Proficient AI Data Generation and Evaluation, is a strong match for my experience. I can develop the complete solution with OpenAI/LLM integration, RAG, secure APIs, automation workflows, database design, and a clear admin dashboard. The focus will be reliable results, clean architecture, and a system ready for real users. Please share your data sources, required integrations, and expected AI output so I can suggest the right implementation plan. Portfolio: https://www.freelancer.in/u/heenafullstacken Thank you, Heena | A Plus IT House
₹2,500 INR in 44 days
0.0
0.0

Your project highlights the critical need for high-quality AI data generation and evaluation, particularly in ensuring models respond accurately and effectively. To address this, I would leverage my skills in prompt engineering and evaluation rubric creation to design targeted prompts that elicit reliable model responses. By systematically analyzing outputs against established benchmarks, I can identify and rectify reasoning errors, ensuring the model's performance aligns with your expectations. My recent work involved developing comprehensive evaluation frameworks for AI models, where I successfully identified weaknesses and enhanced performance through rigorous testing and documentation. What specific outcomes are you hoping to achieve with this AI data evaluation project? ### Why this approach - Directly addresses the client's needs by focusing on prompt engineering and evaluation. - Establishes credibility through relevant past work without exaggeration. - Encourages the client to engage by asking a targeted question about their objectives.
₹2,500 INR in 7 days
0.0
0.0

I am an experienced AI Data Annotation and LLM Evaluation specialist with hands-on expertise in SFT, RLHF, prompt engineering, and AI response evaluation. I have worked on repository-based coding tasks involving Python, C, and C++, validating P2P/F2P test cases, reviewing code quality, and analyzing large codebases using Cursor. My experience includes designing evaluation rubrics, creating golden responses, assessing tool-use workflows, identifying reasoning and factual errors, and ensuring high-quality annotations under strict QA guidelines. I am detail-oriented, reliable, and committed to delivering accurate, consistent, and high-quality work. I would be excited to contribute to your AI model development project.
₹2,500 INR in 40 days
0.0
0.0

Hello, I have seen your are looking AI data evaluation specialist working in Python, Prompt Engineering and Reinforcement Learning I am having 7 years of experienced as AI developer and I will work according to your AI data generation, SFT/RLHF evaluation, prompt writing, golden responses, code task review, tool-use assessment and rubric creation Please share more details of your project. Looking forward to your response Thanks
₹2,500 INR in 40 days
0.0
0.0

It sounds like you're facing challenges in generating high-quality data and evaluating AI models effectively. A critical first step would be to design robust evaluation rubrics tailored to the specific tasks and models involved, ensuring all necessary criteria are covered. I recently helped a tech startup streamline their model evaluations by creating detailed rubrics that exposed weaknesses and improved their dataset quality significantly. Could we discuss the specific models and tasks you're working with to ensure I can provide the most relevant support? — Collen
₹2,500 INR in 7 days
0.0
0.0

Hi, I have solid experience in this type of work and can handle exactly what you’re looking for with clean, professional results. I focus on delivering quality, staying on schedule, and making sure everything meets your expectations. I’m ready to get started and can adapt to your preferred style or requirements. We can discuss the details further in chat to clearly define the scope and make sure everything is aligned. Looking forward to working with you. Thanks,
₹2,500 INR in 40 days
0.0
0.0

Hello, My name is Bharghav, and I bring over 10 years of expertise in Python, C++, and C programming, specifically tailored for advanced AI data generation and model evaluation. I understand you need rigorous evaluation for repository-based tasks, tool-use, and RLHF. I will execute your run scripts, analyze codebase structures, and validate both P2P and F2P test cases in Python and C++ to ensure functional correctness. I will also design precise prompts and evaluation rubrics to identify reasoning errors or hallucinations. Please open the chat so we can discuss your project requirements in more detail. Best regards,
₹4,009 INR in 3 days
0.0
0.0

I am an AI/ML Engineer with hands-on experience in AI data annotation, Prompt Engineering, RLHF, Supervised Fine-Tuning (SFT), and LLM evaluation. I have worked on multiple AI model evaluation projects involving prompt creation, response evaluation, golden response writing, and quality assurance. I also have experience with repository-based evaluation using Cursor, including understanding large codebases, executing run scripts, validating P2P/F2P test cases, reviewing generated code, and ensuring compatibility with existing repositories. My technical skills in Python, C++, C, and Git enable me to analyze, debug, and evaluate complex solutions effectively. I have worked on tool-use and agentic AI projects by creating evaluation prompts, assessing tool selection, writing detailed rubrics, and comparing model outputs for correctness, reasoning, instruction following, safety, and completeness. • RLHF, SFT, and LLM evaluation experience • Prompt engineering and AI response evaluation • Repository-based code validation (P2P/F2P) • Python, C++, C, and Git proficiency • Detail-oriented, reliable, and deadline-focused
₹2,500 INR in 40 days
0.0
0.0

Dear Client, My expertise aligns perfectly with your search for an expert in AI data generation and evaluation, especially for Supervised Fine-Tuning andd Reinforcement Learning from Human Feedback (RLHF), code evaluation, and prompt engineering. I excel at designing high-quality prompts, thoroughly evaluating model responses (identifying factual and reasoning errors), and creating precise "golden responses." My domain knowledge includes complex repository-based tasks using tools like Cursor: analyzing large codebases, executing scripts, performing P2P/F2P validation, and conducting in-depth code quality reviews to ensure compatibility and functionality within existing architectures. What makes me the best candidate for this project? My unique combination of advanced technical skills (Python, C++, C understanding) and proven experience in each of your detailed requirements, coupled with my rigorous, quality-oriented methodology, enables me to deliver results tHat not only meet but exceed your expectations. To further tailor my proposal, could you provide more details on the complexity of the "large codebases" and any preferred LLM evaluation metrics? I am available for a discussion to demonstrate how I can directly contribute to your project's success. Sincerely, Luis ALberto Salas Ortiz
₹2,500 INR in 40 days
0.0
0.0

Hello, I'm interested in this opportunity because my experience closely matches the role's requirements. I have worked on AI data generation, Supervised Fine-Tuning (SFT), RLHF, prompt engineering, LLM response evaluation, and rubric writing. I have evaluated model outputs for reasoning, factual accuracy, instruction following, safety, and tool usage while creating high-quality reference responses. I also have experience with repository-based tasks, code evaluation, test validation, and agentic AI workflows. My attention to detail, analytical mindset, and commitment to producing consistent, high-quality annotations make me confident that I can contribute effectively to your team from day one.
₹2,800 INR in 24 days
0.0
0.0

Hi, I bring hands-on expertise in AI data generation and LLM model evaluation covering SFT and RLHF workflows. I craft high-quality prompts, build standardized evaluation rubrics, and spot hallucinations, flawed reasoning and faulty tool calls.I’m proficient in Cursor for large codebase validation, P2P/F2P test case verification, and agentic AI tool-use assessment. I deliver consistent annotated datasets, fully document evaluation judgments, and align all work with strict quality standards. I can quickly adapt to your project pipelines to boost dataset integrity and overall LLM performance.
₹2,500 INR in 40 days
0.0
0.0

With a deep understanding of AI model development, data mining, and Python, I bring 3+ years of valuable experience to the table. My proficiency lies in not only engineering AI models but also in maximizing their potential through intelligent data generation and evaluation— exactly what your project needs. Having extensively worked with Supervised Fine-Tuning and Reinforcement Learning from Human Feedback, I demonstrated my capability to craft high-quality prompts, analyzed model responses, and created golden responses to serve as reference outputs. My skillset doesn't stop there: I am also well-versed in tool-use and agentic AI projects. I've adeptly designed prompts to evaluate model capabilities and created situational scenarios to expose weaknesses thus refining the machine model even more. One of the hallmarks of my work is meticulously following strict quality guidelines while producing consistent annotations. By documenting my evaluation decisions and collaborating openly during review workflows, I ensure dataset improvements and optimized model performance. Additionally, my penchant for building end-to-end AI solutions means that I am as involved with preprocessing and training as I am with deployment and MLOps. Overall, my comprehensive understanding of the field paired with my strong software engineering background would undoubtedly prove invaluable to your project's timely completion and desired outcomes.
₹2,500 INR in 40 days
0.0
0.0

Bengaluru, India
Member since May 21, 2025
₹3000-4000 INR
₹600-1500 INR
$50-300 USD
₹600-601 INR
₹600-1500 INR
₹1500-12500 INR
$250-750 USD
₹750-1250 INR / hour
$5000-10000 CAD
₹12500-37500 INR
$750-1500 AUD
$10-30 USD
₹750-1250 INR / hour
₹1500-12500 INR
₹100-400 INR / hour
$3000-5000 USD
$10-30 USD
₹600-1500 INR
$10-50 USD
₹1500-12500 INR