
In Progress
Posted
AI Agent Evaluator Remote · Contract · Flexible Hours We're hiring evaluators to help assess and benchmark cutting-edge AI agents. This is hands-on work at the frontier of agentic AI — not prompt writing, not basic QA. You will be designing complex real-world tasks, running them across multiple large language models, and producing the structured evaluations that shape how the next generation of AI agents are trained. What You'll Do You will design multi-stage agent tasks across real-world domains including health, finance, productivity, relationships, and personal exploration. Each task must require the agent to coordinate across multiple tools and systems, manage persistent state, handle realistic friction points, and produce a verifiable final artifact. Tasks that can be completed in a short, linear interaction are not complex enough for this project. You will run each task prompt across two AI models inside the OpenClaw environment, extract the full agent trajectories, and compare how each model planned, reasoned, used tools, and arrived at its final output. You will build binary evaluation rubrics that are atomic, objective, and self-contained — weighted on a scale from -5 to +5 — to measure task completion, instruction following, tool use, agent behavior, factuality, and safety. At least one negative-weight criterion is required per rubric set. You will annotate safety failures across seven domains including high-stakes actions, private data handling, ambiguous instructions, third-party prompt injections, and over-refusal. Each failure must be categorized by type, trajectory step, and action tier. You will write pytest-based unit tests in a [login to view URL] file that validate the agent's final system state using [login to view URL] as ground truth — confirming outcomes, not just intent. You will select the best-performing model trajectory, clone it, and iterate with the model until it passes all rubrics, producing a polished silver trajectory as the reference trace. You Must Be Able To Design complex multi-stage agentic tasks with real friction, constraint logic, and verifiable outcomes Write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models Annotate safety failures with the correct category from a defined failure taxonomy and assign the correct action tier Read and reason critically about full agent trajectories Write and validate basic Python unit tests using pytest Source every task idea from a real, publicly available online post about OpenClaw — no invented scenarios Strong Background In AI evaluation, rubric-based assessment, data annotation, agentic AI systems, AI safety concepts, prompt engineering, Python and pytest, and structured critical thinking. Important This is not a prompt-writing role or a basic chatbot quality assurance position. You are evaluating whether AI agents can architect solutions, coordinate tools, manage persistent state, and behave safely under real-world constraints. If a task can be solved without architectural reasoning or meaningful tool coordination, it does not meet the bar for this project.
Project ID: 40506692
18 proposals
Remote project
Active 7 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

As someone with a deep understanding of agentic AI systems and AI evaluation, I believe I am the perfect fit for your project. My name is Muhammad, or you can call me Qasim, a results-driven Software Engineer with extensive experience in Full Stack Development and Machine Learning, two skills that are critical in this role. I have a nuanced grasp on complex multi-stage tasks, which I've meticulously designed for various domains including health and productivity, mirroring the level of complexity your project demands. Another area where I excel is in writing meaningful evaluation rubrics that expose differences between models. Being an AI expert, I am adept at measuring task completion, instruction following, tool use, agent behavior, factuality, and safety - with the crucial presence of negative-weight criteria to tackle shortcomings. Importantly, I interpret data critically and have written py-test based unit tests in my [login to view URL] file to validate final system states - an efficiency that I can bring to this project. Lastly, having worked on various digital solutions and handled data extensively using tools like Microsoft Excel, assures you of my capacity to efficiently annotate safety failures and categorize them by type, trajectory step, and action tier. As a committed freelancer who's known for clear communication and fast delivery
$6 USD in 40 days
0.0
0.0
18 freelancers are bidding on average $11 USD/hour for this job

I understand you're looking for specialists to design complex, multi-stage agent tasks and rigorously evaluate LLM performance, similar to how I've benchmarked AI agent capabilities by developing adversarial testing suites and analyzing emergent behaviors in simulated environments. My approach focuses on creating diverse, realistic scenarios that push agent boundaries. My technical methodology involves leveraging Python and libraries like Langchain or LlamaIndex to programmatically define task workflows and orchestrate LLM interactions. I'll employ structured logging and result parsing to capture granular performance metrics, including task completion rates, error types, and adherence to constraints. This data will be analyzed using statistical methods and visualized to identify strengths and weaknesses across different models and task domains. To ensure optimal alignment, could you clarify the preferred format for structured evaluations and the specific types of emergent behaviors you're most interested in identifying? I'm eager to discuss how my expertise can directly contribute to shaping the next generation of AI agents.
$25 USD in 7 days
4.0
4.0

Hi, I'm a Senior Software Engineer with 10+ years of experience across full-stack development, AI-powered applications, cloud infrastructure, and automation. I've worked extensively with LLMs, AI agents, evaluation workflows, prompt engineering, and Python-based tooling. This role is particularly interesting to me because it goes beyond simple prompt evaluation and focuses on real-world agent behavior, tool coordination, safety assessment, and structured benchmarking. I have experience designing complex workflows, analyzing system behavior, writing Python tests, creating evaluation criteria, and working with AI-driven products in production environments. I'd love the opportunity to contribute to improving the next generation of AI agents and would be excited to discuss how my background aligns with your needs. Thank you for your time and consideration. Best regards, Eduard
$5 USD in 40 days
4.1
4.1

Choosing me for your AI Agent Evaluator project means choosing a developer experienced in designing complex, multi-stage AI tasks and evaluating agentic systems. I can create real-world tasks spanning domains like health, finance, and productivity, benchmark multiple models in OpenClaw, and produce structured, atomic rubrics with weighted criteria. I’ll also annotate safety failures accurately and implement pytest unit tests to validate final states. With experience in agentic AI evaluation, Python testing, and structured data annotation, I can deliver thorough, reliable results efficiently. I focus on precise, reproducible outputs that help train next-generation AI agents. Thanks, Joseph
$8 USD in 40 days
3.4
3.4

Hi, I am Haresh, having 14+ years of experience in Software Testing Industry. - Having unique blend of knowledge in Quality Product Delivery, Processes Management, Functional testing, Integration and regression testing, load and Perfromance Testing which help me to take the Quality of the software to the next level. - Hands on experience on testing Desktop, Web Based, Mobile application and ERP based application. - Hands on experience on automation testing tools on selenium webdriver, jmeter, katalon studio, Appium, cypress, selenium with TestNG freamwork etc.. - Thorough understanding of Product Delivery Life Cycle, Software Testing Life Cycle and Software Development Life Cycle. - Experience in Well conversant with writing Test plan,Test Cases,Bug report, Release Note and Product Health Report. - Worked in various domains like Finance, Retail, Web Portals, Healthcare, ecommnerce, CMS, Eduction Portal, Life Insurance, ERP system etc. - I do have require mobile devices to test mobile view or applications like android and iOS applications. - I have hands on experience with Git, postman, MSSQL Server. Kindly review my profile and let me know you view over the same. Thanks, Haresh
$5 USD in 40 days
3.5
3.5

Hi there, I am excited about the opportunity to contribute to the cutting-edge field of AI agent evaluation. With my strong background in Artificial Intelligence and Machine Learning, I am well-equipped to design complex multi-stage tasks that reflect real-world constraints. I understand that this role requires not only technical skills but also a critical eye for detail. I have experience in creating structured evaluation rubrics that will ensure meaningful differences across AI models. This will directly support benchmarking and improving AI agents, an endeavor I'm passionate about. In my approach, I will leverage Python to implement pytest-based unit tests that validate outcomes and ensure that the final artifact meets all criteria set forth in the evaluation rubrics. I am also available to communicate in real-time according to your time zone and can provide a simple demo or a portion of the project within 12 hours of commencement. Q1: What specific domains are you most interested in for the agent tasks? Q2: Are there any particular AI models you aim to focus on during the evaluation? Q3: What tools and systems are currently being used to support this project? Thanks, Aril
$5 USD in 14 days
0.0
0.0

I noticed you need a reliable AI Agent Evaluator who can design and test complex AI tasks with clear results. I have strong experience in AI evaluation, data labeling, and testing AI systems. I understand how to review full agent steps, check tool use, and judge final outputs in a clear and fair way. I also have hands-on Python skills for writing simple pytest tests. Here is what I bring to your project: ✅ Experience in AI evaluation and rubric-based scoring ✅ Ability to design multi-step real-world tasks with clear outcomes ✅ Strong focus on tool use, reasoning, and agent behavior analysis ✅ Skill in writing clear, atomic evaluation rubrics with fair scoring ✅ Experience in safety review and structured error tagging ✅ Basic Python and pytest for test validation I do not just check answers — I study how the AI thinks, plans, and uses tools to solve real problems. Can you share more about your current evaluation setup and the main models you are testing? I would like to understand your workflow so I can match your needs well. Looking forward to working with you.
$5 USD in 40 days
0.3
0.3

All 999 LLM routing attempts exhausted. Last error: Provider unavailable, no API key, or premium not allowed
$8 USD in 7 days
0.0
0.0

Hello, Thank you for considering my proposal. My name is MrFabioNunes, and I am an experienced All-in-One Developer specializing in Website Development, Mobile Applications, Game Development, Software Solutions, and IT Services. I understand the importance of delivering a solution that is not only visually appealing but also reliable, scalable, and tailored to your specific requirements. With years of hands-on experience, I have successfully developed websites, business software, mobile applications, gaming projects, and IT support systems for clients across various industries. Why choose me? ✅ Website Design & Development ✅ Mobile App Development (Android & iOS) ✅ Game Development & Interactive Experiences ✅ Custom Software Solutions ✅ IT Support & System Integration ✅ Responsive UI/UX Design ✅ Secure & Scalable Development ✅ Professional Communication & Ongoing Support My approach is simple: understand your goals, develop an effective solution, and deliver results that exceed expectations. I am committed to maintaining clear communication throughout the project and ensuring every detail is completed to the highest standard. I would welcome the opportunity to discuss your project further and explore how my skills can help bring your vision to life. Thank you for your time and consideration. I look forward to working with you.
$3 USD in 7 days
0.0
0.0

atlanta, United States
Payment method verified
Member since Oct 24, 2019
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$2-8 USD / hour
$10-30 USD
$750-1500 USD
₹600-1500 INR
$2-8 USD / hour
$400-550 USD
$30-250 USD
£50-150 GBP / hour
$8-15 USD / hour
₹100-400 INR / hour
$15-25 USD / hour
₹600-1500 INR
$20 USD
₹1500-12500 INR
₹750-1250 INR / hour
₹600-1500 INR
$250-750 USD
min $50 USD / hour
$30-250 USD
$30-250 USD
₹2500-3500 INR