
Closed
Posted
Paid on delivery
Hi, I’m looking for a highly experienced software engineer to collaborate with me on the project. The work involves creating and repairing advanced software-engineering tasks used to evaluate and train AI coding agents. Each task is based on a real GitHub repository, pull request, or commit and must include accurate instructions, strong regression tests, a correct oracle solution, reproducible Docker configuration, and complete metadata. I need someone who is strong in software architecture, Git, Docker, CI/CD, automated testing, and debugging across different programming languages. Experience with AI evaluation, coding agents, benchmark development, or adversarial test design would be especially valuable. The engineer must be able to identify subtle failures—such as mismatches between instructions and tests, incomplete test coverage, broken oracle solutions, non-reproducible environments, and packaging problems—and deliver tasks that pass strict automated and human review. This is detailed, hands-on engineering work for someone who values accuracy, reproducibility, and high-quality technical standards. I’m interested in building a reliable long-term working relationship with the right senior engineer. If this matches your background, I’d be glad to discuss the workflow, expectations, availability, and compensation.
Project ID: 40621309
116 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
116 freelancers are bidding on average $148 USD for this job

⭐⭐⭐⭐⭐ Experienced Software Engineer for AI Coding Agent Tasks ❇️ Hi My Friend, I hope you are doing well. I've reviewed your project requirements and see you're looking for a highly experienced software engineer. Look no further; Zohaib is here to help you! My team has successfully completed 50+ similar projects for software engineering tasks. I will create and repair advanced tasks for evaluating and training AI coding agents, ensuring they meet all your specifications. ➡️ Why Me? I have 5 years of experience in software engineering, focusing on Git, Docker, CI/CD, and automated testing. My expertise extends to debugging across multiple programming languages. I can identify subtle failures and ensure high-quality technical standards in all tasks, providing you with reliable solutions. ➡️ Let's have a quick chat to discuss your project in detail. I’d love to share samples of my previous work and how I can add value to your project. Looking forward to chatting with you! ➡️ Skills & Experience: ✅ Software Architecture ✅ Git ✅ Docker ✅ CI/CD ✅ Automated Testing ✅ Debugging ✅ Programming Languages ✅ AI Evaluation ✅ Benchmark Development ✅ Adversarial Test Design ✅ Regression Testing ✅ Metadata Management Waiting for your response! Best Regards, Zohaib
$150 USD in 2 days
8.1
8.1

Hi, this reads like benchmark engineering rather than ordinary feature work, and that distinction matters because the real risk is hidden inconsistency between task intent, oracle behavior, and reproducibility under automated review. I’ve built several systems where correctness depended on evaluation quality, not just code output. I usually structure this kind of work around three separate checks: task-spec integrity, environment reproducibility, and regression validity. For repository-based tasks tied to GitHub repos, pull requests, or commits, that separation is what prevents false passes and brittle failures. The closest match in my background is Python Bug Localization Using Transformer Models (CodeBERT + TreeBERT), where I had to design evaluation logic, confidence signals, and reproducible handoff around code-level fault analysis. Custom Feature Development & Integration is also relevant because it involved entering an existing codebase, reviewing architecture, and extending test coverage safely. I typically design these workflows so instructions, tests, oracle solution, and Docker behavior are validated independently before they are treated as one artifact. That makes it easier to catch packaging drift, incomplete coverage, and mismatched expectations early. I’d start by reviewing one representative task and mapping the failure surface: spec ambiguity, flaky regression behavior, oracle gaps, and container reproducibility. Thanks, Hercules
$200 USD in 7 days
7.2
7.2

Hello!! I have similar expertise and work experience for the **{{ ADVANCED SOFTWARE ENGINEER FOR AI EVALUATION }}** project. I have worked on complex software engineering projects involving backend systems, automation, testing frameworks, debugging, code quality improvements, and AI-driven development workflows. I have over 10 years of experience in software engineering, with strong expertise in **Python, JavaScript/TypeScript, Git, Docker, CI/CD pipelines, automated testing, software architecture, API development, and scalable system design**. I understand your requirements for creating and repairing advanced software-engineering evaluation tasks based on real GitHub repositories, commits, and pull requests. I can help design reliable tasks with clear instructions, reproducible environments, accurate regression tests, and validation workflows that meet strict quality standards. My experience includes analyzing complex codebases, identifying hidden issues, improving test coverage, debugging failures, setting up containerized environments, and ensuring software solutions are maintainable and reproducible. I focus on delivering high-quality engineering work with attention to accuracy, documentation, testing reliability, and long-term maintainability. I am also comfortable working with AI-assisted development workflows and evaluation environments. Thanks, Christina
$140 USD in 7 days
7.0
7.0

This project matches my experience building and validating production-grade software across Python, Java, Docker, Git, CI/CD, and automated testing. I can create and repair repository-based AI evaluation tasks with precise instructions, regression coverage, a verified oracle solution, reproducible containers, and complete metadata. My process will start by inspecting the source repository, relevant commit or pull request, existing test behavior, and runtime dependencies. I’ll then define the intended behavior, implement or repair the solution, and add focused tests for both normal cases and subtle failures such as instruction/test mismatches, missing edge-case coverage, packaging errors, and environment-specific bugs. I’ll run the full workflow in Docker from a clean checkout, verify deterministic results, check CI commands, and review the final task package as an evaluator would. For Python and Java projects, I can work with pytest, JUnit, Maven/Gradle, GitHub Actions, and Linux-based Docker images. With 10+ years in software development and 249+ delivered projects, I’m comfortable debugging unfamiliar codebases and documenting reproducible results. I can begin immediately and support a dependable long-term workflow. Will you provide the repository/commit and an existing task template for the first evaluation task? Muhammad Saad
$105 USD in 3 days
6.5
6.5

Hello Sir, Benchmark task authoring is unusual work in that the failure modes are subtle by design — a task can look complete and still be broken because the tests pass for the wrong reason, or the oracle solution happens to satisfy an under-specified instruction. My approach: treat each task as needing to fail correctly before it passes. That means verifying the regression tests actually catch the bug the PR fixed, checking the instruction is solvable from the stated context alone without leaking the solution, and confirming the Docker environment rebuilds deterministically rather than working once on my machine. My background covers Python, TypeScript, Docker, CI/CD pipelines, and test automation across API and integration suites — plus hands-on work with LLM tooling. Happy to discuss workflow, availability and compensation. I have worked here with more than 130+ clients. Best regards, Vishruth
$100 USD in 2 days
6.5
6.5

With over *X years of professional experience* and a deep understanding of software engineering, I have what it takes to successfully execute your project scope. As an intricate part of my daily operations, I possess exceptional skills in Docker, Git, CI/CD, debugging as well as automated testing and software architecture. My familiarity with multiple programming languages also ensures that I can effortlessly adapt to any task set before me. What puts me on top for this task is the intersection of my experience with AI evaluation and coding agents alongside bug fixing on real GitHub repositories that aligns with your requirements. Accountability matters to me and I have a proven record of identifying subtle failures and delivering tasks that pass strict automated and human review. I believe our values align perfectly; 'accuracy', 'reproducibility', and 'high-quality technical standards' are all words I live by in my operational framework. What's more? I am committed to building not just a software for you but also a long-term working relationship. Let's discuss the workflow structuring, expectations setting, availability as well as compensation in order to positively impact your software engineering endeavors as effectively as po
$140 USD in 7 days
6.5
6.5

Hi, I can support the creation and repair of repository-based coding-agent evaluation tasks with a strong focus on correctness, reproducibility, and adversarial coverage. My workflow begins by reconstructing the intended behaviour from the source repository, issue, commit, or pull request. I then verify that instructions, tests, oracle changes, metadata, and Docker environment describe the same contract. Regression tests will cover expected behaviour, edge cases, plausible incorrect implementations, and hidden assumptions without overfitting to one solution. I’ll validate tasks from a clean checkout, pin dependencies, remove network or environment ambiguity, inspect packaging and CI behaviour, and confirm that the oracle passes while meaningful mutants fail. Deliverables will use clear Git history and include concise notes explaining defects found, test rationale, and reproducibility checks. I’m comfortable working across unfamiliar codebases methodically and can align ongoing availability with your review cadence. Question 1: Which task format, benchmark schema, programming languages, and review harness are currently used? Question 2: Will assignments focus mainly on authoring new tasks, repairing rejected tasks, or independently auditing completed submissions? Regards, Houssame
$140 USD in 7 days
6.6
6.6

Hi there, I understand you need a senior software engineer to create, review, repair, and validate AI coding evaluation tasks built from real GitHub repositories while ensuring every task is accurate, reproducible, and meets strict engineering and testing standards. My approach will be to first analyze each repository, PR, or commit to understand the intended behavior, then recreate the task with precise instructions, reproducible Docker environments, complete metadata, and comprehensive regression tests. I'll validate the oracle solution, identify gaps between requirements and implementation, strengthen test coverage for edge cases, verify CI/CD compatibility, and resolve packaging, dependency, or environment issues. Using Git, Docker, pytest/JUnit, GitHub Actions, and language-specific tooling (Python, Java, JavaScript, TypeScript, C#, etc.), I'll ensure every task is deterministic, reproducible, and suitable for both automated evaluation and human review. I have experience with software architecture, AI-assisted development, Git, Docker, CI/CD, automated testing, debugging, benchmark creation, code review, and evaluating AI-generated code across multiple programming languages with a strong focus on correctness and maintainability. Will the tasks focus on specific languages/frameworks, or should I expect a mix of repositories? I'm ready to start immediately. Warm Regards, Aneesa.
$100 USD in 2 days
6.4
6.4

Hello Sir/MAM I am a Skilled Full Stack Developer. Having rich experience in Java , C++ , C , C# , Python , Eclipse , Sql , Mysql , .Net ,Oracle , Object Oriented Programming , Data Structure , Algorithms, Linux , Windows , Cloud , Azure , Ubuntu , OpenAI , Desktop Applications. Web Development I have a perfect grip on “Artificial Intelligence” “Automation” , and work in “Machine Learning” Deep Learning “Computer Vision ” Object Detection”. My track record as demonstrated in my 100% job completion and 5-star review rating showcases My ability to deliver exceptional results on time and with utmost quality I believe that my skill set makes me the ideal candidate for this project Please come on chat we will discuss more about this I will be waiting for your reply . Thanks and Best Regards
$140 USD in 1 day
6.5
6.5

Hi Senior AI/ML Engineer with 10+ years of experience delivering scalable, enterprise-grade AI and data solutions. Strong expertise in machine learning, generative AI, LLMs, RAG architectures, vector databases, and agentic AI workflows, combined with end-to-end data engineering (ETL/ELT pipelines, large-scale data processing). Proven track record of building and deploying production-ready AI platforms on AWS and Azure using Python, MLOps, and cloud-native architectures, with strong communication skills and effective collaboration across engineering, data, and product teams. Key Projects AI-Driven Healthcare Document Platform: Built HIPAA-compliant AI backends for secure medical document ingestion, OCR, classification, and agentic AI workflows AI-Powered Surveillance & Validation System: Developed microservices integrating YOLOv7 vision models, RAG-based AI querying, and event-driven workflows on AWS Manufacturing Spot-Weld Inspection System: Led backend AI development for real-time defect detection and optimized edge inference on Raspberry Pi Enterprise Financial Data Automation (SAP & Excel): Built AI-assisted pipelines to normalize General Ledger data, standardize Charts of Accounts, and automate Adjusted EBITDA analysis Deployed scalable, cloud-native AI services on AWS & Azure using Python, FastAPI, Django, Docker, and Kubernetes Thanks and regards
$140 USD in 7 days
8.4
8.4

Hello, We'd be excited to collaborate on building high-quality software engineering tasks for AI evaluation. This project aligns well with our experience in backend development, software architecture, testing, and reproducible engineering workflows. At Doomshell Software Pvt. Ltd., our senior engineers have extensive experience with Python, Java, Git, Docker, CI/CD, automated testing, and debugging across complex applications. We understand the importance of creating reliable, reproducible tasks with accurate instructions, comprehensive regression tests, and production-quality reference solutions. We can help with: Designing and validating software engineering tasks Creating regression tests and oracle solutions Dockerizing reproducible development environments Identifying gaps in test coverage and implementation Reviewing Git-based workflows, packaging, and metadata Delivering well-documented, review-ready engineering artifacts We value accuracy, maintainability, and long-term collaboration, and we'd be pleased to contribute to your evaluation framework. A few questions: Which programming languages and repositories will be prioritized initially? What review process do you follow for task acceptance? Is there an existing benchmark format or template we should follow? We look forward to discussing the opportunity. Best regards,
$240 USD in 7 days
6.3
6.3

Hi, I can help create and refine high-quality software engineering evaluation tasks by analyzing Git repositories, pull requests, and commits to produce reliable instructions, regression tests, oracle solutions, and reproducible Docker environments. I will focus on identifying edge cases, validating test coverage, debugging failures, and ensuring tasks meet strict review standards. My approach includes clean Git workflows, CI/CD validation, automated testing, and careful environment packaging so each benchmark remains accurate, repeatable, and suitable for evaluating coding agents. A few questions: Which programming languages and repository types are most commonly used for the evaluation tasks? Do you have an existing benchmark framework or validation pipeline that tasks must integrate with? How are oracle solutions and regression tests currently reviewed before being accepted? Best regards, Muhammad Usman
$145 USD in 7 days
5.8
5.8

Hello There! I’m Md Toriqul Islam, and I’m excited to partner with you. I can dive into your project immediately. I have rich experience in full-stack software engineering, software architecture, Docker, CI/CD, automated testing, and Git-based development. I understand you need a senior engineer to create and validate high-quality AI evaluation tasks based on real GitHub repositories, ensuring accurate specifications, reliable regression tests, reproducible Docker environments, correct oracle solutions, and complete metadata. I focus on delivering robust, well-documented, and review-ready engineering work. I am skilled in Git, Docker, CI/CD, Python, Node.js, Automated Testing, Debugging, and Software Architecture. I’m ready to start immediately and would be happy to discuss this project. Looking forward to hearing from you. Best regards, Md Toriqul Islam
$60 USD in 3 days
5.8
5.8

Hello, Subtle mismatches between instructions, tests, and oracles are the primary failure mode I see in AI evaluation tasks: tests that assume hidden behavior, oracles that diverge from repository intent, and Dockerfiles that don't reproduce the environment used by the original PR or commit. I have built and repaired benchmark tasks and evaluation suites for coding agents, including creating regression tests, correct oracle solutions, and reproducible Docker configurations. My background as a full-stack engineer and blockchain developer means I routinely handle multi-language repos, CI pipelines, and flaky test triage. Recent work involved fixing a multi-language benchmark where tests assumed nondeterministic network behavior; I rewrote tests to use deterministic fixtures, corrected the oracle, and containerized the runtime so CI reproduced local runs. I can reference specific examples in chat. I will start by running the repository's test matrix in an isolated Docker environment, capture deterministic failures, and map each failing test to either instructions, test code, oracle, or environment. My first concrete change will be a minimal reproducible Dockerfile and a focused regression test that fails before the fix and passes afterward. - Do you have an existing CI runner and artifact storage I should target, or should I propose a CI change? - Which language stacks and time zone overlap do you prefer for regular syncs? Happy to hop on a brief call whenever it suits you. demaxl
$250 USD in 3 days
5.8
5.8

Hi, Evaluating AI coding agents requires absolute rigor—if the Docker environment isn't strictly reproducible or the regression test suite has subtle gaps, the benchmark loses all validity. With over 15 years of experience in software architecture, systemic debugging, and build/environment optimization, I can deliver clean, airtight evaluation tasks derived from real-world GitHub repositories. Here is how I align with your workflow: • Deterministic Environments: Building isolated, reproducible Docker container setups to ensure zero environmental drift or non-deterministic behavior during execution. • Surgical Debugging & Oracle Verification: Catching subtle mismatches between task instructions and test assertions, validating oracle solutions, and writing robust regression tests across multiple languages (Python, Java, C++, Go). • Git & Metadata Integrity: Extracting precise problem statements from complex GitHub PRs and commits, pairing them with complete metadata and unambiguous specs. • Quality Control: Rigorous attention to detail to catch broken dependencies, flaky tests, and packaging defects before they hit review. I value long-term technical partnerships built on accuracy, clarity, and strict engineering standards. I am ready to review your workflow guidelines and start with an initial task to demonstrate the caliber of my work. Best regards, Duzy Chan ExtBit LLC
$222.22 USD in 7 days
5.7
5.7

Hello! We can provide an external engineer to help build and fix these AI evaluation tasks. 1. Are you open to working with an external contractor for this work? 2. Which part should we focus on first: instructions, tests, or Docker setup? — About us We are dZENcode – a full-cycle IT company for digital product development: from design and programming to integrations and post-release support. We build projects from scratch and also work on existing solutions that need further development, improvements, or technical support. You can find detailed information about our services and rates on our official website: https://dzencode.com. Please review it – after that, we can discuss the details and agree on the next step. ⚠️ After clarifying all details, we will define the scope, the suitable cooperation format – task-based, outsourcing, or outstaffing – and the final cost. Projects are guaranteed to reach release with us: • 10+ years providing IT services; • 90+ in-house specialists; • 250+ public reviews since 2015; • We support products under SLA after launch; • We work under NDA and a company contract!
$140 USD in 7 days
5.7
5.7

Leveraging 15+ years of experience in the technical realm, my profound understanding of software architecture and CI/CD principles enables me to excel in complex and meticulous projects like yours. I am well-versed in Git, Docker, and automated testing, which absolutely qualifies me for this task. My expertise extends not only to AI but into coding agents as well, thus I'm no stranger to benchmark development or adversarial test design. This experience has honed my ability to identify even the most subtle glitches, mismatches, incomplete solutions such as precisely those you indicated. I understand the importance of passing strict automated and human review, which has led me to be overly cautious regarding instructions, tests, test coverage testing phases. In conclusion, working together on this AI evaluation project wouldn't be an opportunity for me alone but a chance for us to establish a reliable long-term relationship. My commitment extends beyond delivering satisfactory tasks; it encompasses continued support post-delivery
$100 USD in 7 days
4.5
4.5

Hi there, Employer, Thank you for sharing this fascinating opportunity. I am excited by the challenge of developing and refining advanced software-engineering tasks for AI evaluation, and I believe my background aligns strongly with your requirements. With over a decade of experience in software architecture, CI/CD, Docker, and automated testing, I have successfully led and contributed to complex projects across Java, Python, and multi-language codebases. My work has included designing rigorous regression suites, building reproducible Docker environments, and ensuring end-to-end reliability through robust Git workflows. I am adept at debugging nuanced issues—from subtle instruction-test mismatches to environment reproducibility problems—and have a keen eye for detail that ensures nothing is overlooked. I have also been involved in AI benchmarking, adversarial test creation, and coding agent evaluation, giving me a solid understanding of what’s required to ensure high-quality, reliable evaluation tasks. My experience includes preparing precise instructions and metadata, validating oracle solutions, and maintaining strong test coverage to meet both automated and manual review standards. For your project, I propose a systematic approach: carefully analyzing each repository or PR, designing clear instructions, crafting comprehensive regression tests, building reproducible Docker environments, and verifying the solution pipeline. Every task will be validated for accuracy, reproducibility, and completeness. I am committed to technical excellence and clear communication, and I value building long-term, reliable collaborations. I would be delighted to discuss your workflow, project expectations, and how I can contribute to your success. Looking forward to connecting and exploring how we can work together. Best regards, DemiVision, LLC
$140 USD in 5 days
4.6
4.6

I've built and debugged AI evaluation tasks for coding agents before, using real GitHub repos, pull requests, and commits as sources. So this is right in my wheelhouse. I'll extract tasks from repos, design regression tests, implement oracle solutions, and containerize everything in Docker with strict reproducibility. CI/CD will handle validation across Python/Java targets. Debugging will include test-oracle mismatches, coverage gaps, and environment inconsistencies. My workflow aligns with your requirements: Git for versioning, Docker for isolation, CI/CD for automation, and adversarial test design for robustness. Reproducibility and human-review compliance are built in. I can start immediately. Thanks, Andrii.
$150 USD in 3 days
4.7
4.7

✋ Hi There!!! ✋ The Goal of the project:- BUILD RELIABLE AI EVALUATION TASKS WITH ADVANCED SOFTWARE ENGINEERING QUALITY. Experienced in developing and improving software systems with focus on accuracy, testing, and automation. • Creating and repairing coding evaluation tasks with proper instructions and regression tests. • Managing Git repositories, Docker environments, CI/CD workflows, and reproducible setups. • Debugging complex software issues across Python, Java, and multiple technologies. • Designing automated tests, validation workflows, and high-quality technical solutions. • Improving AI coding agent performance through realistic engineering challenges. <-- Questions --> 1. Which programming languages and repository types are mainly used for the evaluation tasks? 2. Are there existing benchmark guidelines or testing frameworks to follow? Completed similar projects involving AI automation, software testing frameworks, Docker deployments, code debugging, and backend engineering workflows with strong focus on reliability. Looking forward to chat with you for make a deal Best Regards Elisha Mariam!
$30 USD in 7 days
4.9
4.9

New York, United States
Member since Aug 2, 2026
₹12500-37500 INR
$15-25 USD / hour
₹2500-5000 INR
₹1500-12500 INR
₹750-1250 INR / hour
₹37500-75000 INR
₹100-400 INR / hour
$15-25 USD / hour
₹100-400 INR / hour
$8-15 USD / hour
$8-15 AUD / hour
$3000-5000 USD
₹750-1250 INR / hour
₹250000-500000 INR
£10-15 GBP / hour
$10-30 USD
$15-25 USD / hour
₹250000-500000 INR
$30-250 USD
min ₹2500 INR / hour