
Closed
Posted
Multimodal AI Task Designer AI Document-Generation Task Author Overview Design and build evaluation tasks that test large language models' ability to synthesize multi-source, visually-driven business and technical documents into polished professional deliverables (DOCX, PPTX, XLSX, or PDF). Each task is a full pipeline: sourcing, prompt writing, expert-authored reference output, grading criteria, and model testing, aimed at surfacing specific failure patterns in current AI systems. Key Responsibilities Source and curate 4+ real-world documents per task (reports, spreadsheets, slide decks, PDFs, engineering or spec sheets), including at least one distractor file designed to look useful but mislead the model, and at least one file where critical information exists only in a visual element, not extractable as text Write realistic, paragraph-style workplace prompts (no headers or rigid templates) that establish who is asking, why, what needs to be created, and which sources to use, without prescribing the answer's structure Author a Golden Response by hand (no AI assistance) that represents the correct, fully realized deliverable at the exact scope the prompt requires Write structured Grading Guidance covering ground truths, acceptable variation, penalizations, known failure modes, and aesthetic expectations Generate and refine an objective, binary, unstacked grading rubric from the Grading Guidance, including source citations and justifications for each criterion Run iterative model evaluations (Gemini 3.5 Flash) and tune task difficulty until the model reliably scores below a 45% average across three runs (all under 55%), confirming the task exposes a genuine model weakness Pass all required quality gates (Oracle run at 100%, Prompt/Criteria Alignment, Selected Run QC, Human QC) before submitting for Peer Review Revise tasks in response to Peer Review and HDM feedback through to final sign-off Skills Required Strong subject-matter expertise in one or more domains (Business/Consulting, Logistics, Engineering, Architecture, Front-End SWE, or STEM) Excellent technical writing and prompt design skills Sharp attention to detail for cross-referencing data across multiple source documents Working knowledge of DOCX/PPTX/XLSX/PDF authoring and formatting Understanding of how LLMs typically fail at multimodal and multi-document synthesis tasks, so you can design tasks that expose those gaps Compensation Paid per task upon sign-off, at a fixed rate equivalent to 6 hours of work Reviewer duties compensated at 1 hour equivalent Paid twice monthly
Project ID: 40621735
31 proposals
Remote project
Active 4 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
31 freelancers are bidding on average $13 USD/hour for this job

Hello, "Curated Multimodal Evaluation Tasks" - designing multimodal AI evaluation tasks I have built similar data‑curation pipelines, sourcing and cleaning multi‑source information for business leads (see my Polish dental clinics leads project: https://www.freelancer.com/projects/data-collection/Polish-Dental-Clinics-Business-Leads/reviews). This shows I can handle diverse file types and ensure clean, cross‑referenced data. I will also set up OCR extraction for visual‑only elements and create a binary grading rubric that aligns with your quality gates. Do you have preferred file formats for the source documents, and any existing templates for the golden responses? Looking forward to working with you. Artur Giżycki
$21 USD in 40 days
4.5
4.5

Hi there, I understand you need a system to design evaluation tasks that test LLMs on generating professional, multimodal business and technical documents from diverse inputs. I’ll build a structured task framework that assesses AI’s ability to integrate text, visuals, and data into polished DOCX, PPTX, XLSX, and PDF outputs. The approach includes: - Defining realistic business scenarios (e.g., investor pitch decks, technical reports, project plans) - Creating multi-source input sets combining text, diagrams, tables, and metadata - Designing prompt templates that guide model synthesis across formats - Establishing clear evaluation criteria: accuracy, coherence, visual alignment, formatting consistency - Delivering reusable task templates with instructions for automated or manual scoring I’ll ensure tasks are scalable, measurable, and aligned with real-world documentation workflows. All output will be documented in plain, actionable format ready for immediate use. Best Regards, Khorshed Alam, RS Software
$15 USD in 40 days
4.5
4.5

My approach would be: Step 1: Collect 4-6 authentic source documents (reports, spreadsheets, PDFs, presentations, specifications) and include a realistic distractor document plus files where key information exists only in charts, diagrams, or images. Step 2: Write a natural workplace scenario that explains the business context and expected deliverable without revealing the solution or forcing a template. Step 3: Manually create the reference DOCX/PPTX/XLSX/PDF by carefully combining information from all valid sources while intentionally ignoring misleading content. Step 4: Prepare grading guidance covering required facts, acceptable variations, common model mistakes, formatting expectations, and source references. Step 5: Convert the guidance into clear binary grading criteria so every point can be evaluated consistently. Step 6: Run evaluation rounds, analyze failure patterns, refine the task difficulty, and repeat until the target model consistently scores below the required threshold while keeping the task realistic. I pay close attention to detail, cross-check every source, and deliver structured, review-ready work. I would be happy to complete a small sample task to demonstrate my approach before we proceed further.
$8 USD in 40 days
4.1
4.1

The choice is between building the entire task pipeline from scratch or leveraging a pre-built framework, and I would pick building from scratch so I can precisely control the sourcing and prompt writing to surface specific failure patterns as requested. I will construct each task by sourcing 4+ real-world documents, including a distractor file and a file with critical visual information, then I will write prompts that target the model's ability to synthesize this multi-source data. For expert-authored reference output, I will use direct manipulation of common document creation tools like Microsoft Word and Excel, so the output has the correct structure and formatting. Grading criteria will be defined by concrete metrics for factual accuracy and adherence to output format. I would not create a general-purpose prompting system; that would not surface the specific failure patterns this job demands. I am a Preferred Freelancer on Freelancer with a 5.0 rating, 100% on time and 100% on budget. What format will the visual elements within documents typically take, images or charts, and does the AI need to interpret both equally? I need the specific output formats you require for the deliverables before I can provide a fixed quote.
$15 USD in 7 days
3.7
3.7

Hi, This role is a strong fit for someone who can turn messy, multi-source material into precise evaluation tasks. I understand you need polished document deliverables backed by real source files, a carefully written prompt, a hand-crafted gold response, and grading rubrics that expose model failures. I’ve worked on structured document writing, cross-referencing source material, and detail-heavy task design where accuracy and format matter just as much as content. I’m comfortable moving between PDFs, spreadsheets, slides, and specs, and I pay close attention to hidden details in visuals and distractor files. My approach is straightforward: verify the sources, write a natural workplace prompt, build a complete reference answer, and tighten the grading guidance until it is objective and consistent. I’m careful with scope, alignment, and consistency across each stage. If you’d like, I can help produce tasks that are clean, realistic, and ready for review. Best regards, Gabriel
$25 USD in 39 days
0.0
0.0

Hello, Thank you for posting your project. I'm with Jacob Construction LLC, a team specializing in architectural design, structural engineering, construction drafting, and permit-ready drawing packages for residential and commercial projects. We have extensive experience delivering accurate, code-compliant plans that help clients move smoothly from design through permitting and construction. After reviewing your project requirements, I am confident we can provide a practical, efficient, and high-quality solution tailored to your needs. We prioritize clear communication, attention to detail, and on-time delivery to ensure every project is completed to the highest professional standards. We are committed to maintaining close communication throughout the project and are happy to discuss your requirements in detail before getting started. We look forward to the opportunity to work with you and contribute to the success of your project. Best regards, Jacob Construction LLC
$12 USD in 5 days
0.0
0.0

I would love to chat about your project ⚠️THE WORST THAT CAN HAPPEN IS YOU WALK AWAY WITH A FREE CONSULTATION⚠️. At TashriqueTech, we work with a select group of clients to provide focused support and better results. Our multimodal AI evaluation task creation services are tailored to your specific needs. We design comprehensive evaluation tasks that rigorously test large language models' capabilities in synthesizing complex documents. From sourcing real-world documents to crafting realistic workplace prompts, we ensure each task is meticulously outlined. Our process includes developing expert-authored reference outputs, grading criteria, and iterative model evaluations to uncover critical weaknesses in AI systems. IF YOU'RE NOT HAPPY YOU DON'T PAY. So, feel free to message me. Check out our portfolio for examples of successful projects that highlight our ability to deliver precise, high-quality results.
$8 USD in 7 days
0.0
0.0

I will design and deliver full multimodal evaluation tasks at the stated per-task rate within your 8 to 15 USD band. I have produced 18 hand authored multimodal LLM evaluation tasks including human Golden Responses and objective grading rubrics. Each task I deliver will include sourcing and curating four or more real world documents with at least one distractor file and one visual only source, a paragraph style workplace prompt, a hand written Golden Response, detailed Grading Guidance covering ground truths, acceptable variation, penalizations and aesthetic expectations, and a binary unstacked rubric with source citations and justifications. I will run iterative Gemini 3.5 Flash evaluations and tune difficulty until average scores fall below 45 percent across three runs, pass the Oracle run, Prompt Criteria Alignment, Selected Run QC and Human QC, and revise through Peer Review and HDM to final sign off. Which domain should I prioritize first from your list: Business, Logistics, Engineering, Architecture, Front End, or STEM? If you share one sample source set I will return a complete task seed within 48 hours for review, free. Ali Zain
$11.50 USD in 7 days
0.0
0.0

⭐⭐⭐⭐⭐ Design Effective AI Evaluation Tasks for Multimodal Document Generation ❇️ Hi My Friend, I hope you're doing well. I've reviewed your project requirements and see you're looking for a Multimodal AI Task Designer. You don't need to look any further; Zohaib is here to help you! My team has completed 50+ similar projects for AI document generation tasks. I will create tasks that test AI's ability to synthesize multi-source documents, ensuring high-quality deliverables. ➡️ Why Me? I can easily design your evaluation tasks as I have 5 years of experience in technical writing, prompt design, and document synthesis. My skills include data curation, grading rubric creation, and iterative model evaluation. Additionally, I have a strong grip on AI systems and document formatting, ensuring a thorough approach to your project. ➡️ Let's have a quick chat to discuss your project in detail and let me show you samples of my previous work. I'm eager to explore how I can bring value to your project. ➡️ Skills & Experience: ✅ Technical Writing ✅ Prompt Design ✅ Document Curation ✅ Grading Rubrics ✅ Model Testing ✅ Attention to Detail ✅ DOCX Authoring ✅ PPTX Formatting ✅ XLSX Management ✅ PDF Creation ✅ AI Evaluation ✅ STEM Expertise Waiting for your response! Best Regards, Zohaib
$9 USD in 40 days
0.0
0.0

With my extensive background in AI, automation, and digital marketing, I am confident that I am the ideal candidate for your Multimodal AI Evaluation Task Creation. I have been working with AI-driven systems for over a decade, specializing in leveraging AI for lead generation and business growth. This experience has given me a deep understanding of how language models can fail at tasks involving multi-source, visually-driven document synthesis, making me well-equipped to design tasks that highlight these shortcomings. Additionally, my proficiency in technical writing and prompt design aligns perfectly with the responsibilities of this project. I have an eye for detail when it comes to cross-referencing data across multiple sources, which will be crucial in ensuring the tasks are challenging and effectively evaluate the language models' capabilities. Lastly, my broad subject-matter expertise related to business and engineering inclines with your outlined project domains. I have a firm grasp on creating compelling content in various contexts and believe that my ability to think outside rigid templates will lend itself well to constructing realistic prompts that provide meaningful information without giving away the structure of the correct response. Rest assured, with me on board, you'll receive meticulously created tasks that thoroughly evaluate multimodal AI capabilities and consequently improve the overall quality of these models.
$8 USD in 40 days
0.0
0.0

We recently helped a tech company design evaluation tasks that tested the capabilities of large language models, leading to significant insights into AI performance. We will assist in creating comprehensive evaluation tasks for multimodal AI, focusing on synthesizing visually-driven documents into polished deliverables. Your project emphasizes the need for "clean, professional, user-friendly" tasks that reveal specific failure patterns in AI systems. This understanding is crucial for achieving your goals. We have strong expertise in technical writing, prompt design, and document formatting, ensuring a meticulous approach. With 75+ 5-star reviews on similar projects, we rank in the top 1% among 75 million users. It would be our privilege to contribute to your project, and choosing us is a decision you will not regret. Regards, Henco Burger.
$11 USD in 3 days
0.0
0.0

40 hours/week, available for work You can track project progress via the tracker Hi! I specialize in Generative AI, LLM evaluation, and document intelligence, with hands-on experience designing complex AI workflows, RAG systems, and prompt evaluation frameworks. I can create challenging multimodal tasks that accurately measure model reasoning, document synthesis, and instruction-following while exposing real-world failure modes. I can assist you with: 1. Task Design Curate multi-source datasets with distractor and visual-only evidence Write realistic workplace prompts for business and technical scenarios Design tasks that reveal reasoning, retrieval, and synthesis weaknesses 2. Evaluation & Documentation Author expert Golden Responses without AI assistance Create detailed grading guidance, binary rubrics, and scoring criteria 3. Model Analysis Evaluate LLM outputs and identify recurring failure patterns Refine prompts to achieve target difficulty and evaluation consistency My background includes building enterprise AI applications, document-processing pipelines, RAG systems, and evaluation workflows using Python, LLM APIs, and modern AI frameworks. I enjoy designing rigorous benchmarks that measure practical AI capabilities rather than simple prompt-following. I'd be happy to discuss your evaluation framework, quality standards, and target domains to create high-quality multimodal tasks that consistently expose meaningful model limitations. Best regards, Prateek
$10 USD in 40 days
0.0
0.0

Hello, I can help create high-quality multimodal AI evaluation tasks that accurately test large language models across real-world business and technical scenarios. I have 8+ years of experience in technical writing, business analysis, documentation, prompt writing, and AI-focused content development. I specialize in creating structured documentation, detailed prompts, evaluation criteria, and well-organized professional deliverables. I'll develop realistic prompts, organize multi-source documents, prepare accurate reference responses, create objective grading rubrics, and ensure every task meets the required quality standards with strong attention to detail. I focus on accuracy, logical reasoning, clear documentation, and timely delivery. I'm confident I can contribute high-quality evaluation tasks and collaborate effectively throughout the review process. I look forward to discussing your project. Best Regards, Manmohan
$8 USD in 5 days
0.0
0.0

Hi, I have experience designing AI workflows, creating high-quality technical documentation, and developing evaluation datasets for LLM-powered applications. I can create realistic multimodal tasks, curate diverse source documents, write natural business prompts, prepare detailed golden responses, develop objective grading rubrics, and analyze model failure patterns. I'm proficient with DOCX, PPTX, XLSX, PDF authoring, prompt engineering, and structured quality evaluation. My focus is on producing well-documented, challenging tasks that expose reasoning and synthesis limitations while meeting strict quality standards. I am available for long-term collaboration and can consistently deliver high-quality tasks within your review process.
$12 USD in 40 days
0.0
0.0

Hi, I have built my own RAG systems and evaluation metrics MCPs, I know the limitations of LLMs I have also built an SLM, worked as a product manager at AI automation consulting firm, created and deployed a product that was mentioned in silicon India, I also hold an MBA from a Tier-1 institute, I am the perfect candidate, would love to contribute more to this project
$12 USD in 50 days
0.0
0.0

Hey, If you're not happy, you don't pay. As a skilled front-end professional with extensive experience in multimodal AI task design and evaluation, I am confident in my ability to create comprehensive tasks that test large language models' capacity to synthesize multi-source, visually-driven business and technical documents into polished professional deliverables. I'd love to share similar work I've done in my portfolio. Best, Kyle
$11.25 USD in 7 days
0.0
0.0

My experience working with complex financial data management and risk assessment, matched with my knowledge of how LLMs typically fail at multimodal tasks, places me in a unique position to design tasks that expose the gaps. Additionally, my proficiency in Excel-based reporting and data analysis will be invaluable as I generate and refine an objective unstacked grading rubric for your tasks. I'm very comfortable with DOCX/PPTX/XLSX/PDF formatting, which enables me to author technical documents to the highest professional standards. My strong subject-matter expertise across diverse domains such as finance, operations, technology and content creation is a huge asset to this project ⁰
$12 USD in 40 days
0.0
0.0

Hello, I have reviewed your project and would be happy to help. With 5+ years of experience, I can deliver high-quality work with accuracy and on-time delivery. ✔ Quality Work ✔ Fast Communication ✔ Unlimited Revisions ✔ 100% Client Satisfaction I look forward to working with you. Best regards
$8 USD in 40 days
0.0
0.0

I am a skilled researcher, technical writer, and AI project contributor with experience analyzing complex information and producing clear, professional documentation. I have strong attention to detail, allowing me to synthesize data from multiple sources, identify inconsistencies, and create accurate deliverables. My background includes business analysis, content development, prompt writing, and quality evaluation of AI-generated outputs. I am proficient in working with reports, spreadsheets, presentations, and technical documents while maintaining high standards of accuracy and clarity. I am excited about the opportunity to design challenging evaluation tasks that help improve AI performance and reliability across real-world business and technical scenarios.
$12 USD in 40 days
0.0
0.0

I am civil engineer with all knowledge which fits for your requirements.i also know autocad bar bending schedule and estimation
$12 USD in 40 days
0.0
0.0

Atlanta, United States
Member since Aug 14, 2024
$30-250 USD
$750-1500 USD
$30-250 USD
₹75000-150000 INR
₹600-1500 INR
₹750-1250 INR / hour
₹3500-4000 INR / hour
₹750-1250 INR / hour
₹12500-37500 INR
₹750-1250 INR / hour
min $50 USD / hour
$30-250 USD
£20-250 GBP
$250-750 USD
$250-750 USD
₹600-1500 INR / hour
$10-30 USD
£20-250 GBP
$30-250 USD
$30-250 USD
₹600-700 INR
$1500-3000 SGD
£250-750 GBP