Filter

My recent searches
Filter by:
Budget
to
to
to
Type
Skills
Languages
    Job State
    391 multimodal jobs found

    ...legitimately remain on screen for seconds while others may remain for several minutes; - the slide thumbnails in the existing review interface are currently failing to load; - the review step exists, but too many of the AI recommendations need correction. I do NOT want a simple improvement to the equal-duration heuristic or an LLM guessing timestamps. I believe this is closer to a monotonic multimodal sequence-alignment problem. A potential architecture would involve: PowerPoint slide text / notes / visual concepts + WhisperX or equivalent word-level transcription / forced alignment + semantic matching between transcript sections and slides + high-confidence anchor detection + ordered/monotonic sequence alignment across the full deck + natural sentence / pause / silence bound...

    $34 / hr Average bid
    $34 / hr Avg Bid
    130 bids

    ...existing embodied-AI stack and now want to double-down on Vision-Language-Action modelling. The core goal is to design, train and scale a full pipeline that takes raw multimodal data, learns a joint representation and closes the loop all the way to real-time action on a physical robot. You will start from large, messy datasets (images, video clips, proprioception, language annotations) that already sit on our cluster. The job is to craft a new architecture in PyTorch, schedule distributed training, and iterate until the model achieves reliable closed-loop visuomotor reasoning in simulation and on hardware. Robust multimodal representation learning, a world-model/JEPA component, temporal memory, and predictive control need to come together in a single, maintainable codebas...

    $168 Average bid
    $168 Avg Bid
    61 bids

    ...technical challenges. Discuss real-world deployment problems. Explain accuracy, robustness, usability, and speech limitations. D. RQ4 – Emotional Intelligence Explain emotion recognition. Discuss emotion-aware speech generation. Describe models and techniques used. Explain how emotional/prosodic information improves interaction. E. RQ5 – Emerging Technologies Identify future technologies. Discuss multimodal learning. Discuss LLMs. Discuss Edge AI. Discuss affective computing. Discuss advanced computer vision and expressive TTS. Explain future research opportunities. 8. Threats to Validity Discuss limitations of the research methodology. Identify dataset limitations. Discuss evaluation inconsistencies. Explain generalizability issues. Mention possible selection/samplin...

    $77 Average bid
    $77 Avg Bid
    10 bids

    We are building an automated evaluation engine for handwritten student exam papers. We need an experienced AI/Full-Stack engineer to develop an end-to-end pipeline in Node.js that accepts scanned handwritten PDFs, performs OCR/HTR to map word-level coordinates, evaluates the content using Multimodal Vision LLMs, and programmatically renders high-precision visual annotations, margin callouts, checkmarks, and pointer arrows directly onto the output PDF. The Result I want is attached Below in the PDF. This is what I want to achieve.

    $264 Average bid
    $264 Avg Bid
    41 bids

    I want to build an AI-driven system whose sole focus is content creation. I am still flexible on the exact architecture—whether the core ends up being an NLP model that writes long-form text, a generative image engine, or a multimodal blend, the end goal is the same: reliable, original content produced on demand with minimal human intervention. Here is what I need from you: • A clear proposal outlining the model stack you would use (e.g. GPT-4, Llama 2, Stable Diffusion, or your own fine-tuned variant). • A small proof-of-concept that ingests a prompt and returns publish-ready output. • Documentation describing setup, retraining, and how I can scale the system once the POC is approved. Acceptance criteria: the demo must run locally or on a standard cloud n...

    $9 / hr Average bid
    $9 / hr Avg Bid
    24 bids

    ...recognition, guidance, and feedback channels together so the experience is truly usable. Core flow • Real-time positioning: use camera, GPS, IMU and, where available, Wi-Fi/BLE beacons to pinpoint the user’s location and heading, whether in a shopping mall or on a city street. • Path planning & turn-by-turn guidance: generate concise, context-aware instructions that avoid cognitive overload. • Multimodal feedback: combine short audio prompts, optional bone-conduction output, and vibration/haptic cues to overcome noisy environments. • Robust obstacle awareness: fetch depth or stereo data (or infer depth with AI) so the user is warned about hazards at cane-length distance or overhead. • Extensible add-ons: hooks for voice commands, emer...

    $52 / hr Average bid
    $52 / hr Avg Bid
    87 bids

    ...Package,Mid X,Mid Y,Rotation,Layer U1,nRF5340,aQFN-94,12.5000,24.1000,90.0,Top U2,ADPD4200,WLCSP-36,15.2000,8.4000,0.0,Bottom U3,AD5940,WLCSP-56,8.1000,10.2000,270.0,Top D1,760nm LED,0603,16.5000,8.4000,180.0,Bottom bom: Designator,Quantity,Value,Footprint,Manufacturer Part Number,Description U1,1,nRF5340,aQFN-94,nRF5340-CLAA-R,Dual-Core BLE Microcontroller U2,1,ADPD4200,WLCSP-36,ADPD4200BCBZR7,Multimodal Optical AFE U3,1,AD5940,WLCSP-56,AD5940BCBZ-RL7,Bioimpedance AFE (Hydration) U4,1,CALERA,SMD-Module,gSKIN-CALERA,Heat Flux Core Temp Sensor U5,1,BMI270,LGA-14,BMI270,6-Axis IMU (Motion/Steps) U6,1,MAX20303,WLP-56,MAX20303BEWN+T,Wearable Power Management IC U7,1,W25Q128,WSON-8,W25Q128JVSNIQ,128Mb Flash Memory ANT1,1,2.4GHz Ant,1206,2450AT18A100E,Ceramic Chip Antenna D1,2,760nm LE...

    $237 Average bid
    $237 Avg Bid
    17 bids

    ...We are particularly interested in: AI / ML Engineers -Strong expertise in AI/ML, Deep Learning, Computer Vision, Generative AI, Reinforcement Learning or Autonomous Systems -Experience with Python, PyTorch/TensorFlow and modern AI frameworks -Strong understanding of algorithms, model training, optimization and deployment -Experience building AI systems from scratch is highly preferred -LLMs, multimodal AI, agents, robotics AI or edge AI experience is a strong advantage Aerospace / Aeronautical / Drone Engineers -Strong fundamentals in aerodynamics, flight mechanics, propulsion, avionics or UAV systems -Hands-on experience designing, building, testing or modifying drones/UAVs -Experience with flight controllers, sensors, embedded systems, telemetry and autonomous flight -CAD, ...

    $325 Average bid
    $325 Avg Bid
    20 bids

    ...normal breathing and double triggering. 2. **Synthetic dataset creation** with waveform signals, images, labels and physiological parameters. 3. Addition of realistic **noise, artefacts and physiological variability**, followed by diffusion-based augmentation. 4. Development of AI models using **time-series and waveform images**, including Transformers, CNNs, Bi-LSTM and Swin Transformer. 5. **Multimodal fusion** of signal and image features for PVA classification. 6. **Synthetic-to-real domain adaptation** using DANN, CORAL and MMD. 7. **Explainable AI (XAI)** using Grad-CAM, Integrated Gradients, Score-CAM and Attention Rollout, with comparison against expert annotations. 8. **Temporal modelling and uncertainty estimation** to improve clinical reliability. 9. Eventually, develo...

    $88 Average bid
    $88 Avg Bid
    39 bids

    I’m running a systematic investigation into how prompt phrasing influences in-context regression performance. The work spans six separate axes, and for this engagement I want you to concentrate on two of them—Data preprocessing and Evaluation metrics—while keeping the other axes in mind so our findings remain extensible. The core set-up centres on text-based prompts only; no visual or multimodal inputs will appear in this round. All experiments will target polynomial regression tasks, so the prompts, data splits, and metric choices should reflect the non-linear nature of the underlying relationships. Here is what I need from you: • Curate or generate a clean, well-documented dataset suitable for polynomial regression, then outline the preprocessing steps yo...

    $5 / hr Average bid
    $5 / hr Avg Bid
    46 bids

    Multimodal AI Task Designer AI Document-Generation Task Author Overview Design and build evaluation tasks that test large language models' ability to synthesize multi-source, visually-driven business and technical documents into polished professional deliverables (DOCX, PPTX, XLSX, or PDF). Each task is a full pipeline: sourcing, prompt writing, expert-authored reference output, grading criteria, and model testing, aimed at surfacing specific failure patterns in current AI systems. Key Responsibilities Source and curate 4+ real-world documents per task (reports, spreadsheets, slide decks, PDFs, engineering or spec sheets), including at least one distractor file designed to look useful but mislead the model, and at least one file where critical information exists only in a v...

    $13 / hr Average bid
    $13 / hr Avg Bid
    30 bids

    I’m building a deep-learning system that can look at two very different kinds of information—gene-level data and free-form clinical text—and fuse them to predict whether a patient is likely to develop a target disease. The heart of the job is a multimodal architecture that treats gene features and textual features as complementary signals, learns their joint representation, and outputs a binary (disease / no-disease) or probabilistic risk score. Here is what I need from you. First, a clean data-pipeline that ingests my gene expression matrices alongside the associated clinical notes, handles any necessary tokenisation or normalisation, and keeps sample alignment intact. Second, a well-documented model—PyTorch or TensorFlow is fine—that includes separat...

    $203 Average bid
    $203 Avg Bid
    63 bids

    ...Within each platform And I need to build my websites using ai So I am looking for a genuine teacher from any of these platforms Please outline your expertise in either of these platforms in ur reply Do not waste your time. Be PRECISE OK? Scope • Cover the full spectrum of modern AI: classic machine-learning workflows, natural-language processing, computer vision, and emerging multimodal techniques. • Wherever a concrete example fits—image classification, text generation, or small end-to-end projects—walk through the code step by step. Feel free to showcase TensorFlow, PyTorch, Keras, or any other mainstream library if it helps keep explanations clear and practical. • Position every lesson for self-paced study: concise theory, a runna...

    $6 / hr Average bid
    $6 / hr Avg Bid
    68 bids

    Project Description: Project "OmniHub" Cross-Domain Multimodal AI Consulting Ecosystem with Self-Evolving State 1. Project Overview Project OmniHub is an enterprise-grade, multimodal AI consulting hub designed to deliver expert-level, specialized assistance across various domains including Medicine, Law, Agriculture, Marketing, and Education. Unlike standard chat applications, OmniHub features a fully persistent memory tier (tracking preferences, policies, and facts across sessions) and an autonomous self-learning mechanism. The system processes multimodal inputs (Voice, Images, and Video) and orchestrates specialized domain agents via a unified context framework. 2. Core Operational Pillars The platform adapts its operational logic depending on the chosen ...

    $335 Average bid
    $335 Avg Bid
    132 bids

    ...Language Requirement Engagement Type Live social commerce marketplace for the GCC region, featuring interactive real-time shopping via live video, high-concurrency streaming, and integrated multimodal AI pipelines Working hours: Align with GCC Business hours Experience: 7+ yr We are seeking a highly skilled and experienced Senior DevOps / Infrastructure Engineer to architect, scale, and maintain a next-generation live social commerce marketplace optimized for the GCC region. This cutting edge platform integrates high concurrency real time video streaming, live interactive shopping, and robust multimodal AI pipelines (including vision language models and text-to-speech technologies). In this role, you will bridge the gap between complex cloud architectures and reliable o...

    $2356 Average bid
    $2356 Avg Bid
    28 bids

    I’m looking to run a complete multimodal analysis of a body of spoken content that arrives in three different forms: raw audio, corresponding video, and existing text transcripts. My aim is to fuse these streams so we can extract deeper insights than any one modality would reveal on its own. Here’s what I need from you: • Develop or adapt models that synchronise audio, video, and text, producing aligned outputs ready for further study. • Apply speech-to-text (where transcripts are missing or need verification), gauge recognition accuracy, and layer in sentiment or other paralinguistic signals as appropriate. • Package the workflow in reproducible Python notebooks or scripts (PyTorch/TensorFlow, Whisper, Kaldi, or similar libraries are fine), with ...

    $143 Average bid
    $143 Avg Bid
    43 bids

    I want to run my own ChatGPT-style interface—similar to OpenWebUI—on hardware that I own. The backend model is already in place; what I need is the full application layer that lets people sign up, chat, and, when desired, execute small code snippets inside an isolated sandbox. Key workflow • Users reach a multimodal chat UI served from my server. They can paste or drag documents, images, or plain text in-line, and the system will pass them straight to the model. • An embedding side-car should index any uploaded documents so that later queries can reference them. • Each signed-in user gets a throw-away container (Docker or similar) for the “MCP” code sandbox; the container auto-recycles after an extended idle period so disk usage stays...

    $316 Average bid
    $316 Avg Bid
    79 bids

    ...assessment • Speech-to-text validation • Voice data review • Accent evaluation • Dialect evaluation • Pronunciation evaluation • Audio segmentation • Audio classification • Audio prompt evaluation • Image annotation • Image classification • Visual reasoning evaluation • OCR validation • Image quality review • Object detection review • Bounding box validation • Video annotation • Video classification • Multimodal AI evaluation • Safety and policy review • Instruction-following evaluation • Localization review • Translation review • Human preference ranking • Data quality audits • AI training projects • AI fine-tuning support projects • Any other AI...

    $5 / hr Average bid
    $5 / hr Avg Bid
    39 bids

    We are looking for an AI specialist to join our team. We are developing several AI tools with Google. We are currently developing a tool featuring a virtual AI assistant that makes calls. We are using Google's Gemini Live API and Google's native multimodal models Gemini Live 2.5 Flash Native Audio and Gemini 3.1 Flask Live preview

    $11 - $18 / hr
    Urgent Sealed
    $11 - $18 / hr
    101 bids

    ...Blazor, DigitalOcean, PostgreSQL) that publishes video simultaneously to TikTok, YouTube, Facebook, Instagram, LinkedIn, and X, with AI-generated captions via Google's Gemini multimodal API. The core of this role is advanced API integration engineering: Social platform APIs at production depth: Meta Graph API (Business Pages), YouTube Data API, TikTok Direct Post, LinkedIn, X — OAuth token lifecycles and refresh, per-platform rate limits and quotas, upload protocols, webhook/status handling, and per-platform failure isolation (one platform's failure must never block another) AI APIs in production: Google Gemini multimodal (video in, structured JSON out) — Files API lifecycle, generation config and thinking models, response parsing that survives struc...

    $27 / hr Average bid
    $27 / hr Avg Bid
    288 bids

    Project Overview SoftAge AI is recruiting participants for Project ROBO 2, an AI data collection initiative focused on capturing first-person point-of-view (egocentric) videos in home environments. The collected data will support the development of Robotics AI, Embodied AI, Human Action Understanding, and Multimodal AI systems.   Compensation * $7-$25 per hour of approved recording   Supported Devices Participants must use one of the following devices: * iPhone 11 or newer * Google Pixel 6 or newer * Samsung Galaxy S21+ models only   Recording Requirements * The phone must be securely mounted on the forehead using a head strap. * Recordings must be captured from a first-person POV (egocentric perspective). * The camera angle should be approximately 45° downward. *...

    $10 / hr Average bid
    $10 / hr Avg Bid
    2 bids

    Project Overview SoftAge AI is recruiting participants for Project ROBO 2, an AI data collection initiative focused on capturing first-person point-of-view (egocentric) videos in home environments. The collected data will support the development of Robotics AI, Embodied AI, Human Action Understanding, and Multimodal AI systems. Compensation $7-25 per hour of approved recording Supported Devices Participants must use one of the following devices: iPhone 11 or newer Google Pixel 6 or newer Samsung Galaxy S21+ models only Recording Requirements The phone must be securely mounted on the forehead using a head strap. Recordings must be captured from a first-person POV (egocentric perspective). The camera angle should be approximately 45° downward. Hands must remain visible throu...

    $12 / hr Average bid
    $12 / hr Avg Bid
    7 bids

    I need an AI engineer who can architect, train, and iterate on deep-learning models that perform both medical-imaging analysis and diagnostic support. The scope covers X-ray, MRI, and CT data, so you should be comfortable handling multimodal image pipelines and the differing pre-processing each modality demands. You will start from a clean slate: selecting or designing network architectures in Python, building them with PyTorch or TensorFlow, and setting up a repeatable training environment that lets us experiment rapidly. Once a strong baseline is in place, I want to see steady, research-driven improvements—new loss functions, data-augmentation ideas, self-supervised techniques, or anything that reliably drives accuracy upward while keeping the models clinically robust. Dep...

    $164 Average bid
    $164 Avg Bid
    92 bids

    We are developing a portable, handheld research instrument that combines two measurement modalities in a single pen-form device: 1. ELECTRICAL IMPEDANCE SPECTROSCOPY (EIS) 4-wire (tetrapolar) configuration. Frequency sweep 100mHz to 200kHz. Phase angle resolution ≤ 0.1°. Impedance range 10Ω to 10MΩ. Precision analog front-end IC (SPI interface). Shielded cable connector for external probe (Lemo or circular DIN, IP44). 2. OPTICAL SPECTROSCOPY Hamamatsu C12880MA MEMS spectrometer (SPI). Multi-wavelength LED driver: 4 channels (405 / 530 / 625 / 850 nm), sequential, PWM-controlled, current-regulated ±1%. SMA905 fiber optic connector. This project centres on a single-board solution that brings electrical impedance spectroscopy and optical spectroscopy togethe...

    $16 Average bid
    $16 Avg Bid
    7 bids

    We are developing a portable, handheld research instrument that combines two measurement modalities in a single pen-form device: 1. ELECTRICAL IMPEDANCE SPECTROSCOPY (EIS) 4-wire (tetrapolar) configuration. Frequency sweep 100mHz to 200kHz. Phase angle resolution ≤ 0.1°. Impedance range 10Ω to 10MΩ. Precision analog front-end IC (SPI interface). Shielded cable connector for external probe (Lemo or circular DIN, IP44). 2. OPTICAL SPECTROSCOPY Hamamatsu C12880MA MEMS spectrometer (SPI). Multi-wavelength LED driver: 4 channels (405 / 530 / 625 / 850 nm), sequential, PWM-controlled, current-regulated ±1%. SMA905 fiber optic connector. This project centres on a single-board solution that brings electrical impedance spectroscopy and optical spectroscopy togethe...

    $13 / hr Average bid
    $13 / hr Avg Bid
    7 bids

    We are developing a portable, handheld research instrument that combines two measurement modalities in a single pen-form device: 1. ELECTRICAL IMPEDANCE SPECTROSCOPY (EIS) 4-wire (tetrapolar) configuration. Frequency sweep 100mHz to 200kHz. Phase angle resolution ≤ 0.1°. Impedance range 10Ω to 10MΩ. Precision analog front-end IC (SPI interface). Shielded cable connector for external probe (Lemo or circular DIN, IP44). 2. OPTICAL SPECTROSCOPY Hamamatsu C12880MA MEMS spectrometer (SPI). Multi-wavelength LED driver: 4 channels (405 / 530 / 625 / 850 nm), sequential, PWM-controlled, current-regulated ±1%. SMA905 fiber optic connector. This project centres on a single-board solution that brings electrical impedance spectroscopy and optical spectroscopy togethe...

    $700 Average bid
    $700 Avg Bid
    37 bids

    ...(such as DARPA in the US or CERN in Europe), and private corporations like Google, Meta, Huawei, and Samsung. Funding often comes from a mix of public grants, venture capital, and corporate budgets. Key Areas of Breakthrough Technology Research 1. Artificial Intelligence and Machine Learning AI research has exploded since the deep learning revolution of the 2010s. Current frontiers include: Multimodal models that understand text, images, video, and audio simultaneously. AI agents capable of autonomous planning and tool use. Explainable AI (XAI) to make black-box systems more transparent and trustworthy. Energy-efficient AI to address the massive computational demands of training large models. 2. Quantum Computing and Information Science Quantum research aims to harness quantum ...

    $21 / hr Average bid
    $21 / hr Avg Bid
    16 bids

    Full-Stack Engineer Tech stack (must-have) Frontend: React 18, TypeScript (strict), Vite, Mantine 8, CSS Modules, React Router 7 Backend (Node): Node.js, TypeScript, Medplum bots (FHIR vmcontext runtime) Backend (Python): FastAPI, httpx, async Python Data: FHIR R4 (Composition, Encounter, Invoice, Observation, Condition, etc.), PostgreSQL, Redis AI: Google Gemini 2.5 API (Pro multimodal + Flash structured outputs), Anthropic Claude API Infra: Docker, nginx, GitHub Actions Integrations: Stripe, or Jitsi (telemedicine), Stedi,

    $524 Average bid
    $524 Avg Bid
    381 bids

    PROJECT: Refurnish my BTech/MTech research paper to match the exact structure, tone, and detail level of the reference PDF I provide. Topic: Causal Multimodal Diagnostic Agent Combining Chest X-ray Images + Clinical Reports for Thoracic Diseases WHAT I PROVIDE: 1. Reference Paper PDF: You MUST replicate its section headings, table formats, equation style, and writing tone 100%. 2. My Code: , , - ResNet-50 + ClinicalBERT + cross-modal attention + label GNN 3. My Results: with AUROC/F1 per class, , , 4. Dataset Info: MIMIC-CXR subset, 1,495 PA-view image-report pairs, 14 CheXpert labels, patient-level 70/10/20 split. NO full 500GB dataset needed. 5. Target: BTech/MTech final submission + college

    $61 Average bid
    $61 Avg Bid
    8 bids

    Temporal Lesion-Aware Dynamic Gated Multimodal Fusion Framework for DR and DME Analysis Using OLIVES and MMRDR Datasets Framework Overview The proposed framework introduces a Temporal Lesion-Aware Dynamic Gated Multimodal Fusion System for automated analysis of Diabetic Retinopathy (DR) and Diabetic Macular Edema (DME) using multimodal retinal imaging data. The framework combines fundus images, OCT scans, longitudinal retinal information, and optional clinical metadata to improve retinal disease classification, biomarker understanding, and temporal disease progression analysis. Unlike conventional multimodal retinal systems that use static feature fusion, the proposed framework employs a: Dynamic Gated Cross-Modal Fusion Mechanism that adaptively learns the impo...

    $128 Average bid
    $128 Avg Bid
    12 bids

    ...photos/content and the AI: Understands the content Selects template Generates: headlines captions hooks layouts Pushes into Canva drafts ready for approval Canva Design System Needed Please also help create: 5 reel cover styles 5 before/after systems 2 educational carousel systems Fixed typography system Fixed colour palette Reusable templates Preferred Skills Canva automation AI workflows Gemini / multimodal AI Prompt engineering No-code automation Luxury brand/social media design. Instagram Reference:

    $16 / hr Average bid
    NDA
    $16 / hr Avg Bid
    36 bids

    I am putting together a small educational AI/ML project that detects diabetic retinopathy on the publicly-available OLIVES dataset. The goal is not only to build a working multimodal model (fundus images plus any supporting clinical metadata you find useful) but also to showcase clear explainability and rigorous evaluation so the project can be presented in an academic setting. Scope of work • Prepare the OLIVES dataset, handle any class imbalance, and document the preprocessing pipeline. • Design and train a multimodal architecture of your choice in Python—PyTorch, TensorFlow or another modern framework is fine—as long as the code is clean and reproducible. • Produce the quantitative metrics I need: Accuracy, Precision, F1-score, AUC and Coh...

    $46 Average bid
    $46 Avg Bid
    15 bids

    We are building a high-performance team of elite AI annotation and evaluation professionals for advanced AI training and multimodal evaluation projects. This is NOT basic click-task annotation work. We are specifically looking for highly analytical, detail-oriented professionals capable of evaluating complex AI-generated outputs across domains such as: * software engineering * UX/UI and visual design * computer vision * multimodal AI * spreadsheets and documents * presentations and structured data * reasoning and ranking workflows Compensation: • Competitive hourly compensation (high-performing contributors may earn substantial weekly income) • Flexible remote work • Ongoing project opportunities What You Will Do: * Evaluate AI-generated outputs using str...

    $14 / hr Average bid
    $14 / hr Avg Bid
    5 bids

    We are building a high-performance team of elite AI annotation and evaluation professionals for advanced AI training and multimodal evaluation projects. This is NOT basic click-task annotation work. We are specifically looking for highly analytical, detail-oriented professionals capable of evaluating complex AI-generated outputs across domains such as: * software engineering * UX/UI and visual design * computer vision * multimodal AI * spreadsheets and documents * presentations and structured data * reasoning and ranking workflows Compensation: • Competitive hourly compensation (high-performing contributors may earn substantial weekly income) • Flexible remote work • Ongoing project opportunities What You Will Do: * Evaluate AI-generated outputs using str...

    $11 / hr Average bid
    $11 / hr Avg Bid
    11 bids

    ...quality, usability, visual polish, structure, and editability of AI-generated outputs across multiple domains including: * software engineering artifacts * UI/UX and design systems * presentation materials * spreadsheets and documents * multimodal and computer vision outputs Compensation: • Approx. $75/hour • Flexible task-based work • High performers may scale to substantial weekly earnings depending on task availability and quality Open Specializations: 1. Software Engineering Evaluators 2. UX/UI & Visual Design Evaluators 3. Computer Vision / Multimodal Evaluators What You Will Do: * Review AI-generated artifacts and compare multiple responses side-by-side * Rank outputs from best to worst using structured evaluation rubrics * Evaluate usability...

    $11 / hr Average bid
    $11 / hr Avg Bid
    44 bids

    ...build a lightweight internal web tool (POC) that helps validate migrated CMS pages (AEM-based) against predefined design and content standards. This tool will be used by product and content teams to ensure quality, consistency, and completeness of page migrations. ________________________________________ Project Overview We are migrating insurance product pages from an old template to a new “Multimodal Template.” We already have: • Figma designs (as source of truth) • Sample migrated pages (live URLs) • Initial component mapping (Excel) We want to build a tool that: 1. Takes mapping inputs (components, templates, product) 2. Allows users to input one or multiple page URLs 3. Automatically audits these pages 4. Outputs a structured validation report + sc...

    $73 Average bid
    $73 Avg Bid
    17 bids

    I’m building an NLP-driven, multimodal assistant that accepts text, image, and audio inputs, but its replies still drift into hallucination. The goal is straightforward: sharpen response accuracy so the system stays firmly grounded in fact. Right now the core pipeline is a Hugging Face Transformer model wrapped in a Retrieval-Augmented Generation (RAG) layer. I need you to audit the entire flow, diagnose where and why hallucinations appear, and then apply proven mitigation techniques. That could involve prompt engineering, better retrieval logic, truth-focused data augmentation, fine-tuning, or introducing guard-rail frameworks—whatever combination delivers measurably higher factual precision. Deliverables • A revised model or inference pipeline that demonstrabl...

    $71 Average bid
    $71 Avg Bid
    17 bids

    I want to build a custom large-language model that goes beyond text-only chat. The goal is a predictive engine that can read free-form text, combine it with numerical features, and return forward-looking insights. In practice that means designing an architecture able to embed and fuse both modalities, ...training pipeline with documented source code • Trained model weights and reproducible environment files • Evaluation report demonstrating predictive performance on unseen mixed data • Simple inference script or REST endpoint instructions If you have prior experience blending tabular and textual inputs or have leveraged architectures such as TabTransformer, RETAIN-style attention, or multimodal adapters on top of LLM backbones, mention it—those skills w...

    $251 Average bid
    $251 Avg Bid
    34 bids

    ...who can turn it into a publishable manuscript that meets typical IEEE or Springer journal standards. Here’s what I’m after: • Scope. A clear literature review on AI techniques currently used for diagnosing genetic conditions such as cystic fibrosis, sickle-cell, or rare chromosomal abnormalities. Contrast traditional pipelines with cutting-edge deep-learning approaches (CNNs, Transformers, multimodal models, etc.). • Original contribution. Either propose a novel framework or run a small-scale experimental study on an open dataset (e.g., ClinVar, DECIPHER). I’m open to a fresh algorithmic idea, an improvement to an existing model, or a hybrid ensemble—whichever you can substantiate with data. • Methodology & results. Describe data p...

    $85 Average bid
    $85 Avg Bid
    16 bids

    ...execute multiple AI workstreams end to end. This is a single requirement, not multiple specialist roles. The person should be strong across modern AI engineering and capable of taking problems from architecture and prototyping through optimization, deployment, and production readiness. The work may span LLMs / SLMs, recommendation engines, agentic interview workflows, AI-based result assessments, multimodal AI systems, classical ML, deep learning, and MLOps. This role is best suited for someone who is a strong AI generalist with solid engineering discipline and the ability to convert ambiguous problem statements into practical, scalable AI systems. The source role requires 5+ years of experience entirely in the AI/ML domain. What You Will Work On - Design and build AI solutio...

    $12 / hr Average bid
    $12 / hr Avg Bid
    52 bids

    I need a generative-AI developer to build a small-scale MVP of a multimodal chatbot that can act as a daily mentor for students preparing for Chartered Accountancy. The first release must live on a website and support natural, two-way voice conversations—speech-to-text on the way in, text-to-speech on the way out—so learners can talk to it as if they were speaking to a tutor. Core goals • Accurate CA guidance: the bot should answer syllabus-level questions, explain tricky concepts, and suggest study plans. • Fluid voice exchange: latency below two seconds using a stack such as Whisper / Web Speech API for recognition and a neural TTS engine for replies. • Continual engagement: greet students each day, track brief study logs, and offer tailored promp...

    $6 / hr Average bid
    $6 / hr Avg Bid
    22 bids

    ...(OCR, document parsing, transcription from voice/video) • Natural language processing including entity extraction, sentiment analysis, and contextual interpretation • Predictive and pattern-based modelling to generate forward-looking insights from historical data The ideal candidate will have strong expertise in: • Machine Learning / NLP architectures, including transformer-based models and multimodal processing (text, speech, video) • Data engineering and database design, covering both: • Structured systems (e.g. relational databases, data warehouses) • Unstructured data platforms (e.g. object storage, vector databases, knowledge graphs) • Scalable data pipelines for ingestion, processing, and model inference In addition, the candidate...

    $2520 Average bid
    $2520 Avg Bid
    54 bids

    We are UrgentHaul Logistics Europe, a fast-growing logistics provider headquartered in the Netherlands, with operations in Europe and Africa. We are building a high-end, enterprise-grade website that positions us alongside top-tier competitors such as DSV and time matters. The focus is on time-critical logistics, global trade corridors (Europe Africa), multimodal freight solutions, and high-conversion B2B lead generation. We are looking for an experienced WordPress Elementor designer/developer who can translate strategy into a visually premium, conversion-optimised website. Objective: Design and build a fully responsive, SEO-optimised, conversion-focused website using WordPress (Astra Pro) and Elementor (mandatory). The website must reflect a high-trust, enterprise logistics brand ...

    $19 / hr Average bid
    $19 / hr Avg Bid
    306 bids

    ...self-paced learning. 2.3 This role requires practical experience integrating both LLM (Large Language Models) and LMM (Large Multimodal Models) into real-world applications, including text generation and multimodal outputs such as diagrams and visual content. 3. Scope of Work 3.1 Design and develop backend systems using Python (FastAPI or similar) 3.2 Build AI-powered modules including AI tutor, curriculum generator, resource generator, marking system, and question bank system 3.3 Integrate multiple AI APIs (OpenAI, Claude, Gemini or similar) into a unified codebase 3.4 Develop systems using both LLM (text generation and reasoning) and LMM (multimodal outputs such as diagrams and visuals) 3.5 Design structured prompt engineering workflows and JSON-based...

    $11 - $18 / hr
    Sealed NDA
    $11 - $18 / hr
    155 bids

    I’m building a mobile-first AI agent that can fluidly switch between voice commands, standard text input, and basic gesture controls. The core logic, NLP pipeline, and gesture-recognition layer all need to sit inside a single, maintainable codebase that compiles cleanly for iOS and Android. You’ll start by designing the interaction flow: how spoken intent, typed text, or a swipe/pinch maps into the same intent engine. From there, I want the full implementation—speech-to-text, intent classification, gesture mapping, and the reply generation module—wired together behind a unified API so the mobile front end can call one endpoint regardless of modality. I’m comfortable with TensorFlow Lite or PyTorch Mobile for the on-device models and open to using platform-...

    $734 Average bid
    $734 Avg Bid
    182 bids

    ...for all 6,000 SKUs - Store vectors in Pinecone, Qdrant, or similar - Must support fast vector search - Deliver a stable, documented workflow - Clear architecture - Error handling - Confidence thresholds - Bundle logic - API endpoints or modules TECHNOLOGIES (Developer may choose best options) - Object detection: YOLOv8, Grounding DINO, or Vertex AI - Embeddings: OpenAI Vision, Vertex Multimodal, or similar - Vector DB: Pinecone, Qdrant, Weaviate - Automation: - Database: Nocodb WHAT I WILL PROVIDE - 6,000 SKU database - Sample images - Bundle examples - Titles (optional) REQUIREMENTS - Must have proven experience with: - Computer vision - Embeddings - Vector search - Object detection - API integrations - Automation workflows - Must deliver a fully working, production‑ready

    $1106 Average bid
    $1106 Avg Bid
    78 bids

    ...systems Requirements - Strong experience in Machine Learning / Deep Learning - Hands-on experience with Computer Vision - Proficiency in Python and ML frameworks such as PyTorch or TensorFlow - Experience working with large datasets and model training Preferred - Experience with face recognition systems - Familiarity with models such as ArcFace, FaceNet, or similar - Experience withLLMs or multimodal AI systems Application Please send: - Your resume - A short summary of your past Computer Vision or LLM projects - Any GitHub, portfolio, or relevant work (if available) We are prioritizing candidates who are available to start immediately, as the project will begin next week....

    $32 / hr Average bid
    $32 / hr Avg Bid
    77 bids

    ...and a full reference list. I am flexible on the exact style (APA, MLA, Chicago) as long as it is applied consistently and meets university standards. Research approach I only need a thorough, critical literature review—no primary experiments or case studies. The writing should synthesise current peer-reviewed work on computer-vision techniques, GAN identification, forensic audio analysis, multimodal fusion, and emerging AI counter-measures, weaving these strands into a cohesive argument that highlights research gaps and future directions. Originality requirements • Plagiarism score below 5 % on Turnitin. • AI-generated text under 5 % when tested by Turnitin’s AI detection module (or an equivalent recognised tool). • A certificate/report from ...

    $218 Average bid
    $218 Avg Bid
    10 bids

    We are looking for experienced AI Developers to help design and build an advanced AI-powered platform. The role involves developing intelligent chatbots, Retrieval-Augmented Generation (RAG) systems, multimodal AI capabilities, and scalable backend architectures. You will work closely with the founding team to bring innovative ideas to life—from concept to production-ready systems. Key Responsibilities Build and deploy AI chatbots using modern LLM frameworks Design and implement RAG pipelines for document and knowledge-base querying Integrate OCR and Vision models for document and image understanding Implement Text-to-Speech (TTS), Speech-to-Text (STT), and Speech-to-Speech (STS) pipelines Fine-tune LLMs to create offline, self-hostable AI models Architect and develop a...

    $77 Average bid
    $77 Avg Bid
    26 bids

    Top multimodal Community Articles