
Closed
Posted
Paid on delivery
I have an in-house project that needs a fully custom-built data pipeline. The goal is to ingest Parquet files arriving in our file system, apply the required transformations, and move the cleaned data on to our analytics environment. Off-the-shelf ETL tools do not meet our requirements, so I am looking for a bespoke solution designed from the ground up. Scope of work • Design the end-to-end pipeline architecture, choosing the most suitable framework (for example Python, Spark, or similar) while keeping future scalability in mind. • Build robust code that can automatically detect new Parquet files, validate their schema, transform the data as specified, and deliver the output to the target location I will provide once we start. • Add logging, error handling, and simple configuration files so the pipeline can be tweaked without code changes. • Package everything in a repo with clear setup instructions and a lightweight README so I can deploy it internally. Acceptance criteria 1. End-to-end execution succeeds on a supplied Parquet sample set. 2. Any corrupted or schema-drifted file is logged and quarantined without halting the run. 3. All configurable paths, credentials, and parameters sit outside the core code. 4. A short hand-off call and walkthrough of the repository. If you have built similar customer-made pipelines that process Parquet data straight from file systems, I’d like to see an example commit or short demo clip. Let me know your proposed tech stack and estimated timeline, and we can get started right away.
Project ID: 40609649
119 proposals
Remote project
Active 10 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
119 freelancers are bidding on average $464 USD for this job

Hi, creating a system that automatically detects new files and handles unexpected schema changes sounds challenging but interesting. I've worked on projects like a multi-vendor marketplace and a ticketing healthcare app, where I built data pipelines that dealt with file validation and error logging. I will handle the transformations and make sure paths and credentials are kept separate for easy updates. Did you think about using lightweight tools that can adapt to future data changes, especially schema drift? Let’s chat more and plan together to build something impressive. Regards, Nick.
$250 USD in 3 days
9.5
9.5

Hi, I’ve reviewed the project and the main requirement is to build a fully custom data pipeline that ingests Parquet files arriving in your file system, applies transformations, and delivers the cleaned data to your analytics environment. I’ll design end-to-end architecture with scalable data architecture choices, selecting Python and Apache Spark for robust processing. The pipeline will auto-detect new files, validate schemas, apply transforms, log errors, quarantine issues, and keep paths outside core code through configuration files. You’ll get a clean, well-documented repo with setup instructions and a lightweight README so I can deploy it internally. Let’s discuss here now.
$250 USD in 30 days
8.4
8.4

This looks like a great fit, Hi, I will build the Parquet ingestion pipeline with automatic file detection, schema validation, and quarantine logic for corrupted or drifted files, all driven by external config files so you never touch the core code to adjust paths or parameters. On a similar pipeline, separating the schema validation step into its own stage made debugging drift issues significantly faster. I will structure yours the same way. Questions: 1) What is the expected volume of Parquet files per run (tens, hundreds, thousands)? 2) Is the target analytics environment a database, a cloud bucket, or something like a local data warehouse? Share one of your sample Parquet files and I will confirm the schema handling approach today. Looking forward to discussing further. Best regards, Kamran
$277 USD in 10 days
8.6
8.6

I have 7+ years of experience in Full Stack Development, Web Scraping, Python, Excel VBA, SQL, and Automation. I have worked with MetLife GOSC, DXC Technology, and Elite Services. I can deliver a high-quality solution with quick turnaround, clear communication, and on-time delivery. I'm ready to start immediately and would be happy to discuss your requirements. Looking forward to working with you!
$500 USD in 7 days
8.6
8.6

Hi, I’m a senior developer with 20+ years of experience, Top Rated on Freelancer.com with 1,500+ completed projects, and I’ve built custom data pipelines from scratch before. I’ll design a Python-based pipeline using PySpark for parallel processing, ensuring it scales with your data volume. The system will monitor a watched directory for new Parquet files, validate schemas against your expected structure, apply transformations (custom logic you specify), and load results into your analytics environment while logging all operations. I’ll include a config file for paths, credentials, and parameters, add robust error handling to quarantine bad files without stopping the run, and package everything in a Git repo with setup instructions and a README for internal deployment. For the hand-off, I’ll provide a 30-minute walkthrough of the codebase. I can outline the timeline after reviewing your exact transformation rules and target location.
$400 USD in 5 days
7.9
7.9

HI As a full-service development team with extensive experience, we've tackled projects just like yours, building custom data pipelines that extract, transform, and load (ETL) Parquet files. We understand the limitations of off-the-shelf solutions and are well-versed in using Python, Spark, and other frameworks to produce precise, tailored pipelines. We will design your project with scalability in mind and create robust code that automatically detects new Parquet files,tempers schema mismatch and transforms data as specified. Our proficiency extends beyond just the technical aspects of your project; we prioritize logging, error handling, and simple configuration files so tweaks can be made without complex code changes. Our ability to deliver solutions that require minimal maintenance is one reason clients often turn to us for critical projects like yours. Additionally, once your pipeline is developed, we will package it in an organized repository with clear instructions for easy deployment into your analytics environment. Our ISO certifications attest to our commitment to consistently high-quality work while prioritizing client information security. Whether you need web applications or system level tools; rest assured that we make scalability our top priority. Thanks....
$500 USD in 7 days
8.1
8.1

⭐⭐⭐⭐⭐ Build a Custom Data Pipeline for Efficient Data Transformation ❇️ Hi My Friend, I hope you're doing well. I've reviewed your project requirements and see you're looking for a custom-built data pipeline for Parquet files. Look no further; Zohaib is here to help you! My team has successfully completed 50+ similar projects in building data pipelines. I will design and create a robust, scalable pipeline that meets your specific needs using the best-suited technology. ➡️ Why Me? I can easily build your custom data pipeline as I have 5 years of experience in data engineering and pipeline development. My expertise includes working with Parquet files, data transformation, and error handling. Besides, I have a strong grip on Python, Spark, and frameworks that ensure a smooth and efficient process for your project. ➡️ Let's have a quick chat to discuss your project in detail and let me show you samples of my previous work. Looking forward to discussing this with you in our chat. ➡️ Skills & Experience: ✅ Data Pipeline Development ✅ Python Programming ✅ Apache Spark ✅ ETL Processes ✅ Data Transformation ✅ Schema Validation ✅ Logging & Error Handling ✅ Configuration Management ✅ Version Control ✅ Repository Setup ✅ Scalability Solutions ✅ Data Quality Assurance Waiting for your response! Best Regards, Zohaib
$350 USD in 2 days
8.1
8.1

Hello, I checked your "Custom Data Pipeline Engineering" project and it looks like understanding the existing workflow will be important before making any changes. I've worked on similar PHP projects involving php, python, data processing, software architecture, mysql, data architecture, data management, apache spark and prefer delivering work in small milestones so everything stays easy to review and adjust if needed. If you can share a few more details about the current setup and expected outcome, I'll suggest the best approach and provide an accurate timeline. ⭐ 5.0/5 from a recent client: "Project was delivered before Time with Best professional Knowledge One could ever held. Thanks for the support" Final timeline and cost will be confirmed in chat after a complete understanding and documentation of the project expectations in detail.
$450 USD in 9 days
7.7
7.7

Hi, I built a secure CI pipeline with deterministic ingestion, the kind of setup where files get validated on arrival and bad ones are handled without breaking the run. Secure CI Pipeline: matched skills python, 5★ On the quarantine requirement, I'd keep it strict: a schema check on each Parquet file before transform, so any drift gets logged with the file name and moved to a quarantine folder, run keeps going. One question before I plan the architecture: roughly what daily file volume and size are we talking about? That decides plain Python versus Spark. Paths, credentials, and parameters would all live in a config file outside the core code, packaged in a repo with a README you can deploy internally. Happy to structure this as milestones so you release only on a passing run against your sample set. What framework do you lean toward? Adil
$358.16 USD in 7 days
7.5
7.5

Hi there, I understand you need a custom data pipeline to automate the ingestion of Parquet files. Operationally, a watcher service will detect new files arriving in a directory, triggering a job that validates each file's schema against a master definition. The job then applies your required transformations and loads the clean data into the target analytics environment. Any file failing validation is automatically moved to a quarantine location and logged, ensuring the pipeline continues running without interruption. Technical approach: I recommend a Python-based solution using libraries like Pandas and PyArrow for efficient Parquet processing. A file system watcher will trigger the processing script. We will containerize the entire application with Docker for seamless deployment and isolate dependencies. All paths, credentials, and parameters will be managed in an external YAML configuration file. Core modules: - File Ingestor: Monitors the source directory for new files. - Schema Validator: Checks data structure against a predefined template. - Transformation Engine: Applies the specified data cleaning and shaping logic. - Data Loader: Writes the final output to your target system. - Exception Handler: Manages logging and quarantines invalid files. We'll start by building and testing the core transformation logic with your sample Parquet set. Then, we'll build the file-watching and error-handling framework around it, package it, and prepare for a clear hand-off. Regards, Rohit
$250 USD in 10 days
8.0
8.0

Hi there, You need a bespoke end‑to‑end pipeline that watches a directory for incoming Parquet files, validates and transforms them, then ships the clean output to your analytics store while handling schema drift and failures gracefully. The main challenge is reliable file detection and quarantine of corrupted or drifted files without stopping the whole run, plus keeping all secrets and paths out of the code. My approach would be to build the core in Python using PySpark (or Dask if you prefer a lighter footprint) for scalable columnar processing, coupled with watchdog for real‑time file pickup. Schema checks will be done with PyArrow metadata; any mismatch will be logged and the file moved to a quarantine folder. Configuration (paths, credentials, transformation rules) will reside in a YAML file parsed at startup, and logging will be powered by structlog for easy querying. The repo will include Dockerfile, a concise README and a one‑click script to spin the container up on your internal servers. I’ll also provide a short hand‑off call and walkthrough. Do you plan to load the final data into a data warehouse (e.g., Snowflake, BigQuery) or a file‑based lake such as S3/ADLS, and does it require any specific partitioning scheme? Thanks, please get in touch – looking forward to delivering a reliable pipeline for you.
$350 USD in 7 days
7.5
7.5

Hi, I’ve reviewed your pipeline requirements carefully, and I can build a reliable custom solution for detecting Parquet files, validating schema drift, transforming records, and delivering clean output into your analytics environment. This is a strong fit for my Python and Software Architecture background, and I’ll structure the repo with configurable paths, credentials, logging, quarantine handling, and clear deployment instructions so your team can run it internally with confidence. I’d recommend a Python-based pipeline with modular processing, validation layers, and fault-tolerant logging designed for future scaling. I can deliver the full build, test it against your sample set, and include the hand-off walkthrough within 12 days. Will this pipeline run on a single internal server, or should I design for distributed processing from day one? Best regards, KANIKA
$650 USD in 12 days
6.9
6.9

Hello, I have worked on custom file based pipelines where Parquet data needed validation, transformation, logging, and safe delivery into analytics storage. For your project, I would focus first on Software Architecture and Data Architecture so the flow can handle schema drift, quarantine bad files, and scale cleanly. I would likely use Python with Apache Spark if the sample size and growth path justify it, with config files for paths, credentials, and parameters outside the core code. The repo would include setup steps, clear error handling, and a short walkthrough so your team can run and adjust it internally. Best regards, Teo
$500 USD in 5 days
6.5
6.5

Hello I have gone through your specific requirement for Parquet data pipeline. I have chosen PyArrow over fastparquet because schema checks and nested types stay more predictable. I will build a Python worker that watches for new Parquet files then validates, transforms and quarantines bad ones, at least that is where I would start. And I will keep config in YAML with logging built around structlog. Built data pipelines for analytics clients across 8 ingestion projects. Repo and demo clips I can show. Where will the cleaned data land after processing? Need to clear the transform rules and daily file volume first. Free for a quick call this week? Dev S.
$500 USD in 5 days
6.7
6.7

Hi there, Your project for a custom data pipeline sounds like a perfect fit for my expertise. I've built numerous bespoke data processing solutions, including those handling Parquet files directly from file systems, and I understand the limitations of off-the-shelf tools. My proposed tech stack would likely involve Python with libraries like Pandas for efficient data manipulation, and potentially Apache Spark if the scale and complexity warrant it for future scalability. I'll design an architecture that prioritizes robustness, automatic file detection, schema validation, and fault tolerance, ensuring corrupted files are quarantined without halting the process. Logging and configuration will be externalized for easy adjustments. I'll package the entire solution in a clean Git repository with straightforward setup and a concise README. My plan is to: 1. Design the pipeline architecture. 2. Develop the core ingestion, transformation, and delivery logic. 3. Implement robust error handling and logging. 4. Ensure all configurations are externalized. 5. Package and document the code for easy deployment. I'm confident I can deliver a high-quality, custom solution within your timeline and budget. Let's discuss your specific requirements and get started. Best regards, Ashwani
$600 USD in 7 days
6.3
6.3

As a seasoned software engineer with a diverse portfolio, I am exactly who you need for this project. At Fourge, we specialize in creating innovative and scalable solutions tailored to specific business needs. I have extensive experience in designing and engineering bespoke data processing pipelines, utilizing languages such as Python to maximize efficiency and flexibility while keeping scalability in mind. Not only can I develop a custom-built data pipeline that aligns with your unique requirements, but I will also carefully design it with logging mechanisms, error-handling protocols, and configuration files that can make simple modifications without needing significant code changes. Additionally, thanks to my familiarity with Parquet data processing straight from file systems, I can integrate any necessary schema validation and quarantine methods to ensure only clean data is transported to the target location. By leveraging my skills in software architecture and system design, your pipeline will be robust yet easy to use. Furthermore, as a client-oriented professional, I appreciate the importance of clear documentation. Upon completion of the project, along with the deployment-ready repository, I will provide you with comprehensive setup instructions and facilitate a walkthrough to guarantee a seamless post-project transition. Let's discuss your possible timelines and proposed tech stack further so we can get started on this exciting project!
$250 USD in 2 days
6.4
6.4

Hello! We can build a custom data pipeline for your Parquet processing workflow. 1. What should be the main output destination for the cleaned data? 2. Do you already have the transformation rules and sample files ready? — About us We are dZENcode – a full-cycle IT company for digital product development: from design and programming to integrations and post-release support. We build projects from scratch and also work on existing solutions that need further development, improvements, or technical support. You can find detailed information about our services and rates on our official website: https://dzencode.com. Please review it – after that, we can discuss the details and agree on the next step. ⚠️ After clarifying all details, we will define the scope, the suitable cooperation format – task-based, outsourcing, or outstaffing – and the final cost. Projects are guaranteed to reach release with us: • 10+ years providing IT services; • 90+ in-house specialists; • 250+ public reviews since 2015; • We support products under SLA after launch; • We work under NDA and a company contract!
$500 USD in 7 days
6.7
6.7

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Matlab, Python, and similar tools. I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$500 USD in 7 days
6.4
6.4

Your pipeline will fail silently if schema drift occurs mid-batch and you don't implement a quarantine layer with rollback logic. Without proper checkpointing, a single corrupted Parquet file can poison your entire analytics environment. Quick questions - what's your expected daily file volume and average Parquet size? And are you running this on-prem or do you have cloud object storage available? Here is the architectural approach: - APACHE SPARK: Build a distributed PySpark job with schema validation hooks that quarantine malformed files before they reach your analytics layer, using Delta Lake for ACID transactions. - DATA ARCHITECTURE: Implement a three-tier pipeline (landing/staging/curated) with automated schema evolution tracking and dead-letter queues for failed records. - PYTHON: Write orchestration scripts with Watchdog for filesystem monitoring, Pydantic for config validation, and structured logging to Elasticsearch for full audit trails. I've built 4 production Parquet pipelines processing 2TB+ daily across fintech and healthcare clients. Let's schedule a 20-minute architecture review so I can map your transformation logic before writing a single line of code.
$450 USD in 10 days
6.4
6.4

Hi, I can build a custom Parquet pipeline with clear separation between file discovery, schema validation, transformation, delivery, and monitoring. I recommend Python with PyArrow and Polars for efficient single-server processing, with an abstraction that allows Spark adoption if future volume requires distributed execution. The final choice will follow a quick review of file sizes, frequency, transformation complexity, and infrastructure. The pipeline will detect new files, prevent duplicate processing through manifests/checkpoints, validate expected schemas, apply configurable transformations, and write outputs atomically. Corrupted or schema-drifted files will move to quarantine with a detailed error record while valid files continue processing. Paths, credentials, schema versions, and runtime parameters will remain outside the code in YAML and environment variables. I’ll include structured logging, retries, data-quality checks, idempotency, metrics, unit/integration tests, and Docker packaging where appropriate. Delivery includes the complete repository, configuration templates, architecture notes, README, setup commands, tests against your sample data, and a handover walkthrough. I can share relevant code examples where confidentiality permits. The timeline can be confirmed after reviewing the samples and transformation rules. Question 1: What are the daily volume and largest file size? Question 2: What is the target analytics environment? Regards, Houssame
$500 USD in 7 days
6.6
6.6

İzmir, Turkey
Member since Jul 25, 2025
₹1500-12500 INR
$2-8 USD / hour
$2-8 CAD / hour
₹400-750 INR / hour
£20-250 GBP
₹1500-12500 INR
₹1500-12500 INR
₹750-1250 INR / hour
min $50 USD / hour
$10-40 USD
₹12500-37500 INR
$250-750 USD
$25-50 USD / hour
₹750-1250 INR / hour
₹12500-37500 INR
$2-8 USD / hour
$250-750 USD
$30-250 USD
$750-1500 AUD
$10-100 USD