
Completed
Posted
Paid on delivery
Project Description I'm looking for an experienced Azure Databricks / PySpark developer to help convert an existing SQL Server view into a Databricks PySpark transformation. Project Overview I have: An existing SQL view with multiple CTEs. An existing Databricks Asset Bundle notebook that will be used as the template. A source Delta table already available in Unity Catalog. The task is to replace the existing business transformation with a new PySpark implementation while preserving the existing notebook framework. Source The source is a Delta table in Unity Catalog. The notebook should: Read data from the source table. Apply filters. Replicate the SQL CTE logic in PySpark. Perform required aggregations and calculations. Produce the final dataframe. Preserve the notebook's existing audit and validation framework. Write the output Delta table. Create a SQL View on top of the Delta table. SQL Logic The SQL contains: Multiple CTEs Window selection logic Aggregations CASE expressions GROUP BY Date filtering Calculated measures The objective is to produce the same output as the SQL view using PySpark DataFrame APIs. Existing Notebook The notebook already contains: Logging framework Metadata handling Schema validation Audit column generation Exception handling Delta write logic These should remain unchanged. Only the business transformation needs to be replaced. Deliverables Complete PySpark implementation. Clean, optimized, production-ready code. Logic matching the SQL view. Compatible with Databricks Runtime. Delta table creation. SQL View creation. Assistance with testing if required. Required Skills Azure Databricks PySpark Spark SQL Delta Lake SQL Server DataFrame API Window Functions Databricks Asset Bundles (preferred) Nice to Have Experience with Unity Catalog Production ETL development Azure Data Factory knowledge What I'll Provide Existing SQL View Existing Databricks notebook Business logic
Project ID: 40615633
8 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

Converting CTE-heavy SQL views into PySpark is one of those tasks where a literal translation works but performs terribly. I'll restructure the CTE chain into a proper DataFrame pipeline, mapping each CTE to a well-named intermediate DataFrame so the logic stays readable without the performance hit of nested subqueries. I'll read from your Unity Catalog source, replicate every window function, CASE expression, and aggregation using PySpark's DataFrame API, then wire the final output into your existing notebook's audit and Delta write framework, untouched. One thing worth flagging: Spark's window function ordering defaults can differ from SQL Server's, especially with NULLs. I'll make sure the sort behavior matches exactly. 1) Are any of the CTEs referencing each other recursively, or is it a straight linear chain? Happy to talk details in chat. Shayan
₹660 INR in 3 days
0.0
0.0
8 freelancers are bidding on average ₹2,970 INR for this job

Hi , I am from Bangalore. I am certified in databricks. I will get this done.I am data engineer with 7 years of experience. Let’s connect
₹4,900 INR in 1 day
5.9
5.9

Hi, Your Azure Databricks/PySpark migration project is a great match for my experience. I have 18+ years of experience in SQL Server, Azure, Python, and production ETL, with a 4.9 rating and long-term repeat clients. Approach: Convert the SQL Server CTE logic into optimized PySpark DataFrame transformations. Preserve the existing notebook framework, logging, audits, validation, and exception handling. Write the output to Delta Lake and create the SQL View. Validate results against the existing SQL view for functional parity. One question: approximately how many CTEs and source tables are involved in the SQL view? Looking forward to discussing the project!
₹5,000 INR in 2 days
4.8
4.8

Hello, Refreshing the SQL Server view with a PySpark transformation in Azure Databricks is about seamlessly transitioning the existing business logic to a new PySpark implementation while maintaining the current notebook structure. I would start by reviewing the SQL view's CTEs and window selection logic to replicate them accurately in PySpark. The main focus will be on ensuring data integrity, performance, and compatibility with Databricks Runtime. My experience includes similar projects where I modernized SQL transformations into PySpark workflows, emphasizing clean, optimized code and logic alignment. A few questions: - What specific aggregations and calculations should be performed? - Are there any additional requirements for the output Delta table? - Do you have any preferences for testing frameworks? I look forward to discussing the implementation details further. Best regards,
₹2,000 INR in 7 days
2.4
2.4

As a data automation developer, I can convert your SQL view into a PySpark transformation while preserving Databricks audits, Delta output, and SQL view creation. My Python and Spark data work fits this, and I can start now with one CTE converted free for you to judge. Please get in touch.
₹600 INR in 1 day
0.6
0.6

Converting SQL to PySpark isn't just a syntax change—it's about preserving business logic while optimizing for Spark performance. Your existing Databricks framework should remain untouched, with only the transformation layer replaced. That's exactly the type of migration our team has delivered for enterprise data platforms. Hi, We've reviewed your requirements and can seamlessly convert your SQL Server view into an optimized PySpark transformation while preserving your existing Databricks Asset Bundle framework, including logging, metadata, schema validation, audit columns, exception handling, and Delta write logic. Our team has over 10 years of experience working with multiple clients on ETL pipelines, SQL optimization, PySpark, Delta Lake, Databricks, API integrations, and enterprise data solutions. References can be shared if you'd like. We'll replicate the CTEs, window functions, CASE logic, aggregations, and calculated measures using efficient PySpark DataFrame APIs, ensuring the output matches the existing SQL view. We'll also create the Delta table, SQL View, and assist with validation to confirm parity between both implementations. Let's have a short discussion to review the SQL view and notebook so we can deliver a clean, production-ready migration.
₹2,800 INR in 7 days
0.0
0.0

The risk in a SQL to PySpark conversion is not writing the code, it is that the converted version looks correct and quietly returns different rows. Null ordering in window functions, the default frame when you leave out ROWS BETWEEN, and how Spark treats ties all differ from SQL Server, and those gaps only show up on real data. So I would build it, then run both side by side and reconcile row counts and aggregate totals per grouping key before calling it done. If the numbers do not tie out exactly, it is not finished. I would keep your existing logging, schema validation and audit column patterns exactly as they are and slot the transformation into them rather than reworking your framework. Can you share the view definition and roughly how many rows it returns? That tells me whether this is a straight translation or whether the CTEs need restructuring for Spark.
₹5,000 INR in 7 days
0.0
0.0

Hi, there. In a recent project, I converted a complex SQL Server view into a PySpark transformation within an Azure Databricks environment, ensuring data integrity and performance optimization. This involved replicating multiple CTEs, window functions, and aggregations while maintaining the existing auditing and validation framework. To tackle your project, I offer to implement the PySpark logic that mirrors the SQL view's transformations using DataFrame APIs, preserving all existing notebook functionalities such as logging and schema validation. My approach includes reading from the source Delta table, applying necessary filters, and executing the required aggregations and calculations to produce the final DataFrame. Additionally, I will ensure that the output Delta table and SQL view are created correctly, adhering to the specifications provided. My experience with similar transformations has equipped me with the skills to handle complex SQL logic and optimize it for performance in a PySpark context, ensuring reliable and production-ready code. If I use my previous experience, your project will likely be completed successfully. Hope to discuss this in detail. Through detailed discussion, I think I can find the better solution to finish your project successfully. Thank you.
₹2,800 INR in 7 days
0.0
0.0

Hyderabad, India
Payment method verified
Member since Sep 10, 2022
₹600-1500 INR
₹2000-3500 INR
₹600-1500 INR
₹600-1500 INR
$5000-10000 NZD
$1500-3000 SGD
$750-1500 USD
₹75000-150000 INR
₹1500-12500 INR
$15-25 USD / hour
$15-25 USD / hour
$250-750 USD
₹750-1250 INR / hour
₹12500-37500 INR
$30-250 USD
₹12500-37500 INR
₹600-1500 INR
$250-750 USD
₹750-1250 INR / hour
₹750-1250 INR / hour
£750-1500 GBP
₹10000-45000 INR
₹1500-12500 INR
$15-25 USD / hour