BigQuery jobs
...Understanding of responsible and lawful web data collection, including , terms of service, privacy, and applicable data protection requirements. Good to Have OpenCV and image processing NLP / NER LLM-based document extraction Multilingual OCR Indian regional language OCR Entity resolution / record linkage Data lineage and provenance Great Expectations, Soda, Pandera, or similar dbt and dbt tests BigQuery AWS S3, Lambda, Batch GCP Cloud Storage Distributed/batch processing Experience with government, legal, financial, regulatory, or public-sector documents Ideal Candidate We are not looking for someone who has only built basic scraping scripts. The ideal candidate should have experience operating a production system involving: Hundreds of Sources → Crawling → Document...
...geographically. The output feeds an academic project modelled on recent work in the economics of innovation that maps patenting to geography — see [attached paper / link], e.g. the commuting-zone patent maps in Figure 13. That's the kind of geocoded output I'm after. What you'll build Pull US patent records for a set of search terms and CPC classes I provide, from PatentsView, Google Patents (BigQuery), or USPTO bulk data — whichever you can justify. For each patent: patent number, title, abstract, filing and grant dates, assignee(s), inventor(s), and inventor/assignee locations. Clean and disambiguate assignee names (e.g. collapsing "Corning Inc" / "Corning Incorporated" / "CORNING INC." into one entity). Geocode inventor ...
...etc.) built with a modern JS library such as , D3.js, Highcharts or a framework you’re comfortable with. • Ability to filter by date range, device type and referral source right from the dashboard. • A lightweight back-end hookup so the dashboard pulls fresh data automatically from my existing analytics stack (Google Analytics 4 is in place; I can expose an API or give you access to BigQuery if that’s simpler). • Mobile-first, responsive layout that can drop straight into a React component on my current site. • Clear, maintainable code plus a short README so I can re-deploy or extend the visuals later. Acceptance criteria 1. All chosen engagement metrics load under three seconds on a typical broadband connection. 2. Visual elements res...
...• Improve marketing attribution, conversion tracking and campaign measurement. • Translate business requirements into technical requirements and actionable user stories. • Lead initiatives from discovery → implementation → measurement → optimization. Technology Exposure • CRM: Braze, Salesforce Marketing Cloud, Iterable, CleverTap, MoEngage • CDP: mParticle, Segment • Data: Azure, Snowflake, BigQuery, Redshift • Analytics: GA4, Amplitude, Looker, Power BI • Activation: Census / Reverse ETL • Tracking: Google Tag Manager, UTM & event tracking • Channels: Email, SMS, Push, WhatsApp What We’re Looking For • 6+ years in CRM, Lifecycle Marketing, Marketing Automation or MarTech. • Hands-on experienc...
I run a dynamic site that can spin...the drop-off. I want you to pinpoint why, validate it with log-file or server data, and turn that insight into an actionable roadmap. Your deliverable should be a concise technical audit covering crawl behaviour, indexation gaps, orphan pages, and duplicate-content traps—followed by a step-by-step implementation plan I can hand to engineering. If you rely on tools like Screaming Frog, Sitebulb, BigQuery log analysis, or custom scripts, feel free to bring them into the process; just document the key findings and recommended fixes. I’m ready to move quickly and will provide server logs, Search Console exports, and database access (read-only) as needed so you can dive straight into the data. Let’s turn those invisible pages into ...
...static spreadsheets. The build must support: • Real-time data analysis so new encounters, labs, and vitals are visible in minutes, not days. • Customizable reporting templates our analysts can adjust without developer help. • Data visualization tools that present trends clearly for non-technical decision-makers. I’m open to the platform you recommend—whether that’s Snowflake, SQL Server, BigQuery, Redshift, or a comparable HIPAA-compliant stack—as long as you outline why it fits our scale and security needs. Please plan for typical healthcare data standards (HL7 v2, FHIR, or CCD) and include an ETL/ELT pipeline that cleans, normalizes, and de-identifies data where required for compliance. Deliverables expected: 1. Dimensional ware...
...Engineer to join a remote technical project supporting an international client. * It is an open contract and you get paid monthly. ### Requirements • 2–4 years of professional Data Engineering experience • Strong Python and SQL skills • Hands-on experience with Apache Spark • Experience with Airflow, Dagster, or Prefect • Experience with dbt and modern data pipelines • Knowledge of Snowflake, BigQuery, Redshift, or Databricks • Experience with AWS, Azure, or GCP • Strong understanding of data modeling and Git workflows • Experience working with production data pipelines • Strong spoken and written English communication skills ### Nice to Have • Kafka, Kinesis, or Flink • Terraform and CI/CD • Docker and K...
...humidity-logging solution that can capture readings at regular intervals and send them straight to Google Cloud Platform in near-real time. I am open to microcontrollers such as ESP32, Raspberry Pi, or any low-power alternative, provided the final build can connect over Wi-Fi or LTE and push data securely—ideally through Cloud IoT Core or Pub/Sub—into a time-series bucket I can later analyse with BigQuery or Data Studio. What matters most is that the entire path from sensor to cloud is proven and robust. You are free to suggest the specific humidity sensor, the firmware language (Arduino, MicroPython, C, etc.), and the best authentication method for GCP, as long as you keep power consumption modest and the solution easy to reproduce. Please highlight your relevant e...
I have a collection of CSV files that need to live in BigQuery and feed directly into Looker Studio dashboards. The workflow I have in mind is straightforward: create the initial dataset and tables in BigQuery, load the first batch of files, then automate a weekly refresh so each new CSV lands in the right spot without manual intervention. Once the data sits reliably in BigQuery, I’d like clean, visually clear reports in Looker Studio that focus on two areas—sales results and overall operational performance. Think trend charts, variance indicators, and drill-downs that anyone on the team can use without additional training. Deliverables • BigQuery dataset and table schema set up for incoming CSVs • Automated weekly load process (Cloud...
...from different sources, cleaning and transforming it, and storing it in a database or data warehouse. You may also work with APIs, cloud services, ETL workflows, and large datasets. The ideal candidate should have strong experience with Python, SQL, data modeling, and ETL/ELT pipelines. Experience with AWS, GCP, or Azure is preferred. Knowledge of tools such as Airflow, Spark, Kafka, Snowflake, BigQuery, or Databricks would be a plus. Please include the following in your proposal: A short introduction about your Data Engineering experience Examples of similar projects you have completed The technologies and cloud platforms you have used Your availability and expected rate A short video introduction to help me evaluate your communication skills and overall fit This may become ...
...for account changes and client communication. ## Required Technology and Integrations The solution should use a lightweight, scalable stack based around: * Python * Google Ads API * GA4 Data API * OpenAI API and/or Claude API * Google Workspace APIs * Microsoft Graph API * Secure OAuth * Scheduled cloud jobs * Version-controlled source code The architecture should allow future migration to BigQuery and a custom application. ## Google Ads and GA4 The system should securely pull scheduled data across multiple client accounts, including: * Campaign, ad group and keyword performance * Search terms * Spend, clicks, impressions, CTR and CPC * Conversions, conversion value, CPA, CVR and ROAS * Campaign budgets and bid strategies * Impression share and lost impression share * RSA ...
...overdue tasks—so managers see issues before they escalate. • Provide an architecture that leaves room to add classic workflow features such as task management, automated notifications, or progress tracking when we scale. Preferred stack I’m most comfortable with common analytics tools—Python (Pandas, scikit-learn), SQL, Power BI or Tableau for visualization, and a cloud database (PostgreSQL, BigQuery, or similar). If you have an alternative stack that hits the same goals, feel free to suggest it. Deliverables 1. A functioning web-based (or easily hosted) workflow optimization dashboard tied to my three data streams. 2. Annotated source code plus an ETL script or pipeline definition. 3. A short hand-over session or documentation that shows me how...
The goal is to build a Claude-powered agent that can ingest our day-to-day operational data, turn it into clear periodic reports, and surface forward-looking insights through predictive analysis. Data scope • Data sources include production logs, supply-chain metrics, and KPI spreadsheets stored in Google BigQuery and Google Sheets. • Volumes are moderate—about 2 GB daily—so latency is less important than accuracy and clarity of the output. Key capabilities I need baked into the agent 1. Scheduled and on-demand report generation in PDF and Markdown, summarising current operational performance against set KPIs. 2. Predictive analysis that flags potential bottlenecks or capacity issues at least two weeks in advance, using past trends. 3. A simple ...
...Cloud, transforms it, and lands it cleanly in BigQuery. The core stack must be Airflow for orchestration, Dataproc (running PySpark) for heavy transformations, and native BigQuery SQL for final modelling and reporting layers. You will design and implement: • Secure ingestion from the on-prem source into GCS staging • Airflow DAGs that trigger Dataproc jobs, handle retries, logging and alerting • PySpark transformation scripts on Dataproc, tuned for performance and cost • BigQuery SQL models that expose the refined tables • Parameterised configuration so environments can be promoted from dev to prod without code changes Acceptance criteria: – A scheduled Airflow DAG runs end-to-end without manual intervention. &ndash...
...NLP, computer vision, or generative AI is a plus 4. Data Engineer - Build and maintain robust data pipelines and ETL processes - Design and manage data warehouses and data lakes - Work with large-scale distributed systems such as Apache Spark, Kafka, or Airflow - Collaborate with data scientists and analysts to support data needs - Experience with SQL, NoSQL databases, and cloud data services (BigQuery, Redshift, Snowflake) All candidates should have strong communication skills, the ability to work in an agile environment, and a passion for technology and innovation. Please include your portfolio, relevant experience, and expected rate in your proposal. Preferred Candidate Locations: We are specifically looking for developers based in the following countries or regions: - Turk...
...— you focus on teaching and recording. What you'll deliver ~47 video lectures (each ≈ one focused topic, ~60 min) Live coding & tool demos with screen recording Hands-on lab walkthroughs for each module Supporting slides, code samples, and datasets Topics to cover (full syllabus provided) Python & SQL foundations, Linux & Git PostgreSQL, data modeling, star schemas, SCDs Cloud warehousing — BigQuery & Snowflake (partitioning, cost tuning) ELT pipelines with dbt (models, tests, snapshots, lineage) Workflow orchestration with Apache Airflow PySpark for big data + real-time streaming with Kafka & Spark Structured Streaming Docker, CI/CD (GitHub Actions), Terraform basics, Azure Data Factory End-to-end capstone pipeline (ingest → tra...
...network—across multiple instances. The goal is to surface near-real-time insights that can be fed into dashboards or alerts so I can react quickly when workloads spike. The solution should authenticate through a service account, call the Cloud Monitoring API (formerly Stackdriver) or the Compute Engine API where relevant, and output data in a JSON structure that’s easy to forward to Pub/Sub or store in BigQuery later. Lightweight dependency on google-cloud-monitoring / google-api-python-client is fine; please keep the setup straightforward via requirements.txt. Deliverables • Clean, well-documented Python 3.x script (or module) that runs from the command line • README covering setup, service-account configuration, and example execution • Sample ...
...technologies like Kafka, multi-cloud services, auto-scaling using GKE, Load balancers, APIGEE proxy API management, DBT, using LLMs as needed in the solution, redaction of sensitive information, DLP (Data Loss Prevention) etc. ● Work extensively on Google Cloud Platform (GCP) services such as: ○ Data-flow for real-time and batch data processing ○ Cloud Functions for lightweight serverless compute ○ BigQuery for data warehousing and analytics ○ Cloud Composer for orchestration of data workflows (on Apache Airflow) ○ Google Cloud Storage (GCS) for managing data at scale ○ IAM for access control and security ○ Cloud Run for containerized applications Should have experience in the following areas : ○ API framework: Python FastAPI ○ Processing engine: Apache Spark ○ Messaging and stre...
Key Requirements: 3–5 years of hands-on Tableau experience Strong Tableau dashboard development and data source management experience Good SQL skills Exposure to Tableau Server, including publishing, schedules, permissions, and troubleshooting refresh failures Familiarity with AWS and cloud data warehouses such as Amazon Redshift, Snowflake, or BigQuery Strong verbal and written communication skills Comfortable working with unfamiliar datasets and taking direction from internal teams Willingness to work a night shift aligned to US Eastern Time Preferred Experience: Tableau Server administration AWS fundamentals and Tableau connectivity with cloud-hosted data sources Extract performance optimization Managed services or client-facing support environments Exposure to Tableau Cl...
...ETL/ELT solutions using GCP services such as BigQuery, Dataflow, Dataproc, Cloud Storage, Pub/Sub, and Composer. Lead data modernization and migration projects from legacy/on-premise systems to GCP. Implement data integration, transformation, and orchestration processes for large-scale datasets. Ensure data quality, governance, security, and performance best practices. Collaborate with business stakeholders, architects, and cross-functional teams to deliver data-driven solutions. Required Skills 8+ years of experience in Data Engineering and Data Warehousing. Strong hands-on experience with GCP Data Engineering services. Expertise in building batch and real-time data pipelines. Proficiency in SQL, Python, and ETL/ELT frameworks. Experience with BigQuery, Dataflow (Apache B...
...via cached_network_image, and compiles to native for store review. Backend – scale first: Firebase for auth and speed, plus GCP for heavy lifting Auth: Firebase Auth (email, phone OTP, Google, Apple) Database: Firestore for profiles, Cloud SQL Postgres for the immutable ledger Storage: Cloud Storage for media, with signed URLs Functions: Cloud Functions gen2 for ad-event ingestion Analytics: BigQuery + Looker Studio GDPR: data residency in asia-south1 (Mumbai), consent mode v2, data deletion jobs Alternative if you expect >1M DAU quickly: AWS (Cognito, S3, Lambda, Aurora Serverless, CloudFront). Firebase is faster to MVP. CMS/Admin: React + Firebase Admin SDK. Lets you upload, schedule, set payout threshold, and manually adjust balances with audit trail. Ads: not AdM...
...via cached_network_image, and compiles to native for store review. Backend – scale first: Firebase for auth and speed, plus GCP for heavy lifting Auth: Firebase Auth (email, phone OTP, Google, Apple) Database: Firestore for profiles, Cloud SQL Postgres for the immutable ledger Storage: Cloud Storage for media, with signed URLs Functions: Cloud Functions gen2 for ad-event ingestion Analytics: BigQuery + Looker Studio GDPR: data residency in asia-south1 (Mumbai), consent mode v2, data deletion jobs Alternative if you expect >1M DAU quickly: AWS (Cognito, S3, Lambda, Aurora Serverless, CloudFront). Firebase is faster to MVP. CMS/Admin: React + Firebase Admin SDK. Lets you upload, schedule, set payout threshold, and manually adjust balances with audit trail. Ads: not AdM...
...Engineer (AWS) Technical instructor (AWS re/Start) Currently transitioning toward enterprise AI architecture The role focuses on: Microsoft Copilot Copilot Studio Power Automate Power Platform AI agents and workflow automation Enterprise AI solutions GCP / BigQuery integration AI governance and enterprise AI adoption Translating business problems into technical AI solutions What I need help with: Hands-on Copilot Studio Power Automate workflows Enterprise AI solution design AI workflow architecture AI governance basics GCP / BigQuery basics Real-world enterprise AI use cases Ideal mentor: Enterprise architect AI solutions architect Power Platform architect AI automation engineer Someone who has worked in enterprise AI transformation environments I’m NOT looking f...
...Engineer (AWS) Technical instructor (AWS re/Start) Currently transitioning toward enterprise AI architecture The role focuses on: Microsoft Copilot Copilot Studio Power Automate Power Platform AI agents and workflow automation Enterprise AI solutions GCP / BigQuery integration AI governance and enterprise AI adoption Translating business problems into technical AI solutions What I need help with: Hands-on Copilot Studio Power Automate workflows Enterprise AI solution design AI workflow architecture AI governance basics GCP / BigQuery basics Real-world enterprise AI use cases Ideal mentor: Enterprise architect AI solutions architect Power Platform architect AI automation engineer Someone who has worked in enterprise AI transformation environments I’m NOT looking f...
...management policies, slashing infrastructure overhead through data-driven resource sizing and Committed Use Discounts (CUDs). Global Networking Management: Engineered globally distributed infrastructure using GCP Cloud Load Balancing, Cloud Armor DDoS mitigation, and cross-region VPC network peering. Data & AI Infrastructure: Managed resilient backends for data pipelines utilizing Cloud Spanner, BigQuery, and Vertex AI infrastructure, ensuring low-latency data access for analytical applications. Enterprise Kubernetes (GKE / EKS) Orchestration Production-Grade Cluster Operations: Designed and managed multi-region Google Kubernetes Engine (GKE) and Amazon EKS clusters, executing seamless, zero-downtime canary upgrades of production control planes. Advanced Networking & ...
...Required Skills 7–10 years in Data Engineering (3+ years in lead role) Strong in SQL, Python, and Spark (PySpark/Scala) Experience with Snowflake, Databricks, BigQuery, or Redshift Solid understanding of data modeling (Kimball / Data Vault) Hands-on with streaming systems (Kafka / Flink) Familiarity with Terraform, CI/CD, and Cloud (AWS/GCP/Azure) --- Good to Have Experience with Feature Stores (Feast, Tecton) Knowledge of Data Mesh / CDC tools (Fivetran, Airbyte) Exposure to Graph or Vector Databases Open-source contributions or advanced degree --- Tech Stack Python, SQL, PySpark, Scala | Snowflake, Databricks, BigQuery Airflow, dbt | Kafka | AWS/GCP | Docker, Kubernetes, Terraform --- Interested candidates can reach out directly: Nine zero two...
...datasets. Day-to-day you will design and deploy solutions across all three major clouds—AWS, GCP, and Azure—so confidence working in each environment is essential. Our most performance-critical jobs are already running in Spark, and I’m looking for someone who knows how to squeeze every ounce of efficiency from it. If you have Airflow, dbt, Kafka, Kinesis, or warehouse experience with Redshift, BigQuery, or Snowflake, all the better. Native Portuguese is a must, and you’ll collaborate daily in English at a C1/C2 level with distributed stakeholders. The position is full-time and fully remote within Brazil; I’m interested in engineers who enjoy pairing with product managers, setting data quality guardrails, and iterating quickly on business-facing u...
...occupier leads delivered in the first 30 days of live operation. 2. Predictive scores showing ≥70 % precision when back-tested against historical deals we provide. 3. Response and meeting-booking metrics visible on a live dashboard. 4. All components must run without manual intervention beyond initial configuration. Tech flexibility is welcome—whether you prefer Python with scikit-learn, BigQuery pipelines, HubSpot/Zoho integrations, or proprietary tools, just keep the stack transparent and maintainable. Once the machine proves itself, we can discuss scaling it to other Indian cities; right now, Delhi NCR mainly Gurgaon is the only geography that matters....
...Google Sheets) that hit my inbox and a shareable link each morning. • Real-time performance monitoring – live metrics, trend lines, and alert thresholds I can set myself. • Campaign optimization suggestions – surface under-performing ads, wasted spend, or missed keyword opportunities with clear next-step recommendations. Preferred stack is flexible; Google Ads API, GA4 API, Search Console API, BigQuery, Looker Studio or similar are all acceptable as long as the end result is fast and easy to maintain. Authentication should rely on OAuth so I never have to paste tokens manually, and role-based access will keep client views separate from admin views. Acceptance criteria 1. Data from all three platforms populates within 5 minutes of connection. 2. Dai...
...together a set of production-grade data pipelines on Google Cloud Platform and would like an experienced hand to own the build. The job is centred on pipeline development rather than analysis or migration tasks. Scope • Source systems: relational databases we already run in GCP plus batches of parquet and CSV files landing in Cloud Storage. • Destination: curated, partitioned tables in BigQuery with appropriate clustering and cost-efficient storage settings. Work I need from you – Design the end-to-end flow, including ingestion, transformation, and load steps. – Write modular, reusable code or SQL that can be version-controlled and promoted across environments. – Automate orchestration (Cloud Composer, Cloud Functions, or another native...
...(Data 360) Implementation * Design and implement solutions on Salesforce Data Cloud * Configure: * Data Streams (batch & real-time ingestion) * Data Model Objects (DMOs) * Identity Resolution rules - Build Unified Customer Profiles (Golden Records) Data Integration & Ingestion * Integrate data from: * Salesforce clouds (Sales, Service, Marketing) * External systems (APIs, data warehouses, S3, BigQuery, etc.) * Design scalable ingestion pipelines using: * Streaming & batch frameworks * Ensure data normalization and harmonization Identity Resolution & Data Modeling * Design and implement: * Matching & reconciliation rules * Identity graphs * Optimize data models for performance and scalability Segmentation & Activation * Build advanced customer segments u...
I need a BigQuery database created using both the Windsor API and Google Sheets. The cleaned BigQuery database will be linked with Looker Studio for the final output. Requirements: - Integrate data from Windsor API and Google Sheets into BigQuery (Google Sheets native format). - Daily updates of the data. - Perform data processing tasks: - Transformation (Data Sorting, Aggregation, Cleaning) - Validation - Aggregation Ideal Skills and Experience: - Proficiency in Google BigQuery. - Experience with Windsor API integration. - Strong knowledge of Google Sheets. - Data processing and cleaning expertise. - Familiarity with Looker Studio. Please include relevant experience in your application.
I need a BigQuery database created using both the Windsor API and Google Sheets. The cleaned BigQuery database will be linked with Looker Studio for the final output. Requirements: - Integrate data from Windsor API and Google Sheets into BigQuery. - Perform data processing tasks: transformation, validation, and aggregation. Ideal Skills and Experience: - Proficiency in Google BigQuery. - Experience with Windsor API integration. - Strong knowledge of Google Sheets. - Data processing and cleaning expertise. - Familiarity with Looker Studio. Please include relevant experience in your application.
...datasets. Day-to-day you will design and deploy solutions across all three major clouds—AWS, GCP, and Azure—so confidence working in each environment is essential. Our most performance-critical jobs are already running in Spark, and I’m looking for someone who knows how to squeeze every ounce of efficiency from it. If you have Airflow, dbt, Kafka, Kinesis, or warehouse experience with Redshift, BigQuery, or Snowflake, all the better. Native Portuguese is a must, and you’ll collaborate daily in English at a C1/C2 level with distributed stakeholders. The position is full-time and fully remote within Brazil; I’m interested in engineers who enjoy pairing with product managers, setting data quality guardrails, and iterating quickly on business-facing u...
I’m deep into a set of large-scale clinical-trial datasets and need a seasoned partner who can move effortlessly from raw tables to polished Tableau dashboards. The stack is Google Cloud Platform (BigQuery in particular), Amazon Redshift, advanced SQL, and Python, so you should feel at home optimizing long, complex queries, tuning warehouse performance, and scripting repeatable data-quality checks. Where you’ll start I already have a working schema, but it’s straining under growing volume and new KPI requests. You’ll review the existing models, refactor where necessary, and introduce best-practice partitioning and clustering so queries run in seconds, not minutes. From there, we’ll design consumable, high-performance Tableau workbooks that clinicians...
I’m a data analyst working daily with Google BigQuery and Tableau, SQL and to GCP cloud and I’m looking for an experienced mentor who can provide ongoing, hands-on support to help me optimize large-scale SQL queries and improve dashboard performance in real-time projects. My primary need is advanced BigQuery query tuning—profiling execution plans, optimizing partitioning and clustering, and reducing cost and runtime—while also getting guidance on Tableau best practices like efficient layouts, calculated fields, and extract vs. live strategies, along with some light support in Python (Pandas) for data preparation. I’m looking for someone who can work with me through daily, plus deeper weekly sessions, using screen sharing or pair programming, and ...
...batch processing. · Hands-on with Airflow, dbt, or other orchestration tools. · Familiarity with data modeling (OLAP/OLTP), schema evolution, and format handling (Parquet, Avro, ORC). · Experience with hybrid/on-prem and cloud platforms (AWS/GCP/Azure) deployments. · Proficient in working with data lakes/warehouses like Snowflake, BigQuery, Redshift, or Delta Lake. · Knowledge of DevOps practices, Docker/Kubernetes, Terraform or Ansible. · Exposure to data observability, data cataloging, and quality tools (e.g., Great Expectations, OpenMetadata). Shape Good-to-Have · Experience with time-series databases (e.g., InfluxD...
I’m merging every key data source—Shopify, Google Ads, Meta Ads, Snapchat Ads, Adjust (for the mobile app) and our existing GA4 property—into a single reporting flow and need the whole setup built and validated. The work breaks down into two parts. 1. Tracking & data plumbing • Map all Shopify e-commerce events to GA4, making sure purchas...• Data in Looker Studio matches Shopify and ad-platform totals within 2 %. • GA4 debug view shows no duplicate or missing events. • All reports refresh automatically without manual intervention. • A short Loom walkthrough (≤10 min) explains the architecture and how to extend the dashboard. Please outline your estimated timeline and the main tools you’ll rely on (e.g., GTM server-si...
...System Using BigQuery and Machine Learning Job Description: We are seeking an experienced freelancer to develop a Bond Price Research System based on client requirements. Commbined with Machine Learning to enable powerful search, analysis, and insights into bond prices. Project Scope: - Design and build a scalable research platform for searching and analyzing bond price data (government bonds, corporate bonds, etc.). - Store and query large volumes of historical and current bond data in BigQuery. - Integrate Machine Learning models to support price analysis, trend detection, forecasting, and other research features. - Customize the solution according to the client’s specific needs. Key Requirements: 1. Technical Skills (Required): - Strong experience with Goo...
...the transformed results are written back to a target schema (or files, if that proves more efficient). Key points you should know • Source: relational database containing nested JSON / key-value blobs. • Goal: parse, normalize, and flatten these blobs into well-defined columns while preserving relationships and lineage. • Scale: millions of rows, so solutions that leverage Spark, Hadoop, BigQuery, Snowflake, or well-tuned SQL/Python pipelines are welcome—as long as they remain maintainable. Deliverables 1. Transformation code (Python, PySpark, SQL, or Scala) with clear comments. 2. A runnable job definition or workflow file (Airflow DAG, Spark submit script, dbt model, etc.) that shows how to execute the pipeline end-to-end. 3. Simple README exp...
...pipelines (dbt, SQL). Analyze datasets and communicate findings for data-driven decisions. Work with LLM benchmarking and agentic coding workflows. What You Need 4+ years professional experience in DS, ML, or Data Engineering. Expert Python (pandas, NumPy, scikit-learn) and SQL. Proven ability to diagnose ML failure modes and improve model quality. Familiarity with cloud warehouses (Snowflake, BigQuery, or Redshift). Why this role? This is a flexible, remote engagement perfect for those looking to contribute to cutting-edge AI research and work with top-tier industry labs without the commitment of a full-time product role....
...an end-to-end data architecture that can reliably ingest, store, and serve a broad range of clinical-trial assets: patient demographics, clinical trial results, genomic data, COA (Clinical Outcomes Assessment) records, and a growing rater database. What I need from you Design the target architecture and implement the core pipelines—ideally using a modern cloud stack (Snowflake, Databricks, BigQuery, Redshift, or a similar platform; feel free to propose the best fit). Your work should cover raw-to-curated layers, automated metadata capture, and role-based access controls that satisfy typical GxP and HIPAA expectations. Key deliverables • Reference architecture diagram with component rationale • Re-usable ingestion and transformation code (Python, SQL, or...
...help me learn Google Cloud Platform (GCP), Terraform, and DevOps practices from basic to advanced level. The training should be highly practical, hands-on, and focused on real-world implementations, not just theory. Learning Requirements: GCP Topics: GCP fundamentals and architecture IAM (Identity & Access Management) Compute Engine, App Engine, Kubernetes Engine (GKE) Cloud Storage, Cloud SQL, BigQuery Networking (VPC, subnets, firewall rules, load balancing) Monitoring & Logging (Cloud Monitoring, Logging) Security best practices Real-world deployments and use cases Terraform (Infrastructure as Code): Terraform fundamentals and setup Writing and managing Terraform configurations Modules and reusable components Remote state management Provisioning GCP infrastructure usin...
...the bootcamp is hands-on practice: together we will build, deploy, monitor and troubleshoot real ETL flows until I can run them solo with confidence. Primary learning goal My priority is data ingestion, extraction, transformation and storage. We will start with the full ingestion toolkit—DataFlow, Pub/Sub and Cloud Storage—then chain that work into the wider GCP ecosystem: • Storage layers: BigQuery, Bigtable, Cloud Spanner, Cloud SQL, Datastore / Firestore • Transformation & orchestration: DataFusion, DataProc, Cloud Composer, Cloud Scheduler, Cloud Functions • Data quality & cataloging: DataPrep, Data Catalog • Visualisation: Looker, Data Studio Preferred teaching style Live, screen-shared sessions where you demo, then watc...
Project Overview: Over the past few days, I’ve noticed an unexplained surge in event counts in Google Analytics 4, far exceeding our usual traffic trends. The spike likely comes from either aggressive bot traffic or an issue with a tag that’s firing too often. I haven’t made any changes to...the problem. Deliverables: • A concise audit report pinpointing the misbehaving events and their origin • A clean, updated GTM/GA4 configuration (corrected triggers, bot filters, or exclusion rules) • 24-hour validation to ensure the new data is accurate • Brief hand-off notes so I can maintain the fix going forward If you have extensive experience with GA4, GTM, GA Debugger, Looker Studio, or BigQuery, please I can provide immediate access to restore tru...
...transform and load it into our analytical layer. Along the way, I expect you to monitor, tune and refactor the existing Spark jobs so they keep performing efficiently as volumes grow. Because all orchestration happens in Airflow, you should be comfortable creating clear, idempotent DAGs, setting up alerts and handling retries gracefully. Strong knowledge of GCP services such as Cloud Composer, BigQuery and IAM roles will make your life—and mine—easier. I am based in Bangalore; if you are too, that’s a plus for occasional white-board sessions. Telugu fluency is another nice extra, though not mandatory. Deliverables • Production-ready Spark jobs written in Java and deployed on GCP • Airflow DAGs with parameterised configs, logging and alerting ena...
...human-resource information—up to date and error-free. The focus is on resource tracking: I need to see who is available, who is booked, and how utilisation is trending without having to touch a spreadsheet. Here’s the picture: • Source data lives in Excel/CSV files (and, if you can integrate it, a small HRIS API). • I want it validated, normalised, and stored in a structured repository—SQL, BigQuery, even Google Sheets if that’s the simplest. • A refreshed report or dashboard should surface the key tracking metrics for managers automatically, on a schedule we set together. You’re free to choose the tooling—Python, Google Apps Script, Power Automate, Zapier, or something comparable—so long as the workflow is maintaina...
...to jump on a two-to-three-hour working session with me today at 7:30 PM IST. During this slot we will connect marketing, sales and fulfilment data sources, automate key workflows and leave the foundations of a clear reporting layer in place. The live work will revolve around Fivetran for ingestion, Zapier for task automation and orchestration on AWS. We will be pulling events from GA4, GTM and BigQuery, unifying behavioural data from PostHog and HYROS, and blending it with revenue and engagement signals flowing in from ActiveCampaign, Webflow, Facebook Pixel, Google Ads tags, Vidalytics, Typeform, HubSpot CRM, Zoom, Stripe, Hotmart, ClickUp and Circle. Once everything lands in the warehouse I want to shape it with dbt, validate the models in SQL and surface quick insights that p...
...trade-offs while we shape the schema. Your task is to examine the source data I provide, recommend the most appropriate modelling paradigm, then design and implement a clean, scalable structure. Strong familiarity with relational techniques (ER modelling, normalisation, indexing strategies) alongside experience with denormalised or document-oriented patterns would be ideal. If a cloud platform such as BigQuery, Redshift, or Azure SQL turns out to be the right fit, I’ll need your help setting up the environment as well. Deliverables will be: • A clear conceptual and logical model (diagrams or equivalent) • DDL or provisioning scripts ready to deploy in the chosen system • Brief documentation outlining key design choices and future-growth considerations...