All Jobs
No items found.
Senior Data Engineer
Europe
Remote
Who We Are

Participate in AI incubator projects to scout, incubate, and validate client and PwC-internal ideas on a 3–5year horizon; develop technology roadmaps and prototypes that deliver advanced client solutions; champion internally generated concepts; and continually explore, test, and demonstrate cuttingedge AI to create new products, services, and capabilities.

Role Overview:

The Data Engineer will support the design, build, and operation of reliable, production-ready data pipelines for large-scale data and ML use cases. This role requires a detail-oriented engineer with a strong Data Science background who can bridge raw data ingestion, transformation, quality assurance, and downstream machine learning workflows.

The ideal candidate has hands-on experience building scalable data pipelines and supporting the end-to-end Data Science and ML lifecycle, from raw ingestion and hardening through production deployment, monitoring, and continuous improvement.

This role combines strong data engineering discipline with applied Data Science enablement, with particular emphasis on NLP/NLU use cases, Azure data and AI services, Databricks/PySpark processing, CI/CD, and MLOps practices.

Key Responsibilities

1. Build, harden, and maintain reliable ingestion pipelines that move raw data into production-ready data platforms.

2. Design and implement ETL/ELT processes for data cleaning, transformation, schema design, data quality validation, and lineage tracking.

3. Develop large-scale data processing and transformation workflows using Databricks and PySpark.

4. Support Data Science and ML teams by preparing high-quality datasets, features, and pipelines for NLP, NLU, and broader machine learning use cases.

5. Operate and integrate Azure data and AI infrastructure, including Blob Storage, databases, compute resources, MLflow, and Azure AI/ML services.

6. Implement CI/CD and MLOps practices, including automated deployment, testing, monitoring, reliability checks, and promotion gates.

7. Ensure pipeline reliability, observability, performance, and production readiness across the full data and ML lifecycle.

8. Own data quality, traceability, and operational handover so that data products and ML workflows can be trusted in production.

Required Skills

a. Strong Data Engineering experience, including ETL/ELT, pipeline development, data cleaning, transformation, schema design, and orchestration.

b. Hands-on experience building reliable, production-ready data ingestion pipelines with appropriate hardening, validation, monitoring, and operational controls.

c. Strong Python and PySpark skills for large-scale data processing, transformation, and pipeline development.

d. Practical experience with Databricks for scalable data engineering, distributed processing, and production pipeline implementation.

e. Hands-on Azure infrastructure experience, including Blob Storage, databases, compute resources, MLflow, and Azure AI/ML services.

f. Solid understanding of data quality, lineage tracking, observability, reliability, and production readiness for data pipelines.

g. Strong Data Science and ML lifecycle awareness, with the ability to support model development, experimentation, deployment, monitoring, and continuous improvement.

h. Experience supporting NLP and NLU use cases, including preparation of high-quality datasets and pipelines for natural language applications.

Preferred Skills

a. Experience implementing CI/CD and MLOps practices, including automated testing, deployment, monitoring, promotion gates, and release reliability.

b. Ability to bridge Data Engineering and Data Science teams, translating ML workflow needs into reliable data products and production pipelines.

c. Experience with enterprise-grade ML platforms, feature pipelines, experiment tracking, model governance, or production ML operations.

d. Familiarity with regulated or complex enterprise environments where data quality, traceability, security, and operational resilience are critical.

Remote vs Onsite: Fully remote, with possible occasional in person team sessions / workshops / gatherings (i.e. 1x quarter) likely to take place in Prague

US Hours overlap needed: Minimum 2-6pm CET, preferred 2-8pm CET

Role Description
We Expect You to Have:

Apply for this position

Our team will review your application within the next 5 days.

Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Send

Thank you!
We will be in touch shortly

kid giving a thumbs-up while sitting at a desktop table
Done
Oops! Something went wrong while submitting the form.