GCP Data Engineer
France (Paris/Toulouse)
We are looking for an experienced GCP Data Engineer to join a large-scale cloud data engineering programme, helping design, build and optimise enterprise data pipelines and distributed processing solutions on Google Cloud Platform (GCP).
The position requires strong hands-on experience with Java and Apache Spark, alongside a solid understanding of modern GCP data services. You will work within an engineering-focused environment, developing scalable data processing applications and pipelines that support downstream analytics, reporting and AI/ML use cases.
Design, develop and maintain scalable data pipelines and distributed data processing applications on GCP.
Build high-performance data processing solutions using Java and Apache Spark.
Develop and optimise Spark workloads handling large-volume and complex datasets.
Design batch and, where required, real-time/streaming data processing architectures.
Work with GCP services including BigQuery, Dataflow, Pub/Sub, Cloud Storage and Dataproc.
Develop robust ingestion, transformation and integration processes across structured, semi-structured and unstructured data.
Optimise Spark jobs, SQL queries and data pipelines for performance, scalability and cost efficiency.
Design appropriate data models and storage patterns for analytical and operational requirements.
Implement data quality, validation, monitoring, logging and error-handling mechanisms.
Develop reusable engineering components, frameworks and APIs using Java.
Support the deployment and operation of data applications across development, testing and production environments.
Contribute to CI/CD and DevOps practices, including automated testing, deployment and infrastructure management.
Work closely with Data Architects, Data Scientists, Software Engineers and business stakeholders to translate requirements into technical solutions.
Participate in code reviews and contribute to engineering standards around clean code, testing, documentation and software design.
Troubleshoot complex production issues across distributed data processing environments.
Essential:
Strong commercial experience as a Data Engineer / Big Data Engineer / GCP Data Engineer.
Advanced development skills in Java.
Strong hands-on experience with Apache Spark, including Spark SQL and distributed data processing.
Strong knowledge of Google Cloud Platform and its data ecosystem.
Experience with BigQuery for large-scale analytical workloads.
Experience with Dataproc and/or Dataflow.
Experience with Google Cloud Storage (GCS).
Understanding of Pub/Sub and event-driven or streaming architectures.
Strong SQL skills and experience working with large datasets.
Knowledge of data modelling, ETL/ELT patterns and modern data architecture.
Understanding of distributed computing principles, partitioning, parallel processing and Spark performance optimisation.
Experience with Git and CI/CD pipelines.
Highly Desirable:
Python and/or PySpark experience alongside Java.
Experience with Cloud Composer / Apache Airflow.
Experience with Kafka or other streaming technologies.
Infrastructure-as-Code experience with Terraform.
Containerisation using Docker and Kubernetes/GKE.
Experience with Cloud Run, Cloud Functions or other serverless GCP services.
Knowledge of Data Lake, Lakehouse and Data Mesh architectures.
Experience supporting downstream Machine Learning, AI or advanced analytics workloads.
Familiarity with monitoring and observability across production data platforms.