Job Description
As a DevOps Engineer in Data Platform Unit, you will play a key role in shaping the reliability, scalability, and efficiency of the data platforms that enable technology and business teams across the bank to build data-driven solutions. You will combine software engineering, cloud infrastructure, automation, and observability expertise to build resilient systems and streamline operations for a modern data platform ecosystem leveraging Kubernetes, Kafka, Flink, Spark, Iceberg, AWS services, and observability tooling such as Grafana and Prometheus.
Working closely with cross-functional teams, you will champion DevOps best practices, enhance platform performance and reliability, and help foster a culture of operational excellence.
How You Will Make an Impact:
- Develop software tools and automation solutions to improve operational efficiency, reliability, and developer experience across DevOps, ITOps, and engineering teams.
- Design, build, and manage cloud-native infrastructure on AWS using Infrastructure as Code and automation best practices.
- Implement and maintain monitoring, alerting, and observability capabilities using tools such as Grafana and Prometheus to ensure platform health, reliability, and performance.
- Create and maintain dashboards, metrics, and alerts that provide actionable insights into platform operations and enable proactive issue detection.
- Proactively identify areas for improvement within platform services, infrastructure, and operational processes, and implement enhancements.
- Support incident response, troubleshooting, root cause analysis, and post-incident reviews to continuously improve platform stability and resilience.
What Makes You a Great Fit:
- Proven experience in a DevOps, Platform Engineering, Site Reliability Engineering (SRE), or similar role, with a strong track record of managing cloud infrastructure and ensuring system reliability, scalability, and performance.
- Strong experience with Amazon Web Services (AWS), including designing, building, and operating cloud-native infrastructure.
- Experience with Infrastructure as Code tools such as Terraform and Terragrunt, as well as CI/CD platforms and GitOps practices using GitLab and Argo CD.
- Hands-on experience with monitoring, observability, and alerting tools such as Grafana, Prometheus, CloudWatch, or similar solutions.
- Familiarity with containerization and orchestration technologies such as Docker and Kubernetes, including managing workloads in Amazon EKS.
- Proficiency in one or more scripting or programming languages such as Python, Java, Bash, or PowerShell, with a focus on automation and operational tooling.
- Experience managing application and infrastructure deployments across production and non-production environments using automated delivery pipelines, Helm charts, and GitOps approaches.
- Experience supporting distributed data platforms and technologies such as Kafka, Flink, Spark, Iceberg, or similar large-scale data processing systems is considered a strong advantage.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru