Job Description
Responsibilities
- Design and operate the SRE practice for Managed oferings, including on-call processes, SLA frameworks, incident response playbooks, and post-incident review (PIR) processes.
- Build and maintain observability infrastructure: centralised logging (correlation IDs), metrics dashboards, distributed tracing, and alerting for the Predator/Instinct platform stack.
- Define and track SLOs (Service Level Objectives) and error budgets for real-time transaction processing pipelines, targeting high TPS and low round-trip latency.
- Manage cloud infrastructure provisioning and configuration using IaC tooling (Terraform, Helm), supporting both AWS/Azure cloud deployments and on-premises customer environments.
- Implement and maintain CI/CD pipelines for GFS solutions (Jenkins, etc.)
- Work with Engineering teams to ensure security and compliance readiness for Managed services ā including PCI DSS, ISO 27001, SOC 1/2/3, PDPA/GDPR ā in close coordination with InfoSec teams.
- Drive platform resilience improvements: high availability, auto-scaling, disaster recovery, backup/restore procedures, and chaos engineering practices.
- Manage secrets, certificate rotation, identity/access controls (OAuth/RBAC), and vulnerability management for the hosted environment.
- Support performance testing methodology and baseline establishment for our products.
- Contribute to the Architecture Review Committee (ARC) with SRE and operational perspectives on technology choices.
- Collaborate with engineering squads to embed reliability and DevSecOps practices across the SDLC.
Are you interested in this position?
Apply by clicking on the āApply Nowā button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru