Job Description
Responsibilities
- Embed with product teams, helping them improve their operational maturity by setting up and refining practices around on-call, monitoring, alerting, and run books
- Run regular game day exercises with product teams, helping them feel more prepared to investigate and quickly remediate incidents
- Jump into application code written in Haskell & TypeScript and implement reliability techniques such as retries, better error handling, better logging, circuit breaking, etc
- Steer SLOs towards meaningful customer outcomes that product teams are accountable for. Those SLOs become a strong signal for whether the product is working as intended
- Champion reliability practices through design document and code reviews
- Identify observability gaps that hinder debugging, incident response, and business intelligence and help close those
- Advocate for longer-term improvements that non-product engineering teams can drive
- Participate in the product team’s on-call rotation while embedding or as part of a more general engineering rotation, helping us improve processes and how we learn from incidents
The ideal candidate for the role:
- Has past Site Reliability Engineering or DevOps experience
- Has measurable examples of influencing an organization towards greater reliability
- Has significant experience with PostgreSQL
- Has authored and operated Temporal workflows
- Has experience with observability platforms like Grafana or Honeycomb
- Has familiarity with OpenTelemetry
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru