Job Description
Responsibilities
Observability and Datadog Ownership
- Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration.
- Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases.
- Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product.
- Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook.
Automation and Toil Elimination
- Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting.
- Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free.
- Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning.
Security, Documentation, and Mentorship
- Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning.
- Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response.
Skills and Experience Needed
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems.
- Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs.
- Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate.
- Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures.
- Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud.
- Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred).
- Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews.
- Experience operating multi-region, multi-tenant systems.
- Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation.
- Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals.
- Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints.
- Nice to have: experience migrating from a legacy monitoring stack to a consolidated observability platform; relevant Azure, Kubernetes, or Datadog certifications.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru