Engineer

September 2, 2026
Application ends: December 1, 2026
Apply Now

Job Description

Responsibilities

Observability and Datadog Ownership

  • Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration.
  • Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases.
  • Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product.
  • Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook.

Automation and Toil Elimination

  • Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting.
  • Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free.
  • Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning.

Security, Documentation, and Mentorship

  • Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning.
  • Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response.

Skills and Experience Needed

  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems.
  • Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs.
  • Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate.
  • Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures.
  • Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud.
  • Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred).
  • Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews.
  • Experience operating multi-region, multi-tenant systems.
  • Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation.
  • Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals.
  • Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints.
  • Nice to have: experience migrating from a legacy monitoring stack to a consolidated observability platform; relevant Azure, Kubernetes, or Datadog certifications.

Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru