Engineer

July 22, 2026
Application ends: October 22, 2026
Apply Now

Job Description

Responsibilities

  • Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
  • Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
  • Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
  • Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.

Service-Level Optimization & Reliability

  • Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers — profile before guessing, measure every fix.
  • Define, implement, and own SLOs, SLIs, and error budgets for critical customer-facing services, and drive improvements against them.
  • Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
  • Lead reliability and performance root-cause analysis and drive durable fixes.
  • Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.

Operations & Cross-Team Support

  • Participate in a shared on-call rotation with solid runbooks and blameless post-incident reviews.
  • Continuously reduce manual toil through automation and better tooling.
  • Collaborate with product squads to instrument services, define meaningful SLIs, and surface the right signals.
  • Document runbooks, dashboards, and operational procedures so any engineer can respond to incidents with clear guidance.

What You Bring

  • 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability-focused DevOps work.
  • Strong hands-on experience with centralized log management platforms — ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
  • Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure.
  • Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths.
  • Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
  • Experienced defining and operating SLOs, SLIs, and error budgets for customer-facing services.
  • Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
  • Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
  • Hands-on experience operating services on Kubernetes, EKS preferred.
  • Familiar with Infrastructure-as-Code tooling such as Terraform for collaboration with the DevSecOps team.
  • Evidence-driven — profiles and measures before guessing; validates every optimization against before/after data.
  • Reliability-oriented — treats detection speed and signal quality as first-class engineering problems.
  • Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.

X-Factor: AI-Native Engineering

  • You actively use modern AI agentic workflows daily — not limited to Copilot autocomplete.
  • You are proficient with Claude Code, Cursor, Windsurf, or equivalent tools.
  • You are comfortable with project-level AI configuration (CLAUDE.md, rules files), agentic task delegation, and AI-driven code review.
  • You think in terms of 5–10x productivity through AI-augmented development — and you can demonstrate it.

Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru