Job Description
Responsibilities
- Own and operate the centralized log management platform — ingestion, parsing, structured logging standards, and retention across all services.
- Build and tune alerting with tiered thresholds — catching real problems early while minimizing noise and alert fatigue.
- Perform log analysis across Linux and Windows systems to diagnose incidents and surface root causes.
- Drive MTTD under 15 minutes through better signals, dashboards, and runbooks.
Service-Level Optimization & Reliability
- Identify, diagnose, and optimize service latency and inefficiency across edge, application, and backend layers — profile before guessing, measure every fix.
- Define, implement, and own SLOs, SLIs, and error budgets for critical customer-facing services, and drive improvements against them.
- Build deep observability with Datadog — APM, dashboards, monitors, log management, and SLO tracking.
- Lead reliability and performance root-cause analysis and drive durable fixes.
- Support load testing and capacity planning — identify breaking points before traffic growth causes production issues.
Operations & Cross-Team Support
- Participate in a shared on-call rotation with solid runbooks and blameless post-incident reviews.
- Continuously reduce manual toil through automation and better tooling.
- Collaborate with product squads to instrument services, define meaningful SLIs, and surface the right signals.
- Document runbooks, dashboards, and operational procedures so any engineer can respond to incidents with clear guidance.
What You Bring
- 5–7 years of professional engineering experience, with at least 3 years in SRE, Platform Engineering, or strongly reliability-focused DevOps work.
- Strong hands-on experience with centralized log management platforms — ingestion, parsing, structured logging, and retention using ELK/OpenSearch, Datadog Logs, Loki, or similar.
- Able to diagnose incidents through log analysis across both Linux and Windows environments, isolating root causes under pressure.
- Designs alerting systems with tiered thresholds that minimize noise while catching real problems early, with clear escalation paths.
- Proficient with Datadog across APM, dashboards, monitors, log management, and SLO tracking.
- Experienced defining and operating SLOs, SLIs, and error budgets for customer-facing services.
- Can isolate and resolve latency and inefficiency across edge, application, and backend layers — profiles before guessing, measures every fix.
- Comfortable scripting in Python, Bash, or Go to automate alerting, diagnostics, and toil reduction.
- Hands-on experience operating services on Kubernetes, EKS preferred.
- Familiar with Infrastructure-as-Code tooling such as Terraform for collaboration with the DevSecOps team.
- Evidence-driven — profiles and measures before guessing; validates every optimization against before/after data.
- Reliability-oriented — treats detection speed and signal quality as first-class engineering problems.
- Good communicator — works with product squads to define SLIs and explains reliability constraints clearly.
X-Factor: AI-Native Engineering
- You actively use modern AI agentic workflows daily — not limited to Copilot autocomplete.
- You are proficient with Claude Code, Cursor, Windsurf, or equivalent tools.
- You are comfortable with project-level AI configuration (CLAUDE.md, rules files), agentic task delegation, and AI-driven code review.
- You think in terms of 5–10x productivity through AI-augmented development — and you can demonstrate it.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru