Job Description
Responsibilities
The Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help teams ship secure, reliable software quickly. You will build and operate scalable microservices and the infrastructure behind reliable, secure APIs, improve the software delivery experience, and partner with Product and Engineering teams to design, test, deploy, and operate production systems.
You will also help teams safely adopt and operate AI-enabled tooling and workflows by engineering features for and maintaining autonomous coding-agent orchestration platform. You will understand the operational characteristics and risks of AI services, integrate them into developer and operational workflows, and ensure appropriate reliability, observability, access controls, auditability, and cost management. You will troubleshoot production issues, document operational standards, and continuously improve the infrastructure and developer experience that enable teams to move quickly with confidence.
What you’ll work on
- Build and operate scalable microservices, platform capabilities, and shared engineering libraries.
- Design, deploy, and maintain secure, reliable cloud infrastructure and Kubernetes-based services.
- Improve CI/CD pipelines, developer workflows, and software delivery automation across engineering teams.
- Engineer and operate autonomous coding-agent orchestration platform and AI-enabled developer tooling.
- Implement observability, monitoring, incident response, and operational best practices for production systems and AI services.
- Partner with Product and Engineering teams to design resilient architectures and improve platform reliability.
- Continuously improve security, access controls, auditability, and cost optimization across cloud infrastructure and AI-powered workflows.
What you’ll bring to
Core Requirements
- 3–6 years of professional software development experience using languages such as Golang, Java, JavaScript/TypeScript, Python, or Rust.
- Hands-on experience building and operating cloud-native applications on AWS or GCP, including Kubernetes-based infrastructure.
- Experience engineering and operating production workflow orchestration systems, developer platforms, or autonomous agent solutions.
- Experience integrating AI capabilities through APIs, SDKs, or workflow tooling while applying operational guardrails such as observability, access controls, auditability, evaluation, monitoring, and human approval where appropriate.
- Strong understanding of distributed systems, RESTful API design, SQL database design, schema modeling, and query optimization.
- Demonstrated commitment to writing clean, maintainable, well-tested code that supports continuous delivery and operational excellence.
- Bachelor’s degree in Computer Science or a related technical field, or equivalent practical experience.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru