Senior Site Reliability Engineer

@ PB consulting
PB consultingpbconsulting.com

Senior Site Reliability Engineer

Cramerton, North Carolina
Posted today

About the job

We develop cloud-native platforms focusing on reliability, automation, and security. The role emphasizes Kubernetes, AWS, observability, and AI tools for troubleshooting.

Requirements

  • 10+ years SRE/DevOps experience
  • AWS cloud services expertise
  • Kubernetes and Docker experience
  • Terraform and IaC skills
  • Python and Bash scripting

Qualifications

  • Strong incident management skills
  • Knowledge of SLO/SLI governance
  • Experience with cloud security
  • AI/LLM tools familiarity
  • Operational troubleshooting ability

Full job description

Job Summary

We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS, container platforms, observability, infrastructure automation, and AI-assisted troubleshooting. The role will focus on platform reliability, incident management, SLO/SLI governance, and automation across cloud-native environments.

Roles and Responsibilities
  • Design, maintain, and troubleshoot Kubernetes and container platforms including EKS, ECS, and Docker.
  • Manage AWS services including EC2, ECS, EKS, Lambda, RDS, S3, and IAM.
  • Implement infrastructure automation using Terraform.
  • Monitor and improve platform reliability using Dynatrace, Splunk, and Grafana.
  • Develop automation and troubleshooting scripts using Python and Bash.
  • Lead incident response, root cause analysis, and production issue resolution.
  • Define and govern SLOs, SLIs, and reliability standards.
  • Apply DevSecOps practices across cloud and containerized environments.
  • Leverage AI/LLM tools, including Claude AI, for pipeline troubleshooting and operational problem solving.
  • Improve system availability, performance, scalability, security, and operational efficiency.
Required Skills & Experience
  • 10+ years of experience in SRE/DevOps.
  • Strong experience with AWS cloud services, including ECS, EKS, EC2, Lambda, RDS, S3, and IAM.
  • Hands-on experience with Kubernetes, Docker, EKS, and ECS.
  • Strong experience with Terraform and Infrastructure as Code.
  • Experience with Dynatrace, Splunk, and Grafana.
  • Strong Python and Bash scripting skills.
  • Must have experience using Claude AI or similar AI/LLM tools for pipeline troubleshooting.
  • Strong incident management and production support experience.
  • Experience with SLO/SLI governance and site reliability practices.
  • Strong understanding of DevSecOps, cloud security, automation, and CI/CD.
Show full description