Senior Site Reliability Engineer – AI & Automation.

@ Neshent Technologies
Neshent Technologiesneshenttechnologies.com

Senior Site Reliability Engineer – AI & Automation.

Fontainebleau, Florida
Posted today

About the job

Join our team as a Senior Site Reliability Engineer to design AI solutions, automate operations, support cloud infrastructure, and improve system reliability at scale.

Requirements

  • 5+ years SRE or DevOps experience
  • Strong Python scripting skills
  • Experience with AI agents and LLMs
  • Knowledge of cloud platforms AWS, Azure, GCP
  • Hands-on with Terraform, CloudFormation, CDK

Qualifications

  • Experience supporting production systems
  • Excellent communication skills
  • Ability to mentor engineers
  • Experience in incident response and troubleshooting

Full job description

Primary Responsibilities

*Design and implement AI agents, LLM-based solutions, and operational automation for SRE and DevOps environments.
*Support production systems, perform incident response, troubleshooting, and 24x7 operational triage.
*Develop automation and integrations using Python.
*Build and maintain monitoring and observability solutions using Splunk, Datadog, New Relic, or AppDynamics.
*Design and manage cloud infrastructure across AWS, Azure, and/or GCP.
*Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, and/or AWS CDK.
*Apply AI/ML, prompt engineering, and agentic AI concepts to operational use cases.
*Develop runbook automation, anomaly detection, and intelligent incident management capabilities.
*Collaborate with engineering, operations, and business stakeholders to improve system reliability and operational processes.
*Mentor engineers and help establish SRE, DevOps, automation, and observability best practices.

Required Skills

*5+ years of SRE, DevOps, or Site Reliability Engineering experience.
*Strong Python development skills for automation, tooling, and integrations.
*Experience with AI agents, LLMs, and agentic AI frameworks.
*Experience with Splunk, Datadog, New Relic, AppDynamics, or similar observability platforms.
*Hands-on experience with AWS, Azure, and/or GCP.
*Experience with Terraform, CloudFormation, and/or CDK.
*Strong production support and incident response experience.
*Knowledge of AI/ML, prompt engineering, and operational AI automation.
*Strong communication, stakeholder management, and mentoring skills.
*Ability to work effectively in a 24x7 operational environment.

Preferred Skills

*Experience with Adobe Experience Manager (AEM).
*Experience supporting CMS platforms and digital applications.
*Knowledge of incident management, runbook automation, and anomaly detection.
*Familiarity with Atlassian Rovo, AI operational tooling, and modern observability platforms.

Show full description