Senior Site Reliability engineer
@ FilevineSenior Site Reliability engineer
About the job
Filevine is a Legal AI company delivering Legal Operating Intelligence, transforming legal work with a unified platform and AI-powered insights to improve efficiency and reliability.
Requirements
- 8+ years software engineering experience
- Experience with cloud infrastructure
- Proficiency in Python, Go, Bash
- Knowledge of Kubernetes/AWS
- Leadership in incident response
Qualifications
- Bachelor's degree or equivalent
- Strong troubleshooting skills
- Ability to mentor engineers
- Experience applying AI/ML in operations
- Effective communication skills
Full job description
Responsibilities
Observability & Alerting: Design and improve monitoring, logging, tracing, dashboards, and SLI/SLOs for production visibility.
Automation & CI/CD: Build internal tools and delivery pipelines to boost efficiency, eliminate toil, and ensure reliable deployments.
Reliability & System Quality: Drive continuous improvements in system performance, scalability, and security to mitigate customer impact.
Incident Management: Lead production incident response from triage to resolution, turning lessons into durable runbooks and preventatives.
Technical Leadership: Guide major technical initiatives, align cross-team engineering efforts, and manage technical risks.
Mentorship: Elevate SRE team capability through design reviews, paired problem-solving, and incident post-mortems.
Operations & On-Call: Join the on-call rotation while leading capacity planning and disaster recovery readiness.
AI/ML Operationalization: Leverage operational AI/ML tools to forecast capacity risks, detect patterns, and automate system health.
Qualifications
Experience: 8+ years in software/platform engineering or DevOps, including 5+ years dedicated to Site Reliability Engineering.
Infrastructure & Observability: Expertise in cloud platforms (AWS), Kubernetes, IaC, and full-stack observability (tracing, logging, SLI/SLOs). Grafana, Datadog, NewRelic
Automation & Scripting: Proficient in Python, Go, or Bash for building CI/CD pipelines, production tools, and toil-reducing automation.
Incident Leadership: Proven track record in root cause analysis, high-severity incident response, and long-term reliability engineering.
Leadership & Mentorship: Strong communication skills with a history of mentoring engineers, leading cross-functional projects, and setting technical strategy.
Operational AI/ML: Hands-on experience using AI/ML on telemetry data to predict capacity risks, spot anomalies, and optimize system workflows.
Similar jobs in Columbus, OH
- J
HVAC Controls Engineer – Building Automation
Jobot · Columbus, OH, United States
Posted 5 days ago - A
Site Strategy Manager, Data Center Sites - Remote (U.S.)
AECOM · Columbus, OH, United States
Posted 2 weeks ago - A
Senior Product Test Engineer
Anduril Industries · Ashville, Ohio, United States
Posted 3 weeks ago - A
Senior Civil / Environmental Engineer
AECOM · Columbus, OH, United States
Posted 2 weeks ago - A
Senior Product Quality Engineer
Anduril Industries · Ashville, Ohio, United States
Posted 3 weeks ago - S
Seeking experienced babysitter for toddler with own car & reliability!
Sittercity · Columbus, OH, United States
Posted 2 days ago