ML-Ops / Platform Engineer

@ Long Finch Technologies
Long Finch Technologieslongfinchtechnologies.com

ML-Ops / Platform Engineer

Charlotte, North Carolina
Posted 1 week ago

About the job

A company specializing in cloud-native solutions, AI platform enablement, and cloud infrastructure. The role focuses on building and operating scalable, secure platforms on AWS and Azure, supporting GenAI and enterprise applications.

Requirements

  • Cloud-native application experience
  • Kubernetes and Docker expertise
  • DevOps and automation skills
  • AI platform operational knowledge
  • Proficiency in Python/Java

Qualifications

  • Strong cloud architecture background
  • Experience with AI models and frameworks
  • Knowledge of cloud security practices
  • Prior experience in enterprise environments

Full job description

Must have skills: MLOps, AWS/Azure, Kubernetes, Docker, Python, Terraform, CI/CD, GenAI/LLM, RAG, Bedrock/Azure OpenAI, Vector DB, Observability

Responsibilities:

· Designed, deployed, and operated enterprise-grade MLOps/GenAI platforms across AWS and Azure, leveraging AWS Bedrock, SageMaker, Azure OpenAI, Azure AI Foundry, model serving, embeddings, RAG, vector databases, AI gateways, guardrails, and agentic frameworks.

· Built and automated cloud-native ML infrastructure using Terraform, Kubernetes (EKS/AKS/OpenShift), Docker, GitHub Actions, Azure DevOps, Jenkins, GitOps, and ArgoCD, enabling scalable model deployment, CI/CD, versioning, and release management.

· Implemented secure and highly available ML/AI platforms using AWS IAM, Azure IAM/RBAC, VPC/VNet, Key Vault, Secrets Manager, API Gateway, load balancers, ingress controllers, service mesh, autoscaling, and multi-account/subscription architectures.

· Developed and operationalized ML/GenAI workloads using Python, REST APIs, microservices, MongoDB, PostgreSQL, Redis, and vector databases, implementing model evaluation, prompt engineering, state management, caching, and high-throughput inference capabilities.

· Monitored, troubleshot, and optimized production ML/AI workloads using observability, logging, monitoring, SRE practices, performance tuning, resiliency, disaster recovery, and cost optimization, while collaborating with application, platform, infrastructure, and security teams.

Show full description