ML-Ops / Platform Engineer

@ Long Finch Technologies
Long Finch Technologieslongfinchtechnologies.com

ML-Ops / Platform Engineer

Mint Hill, North Carolina
Posted 1 day ago

About the job

Join a team focused on building cloud-native applications and GenAI platforms on AWS and Azure. Ensure high availability, security, and scalability while enabling advanced AI integrations and infrastructure.

Requirements

  • Experience with cloud-native applications
  • Kubernetes and Docker expertise
  • DevOps and automation skills
  • Knowledge of AI platforms and models
  • Proficiency in Python or Java

Qualifications

  • Bachelor's degree preferred
  • Strong problem-solving skills
  • Team collaboration ability
  • Experience with cloud security
  • Understanding of microservices

Full job description

  • Strong hands-on experience building and operating cloud-native applications and platforms on AWS and Azure, including VPC/VNet, IAM, Load Balancers, API Gateway, Lambda/Functions, EKS/AKS, ECS, App Services, Storage, Key Vault/Secrets Manager, and cloud networking.
  • Deep expertise in AWS and Azure architecture, including multi-account/subscription design, landing zones, cloud security, high availability, disaster recovery, scalability, and cost optimization.
  • Experience enabling and operationalizing Enterprise GenAI platforms, including model onboarding, AI gateways, inference platforms, vector databases, RAG, guardrails, and AI application enablement.
  • Strong knowledge of Azure AI Foundry, Azure OpenAI, AWS Bedrock, SageMaker, model serving, embeddings, prompt engineering, AI evaluations, and agentic frameworks.
  • Expert-level experience with Docker, Kubernetes/OpenShift, AKS, EKS, service mesh, ingress controllers, autoscaling, and multi-cluster platform operations.
  • Strong DevOps and Platform Engineering experience, including GitHub Actions, Azure DevOps, Jenkins, GitOps, ArgoCD, Terraform, Ansible, automated testing, and release management.
  • Proficiency in Python, Java, REST APIs, microservices, and distributed systems development.
  • Experience with MongoDB, Redis, PostgreSQL, Vector Databases, caching strategies, state management, and high-throughput data platforms.
  • Strong understanding of cloud security, IAM, secrets management, observability, monitoring, logging, SRE practices, and production support.
  • Experience troubleshooting and optimizing cloud-hosted, containerized workloads for performance, resiliency, scalability, and cost efficiency.
  • Ability to partner with application, platform, infrastructure, and security teams to accelerate GenAI and cloud modernization initiatives.
Show full description