ML-Ops / Platform Engineer

Long Finch Technologies

Long Finch Technologieslongfinchtechnologies.com

ML-Ops / Platform Engineer

Paw Creek, North Carolina
Posted 1 week ago

About the job

Join a team focused on building cloud-native applications and GenAI platforms on AWS and Azure. Ensure high availability, security, and scalability while enabling advanced AI integrations and infrastructure.

Requirements

  • Experience with cloud-native applications
  • Kubernetes and Docker expertise
  • DevOps and automation skills
  • Knowledge of AI platforms and models
  • Proficiency in Python or Java

Qualifications

  • Bachelor's degree preferred
  • Strong problem-solving skills
  • Team collaboration ability
  • Experience with cloud security
  • Understanding of microservices

Full job description

  • Strong hands-on experience building and operating cloud-native applications and platforms on AWS and Azure, including VPC/VNet, IAM, Load Balancers, API Gateway, Lambda/Functions, EKS/AKS, ECS, App Services, Storage, Key Vault/Secrets Manager, and cloud networking.
  • Deep expertise in AWS and Azure architecture, including multi-account/subscription design, landing zones, cloud security, high availability, disaster recovery, scalability, and cost optimization.
  • Experience enabling and operationalizing Enterprise GenAI platforms, including model onboarding, AI gateways, inference platforms, vector databases, RAG, guardrails, and AI application enablement.
  • Strong knowledge of Azure AI Foundry, Azure OpenAI, AWS Bedrock, SageMaker, model serving, embeddings, prompt engineering, AI evaluations, and agentic frameworks.
  • Expert-level experience with Docker, Kubernetes/OpenShift, AKS, EKS, service mesh, ingress controllers, autoscaling, and multi-cluster platform operations.
  • Strong DevOps and Platform Engineering experience, including GitHub Actions, Azure DevOps, Jenkins, GitOps, ArgoCD, Terraform, Ansible, automated testing, and release management.
  • Proficiency in Python, Java, REST APIs, microservices, and distributed systems development.
  • Experience with MongoDB, Redis, PostgreSQL, Vector Databases, caching strategies, state management, and high-throughput data platforms.
  • Strong understanding of cloud security, IAM, secrets management, observability, monitoring, logging, SRE practices, and production support.
  • Experience troubleshooting and optimizing cloud-hosted, containerized workloads for performance, resiliency, scalability, and cost efficiency.
  • Ability to partner with application, platform, infrastructure, and security teams to accelerate GenAI and cloud modernization initiatives.
Show full description
Ask Duey about this job

Duey AI may make mistakes. Please double-check key information.