Member of Technical Staff - ML Performance
@ ModalMember of Technical Staff - ML Performance
This job is still taking applications, but it's been up a while.
About the job
Modal builds infrastructure for AI workloads, serving top clients like Ramp and DoorDash. Raising funds and growing rapidly, they focus on GPU access, low-latency inference, and open-source contributions. The role involves enhancing ML system performance.
Requirements
- 5+ years high-performance coding
- Experience with torch and ML frameworks
- Knowledge of Nvidia GPU and CUDA
- ML performance engineering skills
- Familiarity with inference engines
Qualifications
- Strong programming skills
- Experience in ML system optimization
- GPU architecture knowledge
- Problem-solving in performance tuning
- Open-source contributions preferred
Full job description
About Us:
AI needs a new infrastructure layer. We're building it at Modal.
Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.
Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.
Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.
The Role:
We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!
Requirements:
5+ years of experience writing high-quality, high-performance code.
Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
Familiarity with Nvidia GPU architecture and CUDA.
Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).
Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).
Similar jobs in New York, New York
- C
Member of Technical Staff
Cockroach Labs · New York, NY, United States
Posted 4 days ago - A
Member of Technical Staff, Causality
Ataraxis AI · New York HQ
Posted 3 weeks ago - A
Member of Technical Staff, Survival Analysis
Ataraxis AI · New York HQ
Posted 3 weeks ago - V
Member of the Technical Staff - Next.js
Vercel · New York, New York, United States
Posted 2 days ago - A
Member of Technical Staff, Bayesian Statistics
Ataraxis AI · New York HQ
Posted 3 weeks ago - A
Staff Technical Architect
Appnovation Technologies · New York, New York, United States
Posted today