Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

Snowflake
Bellevue
Workplace: OnsiteFull timeUSD 200,000 - 287,500 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Collaboration","Curiosity"]

Senior Software Engineer to advance Snowflake’s Cortex Training LLM post-training platform. You will design end-to-end ML infrastructure across APIs, control plane, and GPU data plane; scale multi-tenant GPU scheduling and regional pools; optimize training/inference/ RL loops under heavy load; and productionize research components for enterprise-grade reliability and performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Snowflake
Snowflake
3 months ago

Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Senior Software Engineer to advance Snowflake’s Cortex Training LLM post-training platform. You will design end-to-end ML infrastructure across APIs, control plane, and GPU data plane; scale multi-tenant GPU scheduling and regional pools; optimize training/inference/ RL loops under heavy load; and productionize research components for enterprise-grade reliability and performance.
Location: Bellevue
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
  • •Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
  • •Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
  • •Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.

Pay and Benefits

Salary: USD 200,000 - 287,500 annually

Key Requirements

  • •5+ years building and shipping production ML systems.
  • •Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production.
  • •Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.
  • •Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
  • •BS in Computer Science or a related field (MS/PhD a plus).
Experience:5+ yearsMachine learningDistributed systemsLLMGPUInfrastructure
Education:Bachelor's
Skills:Problem-solvingCollaborationCuriosity
Tech Stack:PythonPyTorchDeepSpeedFSDPRayCUDANCCLVLLMKubernetes

Company Brief

Snowflake
Provides a cloud-native data platform for data warehousing, data engineering, data science, and analytics, enabling organizations to store, process, and share large-scale data across multiple cloud providers with elastic scalability and performance.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Bozeman, United States
Founded: 2012
WebsiteLinkedIn