Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

Snowflake
Bellevue
Workplace: HybridFull timeUSD 200,000 - 287,500 annuallyFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Reliability focus","Performance orientation","Experimental mindset","Low-ego collaboration"]

Build and scale Snowflake’s Cortex Training platform so customers can run demanding post-training LLM workloads inside Snowflake. You’ll design and ship full-stack components spanning public training APIs/SDKs and the control plane to the GPU data plane, powering multi-tenant scheduling, orchestration, and fault-tolerant distributed training and inference. Partner with research to productionize state-of-the-art training techniques into reliable, composable services at enterprise scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Snowflake
Snowflake
1 day ago

Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live
Reposted: similar role first listed 3 months ago

Job Summary

Build and scale Snowflake’s Cortex Training platform so customers can run demanding post-training LLM workloads inside Snowflake. You’ll design and ship full-stack components spanning public training APIs/SDKs and the control plane to the GPU data plane, powering multi-tenant scheduling, orchestration, and fault-tolerant distributed training and inference. Partner with research to productionize state-of-the-art training techniques into reliable, composable services at enterprise scale.
Location: Bellevue
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and build across the full stack, from public training APIs and SDKs through the control plane to the GPU data plane.
  • •Scale distributed systems for GPU compute serverless, including multi-tenant scheduling, placement, capacity-aware routing, and fault tolerance.
  • •Drive end-to-end performance at scale by keeping training, inference, and RL loops fast and the data plane responsive under heavy concurrent load.
  • •Productionize research building blocks by partnering with Snowflake Research to turn training and inference techniques into reliable, composable components.
  • •Ensure GPUs stay saturated by optimizing throughput, reliability, and cost efficiency for concurrent workloads.

Pay and Benefits

Salary: USD 200,000 - 287,500 annually

Key Requirements

  • •3+ years (Intermediate) or 6+ years (Senior) building and shipping production ML systems.
  • •Strong distributed systems and infrastructure foundation, including designing scalable fault-tolerant services.
  • •Operate production services on Kubernetes and have hands-on debugging across data, infrastructure, and GPU layers.
  • •Familiarity with GPU and LLM infrastructure (e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM).
  • •BS in Computer Science or a related field (MS/PhD a plus).
Experience:3+ yearsML systemsLLMDistributed systemsGPU infrastructure
Education:Bachelor's
Skills:Reliability focusPerformance orientationExperimental mindsetLow-ego collaboration
Tech Stack:KubernetesPyTorchDeepSpeedFSDPRayCUDANCCLVLLMGPUsLLMsDistributed systemsSchedulingOrchestrationMulti-node trainingInferenceFault toleranceThroughputServerless

Company Brief

Snowflake
Provides a cloud-native data platform for data warehousing, data engineering, data science, and analytics, enabling organizations to store, process, and share large-scale data across multiple cloud providers with elastic scalability and performance.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Bozeman, United States
Founded: 2012
WebsiteLinkedIn