Senior/Staff Software Engineer – LLM Inference & Reinforcement Learning Platform

Snowflake
Bellevue
Full timeUSD 236,000 - 330,000 annuallyFunction: Research & Scientific (R&D)Experience: 5+ yearsSkills: ["Independent problem solving","Cross-functional collaboration","Communication","Experimentation","Performance optimization"]

Build and optimize high-performance LLM inference systems across distributed serving, runtime execution, and GPU performance-critical kernels. Develop new techniques to improve latency, throughput, memory efficiency, and cost, including speculative/parallel decoding, KV-cache management, scheduling and batching, and quantization. Create adaptive inference systems that automatically profile workloads, identify bottlenecks, and tune configurations with minimal manual effort. Collaborate with model researchers and engineering teams to deploy research innovations in production.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Snowflake
Snowflake
3 days ago

Senior/Staff Software Engineer – LLM Inference & Reinforcement Learning Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and optimize high-performance LLM inference systems across distributed serving, runtime execution, and GPU performance-critical kernels. Develop new techniques to improve latency, throughput, memory efficiency, and cost, including speculative/parallel decoding, KV-cache management, scheduling and batching, and quantization. Create adaptive inference systems that automatically profile workloads, identify bottlenecks, and tune configurations with minimal manual effort. Collaborate with model researchers and engineering teams to deploy research innovations in production.
Location: Bellevue
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and develop high-performance LLM inference systems across distributed serving, runtime systems, GPU execution, and performance-critical kernels.
  • •Improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost through novel techniques.
  • •Explore advanced inference methods such as speculative/parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, and KV-cache management.
  • •Develop adaptive inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.
  • •Apply AI-driven and AI-native systems engineering to automate profiling, bottleneck detection, configuration search, experimentation, runtime strategy selection, debugging, and performance tuning.

Pay and Benefits

Salary: USD 236,000 - 330,000 annually

Key Requirements

  • •Bachelor’s degree in Computer Science, Electrical Engineering, or related field; Master’s degree or PhD preferred.
  • •5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
  • •Strong understanding of modern LLM inference architectures and performance tradeoffs for large-scale serving.
  • •Hands-on experience with LLM inference and serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar systems.
  • •Hands-on GPU programming experience using CUDA and Triton (or similar), plus strong skills profiling and optimizing system performance end-to-end.
Experience:5+ yearsLLM inferenceDistributed AI systemsGPU systemsHigh-performance computing
Education:
Skills:Independent problem solvingCross-functional collaborationCommunicationExperimentationPerformance optimization
Tech Stack:LLM inferenceVLLMSGLangTensorRT-LLMCUDATritonCUTLASSCuBLASCuDNNNsight SystemsNsight Compute

Company Brief

Snowflake
Provides a cloud-native data platform for data warehousing, data engineering, data science, and analytics, enabling organizations to store, process, and share large-scale data across multiple cloud providers with elastic scalability and performance.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Bozeman, United States
Founded: 2012
WebsiteLinkedIn