ML Engineer, Inference & Optimization

Pika
Palo Alto
Workplace: OnsiteFull timeUSD 185,000 - 250,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Collaboration","Communication","Problem-solving","Mentoring"]

Senior Inference Engineer at an AI startup accelerating inference for video and language models. You will design high-performance inference pipelines, optimize GPU parallelism, implement acceleration techniques (quantization, attention), and deploy state-of-the-art videogen/LLM models at scale, collaborating across research and engineering to push real-time AI capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pika
Pika
2 months ago

ML Engineer, Inference & Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Senior Inference Engineer at an AI startup accelerating inference for video and language models. You will design high-performance inference pipelines, optimize GPU parallelism, implement acceleration techniques (quantization, attention), and deploy state-of-the-art videogen/LLM models at scale, collaborating across research and engineering to push real-time AI capabilities.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Manager level

Key Responsibilities

  • •Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
  • •Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
  • •Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
  • •Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
  • •Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
  • •Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.

Pay and Benefits

Salary: USD 185,000 - 250,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceMonthly StipendsCompany Retreats

Key Requirements

  • •5+ years engineering experience with inference acceleration and model deployment at scale.
  • •Proven expertise in inference optimization (quantization, attention acceleration, deep learning compiler stacks).
  • •Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP for distributed inference.
  • •Familiarity with video generation (videogen) models and large language models (LLMs).
  • •Strong cross-discipline collaboration skills; ownership mindset and ability to work in a fast-paced startup environment.
Experience:5+ yearsAIVideo generationLLMsDistributed computing
Skills:CollaborationCommunicationProblem-solvingMentoring
Languages:English
Tech Stack:CUDANCCLGPUQuantizationAttentionTPSPPPDistributed

Company Brief

Pika
Pika (Pika Labs) builds AI-driven tools to generate and edit videos from text prompts and images, enabling creators to produce cinematic, 3D, anime and stylized videos with a web-based platform and APIs.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Valuation: USD 250M to 500M
Funding: Series B
Headquarters: Palo Alto, United States
Founded: 2023
WebsiteLinkedIn