Member of Technical Staff - AI Training Platform

Unconventional AI
Mountain View, Seattle
Workplace: OnsiteFull timeFunction: Education & TrainingExperience: 5+ yearsEducation: bachelorsSkills: ["Cross-functional collaboration","Communication","Problem-solving","Technical ownership","Translating requirements"]

Build and scale the end-to-end tooling that powers AI model training, evaluation, and benchmarking. Own core training infrastructure including multi-node distributed systems, elastic sharding, data streaming, and checkpointing/recovery. Develop and optimize performance-critical kernels with CUDA and Triton, create benchmarking suites, and build hardware evaluation tooling. Partner with researchers and infrastructure/hardware teams to translate algorithmic trade-offs into robust system specifications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Unconventional AI
Unconventional AI
3 days ago

Member of Technical Staff - AI Training Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build and scale the end-to-end tooling that powers AI model training, evaluation, and benchmarking. Own core training infrastructure including multi-node distributed systems, elastic sharding, data streaming, and checkpointing/recovery. Develop and optimize performance-critical kernels with CUDA and Triton, create benchmarking suites, and build hardware evaluation tooling. Partner with researchers and infrastructure/hardware teams to translate algorithmic trade-offs into robust system specifications.
Location: Mountain View, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Architect, scale, and maintain the AI/ML/RL training platform infrastructure, including multi-node distributed training, elastic sharding, and robust data streaming pipelines.
  • •Implement robust model checkpointing and recovery mechanisms for fast, large-scale iteration.
  • •Maintain and expand the proprietary training framework used for internal model training.
  • •Develop and optimize GPU kernels and low-level performance work using CUDA and Triton; build benchmarking suites to track MFU, memory bandwidth, and convergence stability.
  • •Collaborate with researchers and engineering teams to translate algorithmic trade-offs into infrastructure and hardware specifications, and build tooling to evaluate novel hardware against benchmarks.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquity401kPaid LeaveMeal Allowance

Key Requirements

  • •5+ years of AI/ML engineering experience with mastery of modern ML software stack and distributed model training (e.g., transformers, Mixture of Experts, diffusion models).
  • •Strong ability to map state-of-the-art model architectures to system performance implications, including cluster partitioning and communication primitives.
  • •Proven production experience implementing, debugging, and maintaining training frameworks such as Megatron-LM, DeepSpeed, Ray, and PyTorch Lightning.
  • •Strong programming skills in Python or C++ with demonstrated ability to build reliable training infrastructure.
  • •BS in Computer Science, Physics, Electrical Engineering, or Applied Math.
Experience:5+ yearsAI/MLDistributed trainingLarge language modelsGenerative modelsHigh-performance computing
Education:Bachelor's in Computer Science, Physics, Electrical Engineering, or Applied Math
Skills:Cross-functional collaborationCommunicationProblem-solvingTechnical ownershipTranslating requirements
Languages:English
Tech Stack:PythonC++CUDATritonCUDA kernelsGPU kernelsKubernetesNCCLPyTorch LightningMegatron-LMDeepSpeedRayVLLMSGLangVeRLTRLElastic shardingDistributed trainingModel checkpointingData streaming pipelines

Company Brief

Unconventional AI
Builds customizable AI copilots and fine-tuned large language model solutions to help teams automate workflows, surface actionable insights, and integrate generative AI into business processes across products and operations.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Website