AI Systems, Model Optimization

Unconventional AI
Palo Alto
Workplace: RemoteFull timeFunction: Data Science & Machine LearningEducation: mastersSkills: []

Develop the end-to-end path from AI model architecture to physical silicon, creating training techniques, optimization strategies, and infrastructure so models run efficiently on Unconventional’s novel compute substrates. Build performance models and drive hardware-aware mapping, including quantization, sparsity/pruning, and distillation. Optimize and debug ML kernels using CUDA/Triton/CUTLASS and profile bottlenecks, collaborating closely with AI and hardware/infrastructure teams to enable tapeouts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Unconventional AI
Unconventional AI
1 month ago

AI Systems, Model Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Develop the end-to-end path from AI model architecture to physical silicon, creating training techniques, optimization strategies, and infrastructure so models run efficiently on Unconventional’s novel compute substrates. Build performance models and drive hardware-aware mapping, including quantization, sparsity/pruning, and distillation. Optimize and debug ML kernels using CUDA/Triton/CUTLASS and profile bottlenecks, collaborating closely with AI and hardware/infrastructure teams to enable tapeouts.
Location: Palo Alto
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Develop performance models to evaluate compute, memory, and energy trade-offs and track pareto-optimality across models and hardware configurations.
  • •Partition and map complex AI models down to hardware using optimization strategies such as custom quantization, sparsity/pruning, and distillation.
  • •Create and apply hardware-aware training techniques including Quantization-Aware Training (QAT), noise-aware training, and sparsification for analog compute constraints.
  • •Develop and optimize GPU kernels using low-level tools like CUDA, Triton, or CUTLASS; profile and debug ML code to resolve training and inference bottlenecks.
  • •Translate between AI model architects and hardware/infrastructure teams by converting model requirements into specifications and codifying learnings for tapeouts.

Pay and Benefits

Perks:Health Insurance401kPaid LeaveMeal Allowance

Key Requirements

  • •MS/PhD (or equivalent research/project experience) in a quantitative field such as AI/ML, Computer Science, Physics, Electrical Engineering, or Applied Math.
  • •Deep, practical understanding of the modern AI/ML stack and optimized compilation/execution on modern GPU systems, including profiling to resolve performance bottlenecks in complex ML codebases.
  • •Ability to map state-of-the-art model architectures (e.g., Transformers, Mixture of Experts, diffusion models) to system performance implications and apply efficiency techniques like sparsity, quantization, and distillation.
  • •Deep experience with PyTorch, including its internals, torch.compile, and distributed data parallel (DDP) / fully sharded data parallel (FSDP).
Education:Master's
Languages:English
Tech Stack:PyTorchTorch.compileDDPFSDPCUDATritonCUTLASSMegatron-LMDeepSpeedTransformersMixture of ExpertsDiffusion modelsQuantization-Aware Training (QAT)SparsityPruningDistillationSparsificationOptimizationProfilingTorch

Company Brief

Unconventional AI
Builds customizable AI copilots and fine-tuned large language model solutions to help teams automate workflows, surface actionable insights, and integrate generative AI into business processes across products and operations.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Website