MTS, AI Engineering, SMAI

Micron Technology
Taiwan
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Analytical thinking","Communication","Prioritization","Mentoring","Interpersonal skills"]

Build and deploy scalable AI/ML solutions for Micron’s Smart Manufacturing and AI team, focusing on GPU performance for custom model training, fine-tuning (SFT/RLHF), and inference. Architect and optimize distributed training on multi-node, multi-GPU clusters, develop high-performance GPU kernels (CUDA/HIP/PTX/SASS), and create performance regression testing. Collaborate with hardware and ML teams to characterize workloads and improve throughput, memory efficiency, and latency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Micron Technology
Micron Technology
3 months ago

MTS, AI Engineering, SMAI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and deploy scalable AI/ML solutions for Micron’s Smart Manufacturing and AI team, focusing on GPU performance for custom model training, fine-tuning (SFT/RLHF), and inference. Architect and optimize distributed training on multi-node, multi-GPU clusters, develop high-performance GPU kernels (CUDA/HIP/PTX/SASS), and create performance regression testing. Collaborate with hardware and ML teams to characterize workloads and improve throughput, memory efficiency, and latency.
Location: Taiwan
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Architect and execute large-scale custom model training and fine-tuning jobs (SFT, RLHF) on multi-node, multi-GPU clusters.
  • •Optimize training throughput and memory efficiency using distributed training strategies and mixed-precision techniques.
  • •Design and develop autonomous AI Agents for multi-step reasoning, planning, and tool execution to automate manufacturing workflows.
  • •Analyze and profile workloads (e.g., LLM training, rendering pipelines) to identify compute, memory bandwidth, and latency bottlenecks.
  • •Write and optimize high-performance GPU kernels using CUDA/HIP/custom assembly and build performance regression testing suites to detect degradations.

Key Requirements

  • •Minimum 5+ years of experience in performance optimization, parallel computing, or low-level systems programming.
  • •Strong understanding of GPU architecture (memory hierarchy, tensor cores, NVLink/interconnects) and managing GPU resources across cloud and on-prem.
  • •Hands-on experience with distributed training/model parallelism (DDP, FSDP, and other model parallelism techniques).
  • •Proficiency fine-tuning LLMs using PEFT (LoRA, QLoRA) and optimizing inference engines (vLLM, TensorRT-LLM).
  • •Proficiency in programming with Python or Java, plus software development skills for CI/CD and cloud/container environments (Jenkins, Git, Docker, Kubernetes).
Experience:5+ years
Skills:Analytical thinkingCommunicationPrioritizationMentoringInterpersonal skills
Tech Stack:GPU architectureMulti-node clustersMulti-GPUFSDPDeepSpeedMegatron-LMMixed precisionFP16BF16CUDAHIPPTXSASSC++DDPLoRAQLoRAPEFTVLLMTensorRT-LLM

Company Brief

Micron Technology
Designs and manufactures semiconductor memory and storage solutions, including DRAM, NAND, and NOR flash, for computing, networking, mobile, and automotive markets worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Boise, United States
Founded: 1978
Glassdoor
Glassdoor: 3.7
WebsiteLinkedIn