Member of Technical Staff (MTS), Machine Learning, SMAI

Micron Technology
Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Analytical thinking","Interpersonal communication","Oral communication","Written communication","Prioritization"]

Build and deploy scalable AI/ML and agentic GenAI solutions for Micron’s Smart Manufacturing and AI mission. Architect and execute large-scale custom model training and fine-tuning on multi-node, multi-GPU clusters, optimize distributed training performance, and develop autonomous agents to automate manufacturing workflows. Create data/solution pipelines, design data models in Snowflake and Google Cloud, and maintain ML/AI CI/CD pipelines while collaborating with data scientists and engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Micron Technology
Micron Technology
2 weeks ago

Member of Technical Staff (MTS), Machine Learning, SMAI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build and deploy scalable AI/ML and agentic GenAI solutions for Micron’s Smart Manufacturing and AI mission. Architect and execute large-scale custom model training and fine-tuning on multi-node, multi-GPU clusters, optimize distributed training performance, and develop autonomous agents to automate manufacturing workflows. Create data/solution pipelines, design data models in Snowflake and Google Cloud, and maintain ML/AI CI/CD pipelines while collaborating with data scientists and engineers.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Architect and execute large-scale custom model training and fine-tuning jobs (SFT, RLHF) on multi-node, multi-GPU clusters.
  • •Optimize training throughput and memory efficiency using distributed training strategies (FSDP, DeepSpeed, Megatron-LM) and mixed-precision (FP16/BF16).
  • •Design and develop autonomous AI agents that perform multi-step reasoning, planning, and tool execution to automate manufacturing workflows.
  • •Implement agentic frameworks (e.g., LangChain, LangGraph, CrewAI) to orchestrate LLM interactions with internal APIs, databases, and software tools.
  • •Profile and debug GPU performance bottlenecks and build/maintain data and CI/CD pipelines for ML and GenAI applications.

Key Requirements

  • •10+ years building scalable ETL pipelines and 10+ years in big data processing and/or developing applications and data sources.
  • •Deep understanding of GPU architecture and GPU resource management in cloud and on-prem environments.
  • •Hands-on distributed training/model parallelism experience (DDP, FSDP, and model parallelism).
  • •Proficiency fine-tuning LLMs using PEFT (LoRA, QLoRA) and optimizing inference engines (vLLM, TensorRT-LLM).
  • •Strong programming/scripting skills in Python (or Java) and CI/CD tooling (Jenkins, Git, Docker, Kubernetes).
Experience:Big dataGenAICloudDistributed systemsGPU computing
Education:
Skills:Analytical thinkingInterpersonal communicationOral communicationWritten communicationPrioritization
Tech Stack:PyTorchTensorFlowScikit-learnPythonJavaCUDATritonC++LoRAQLoRAPEFTVLLMTensorRT-LLMLangChainLangGraphCrewAILlamaIndexAutoGenDDPFSDP

Company Brief

Micron Technology
Designs and manufactures semiconductor memory and storage solutions, including DRAM, NAND, and NOR flash, for computing, networking, mobile, and automotive markets worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Boise, United States
Founded: 1978
Glassdoor
Glassdoor: 3.7
WebsiteLinkedIn