Staff Machine Learning Engineer - ML Frameworks

Adobe Systems
San Jose, Seattle, San Francisco
Workplace: OnsiteFull timeUSD 172,500 - 306,625 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsEducation: phdSkills: ["Critical thinking","Analytical problem-solving","Quantitative problem-solving","Communication","Teamwork"]

Design, develop, and maintain AI/ML infrastructure for Firefly model training and deployment at large scale. Build distributed training frameworks that use GPUs to improve performance, scalability, resiliency, and elasticity, including support for FSDP and model parallelism. Partner with researchers and data scientists to streamline the model training pipeline, improve orchestration and scheduling, and accelerate experimentation through AutoML and related tooling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Adobe Systems
Adobe Systems
1 week ago

Staff Machine Learning Engineer - ML Frameworks

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design, develop, and maintain AI/ML infrastructure for Firefly model training and deployment at large scale. Build distributed training frameworks that use GPUs to improve performance, scalability, resiliency, and elasticity, including support for FSDP and model parallelism. Partner with researchers and data scientists to streamline the model training pipeline, improve orchestration and scheduling, and accelerate experimentation through AutoML and related tooling.
Location: San Jose, Seattle, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, develop, and maintain AI/ML infrastructure solutions for training and deployment of large-scale AI models.
  • •Implement and improve distributed training frameworks using GPUs to enhance performance, scalability, resiliency, and elasticity.
  • •Improve orchestration and scheduling to scale training jobs and accelerate experimentation with AutoML and similar approaches.
  • •Collaborate with data scientists and ML researchers to streamline the model training pipeline and ensure efficient resource utilization.
  • •Drive innovation in infrastructure practices to support machine learning research and development.

Pay and Benefits

Salary: USD 172,500 - 306,625 annually
Equity and Bonus:Equity

Key Requirements

  • •PhD or Master’s in computer science or a related field, plus 5+ years of hands-on industry experience.
  • •Proficiency in Python and experience developing systems, frameworks, and SDKs.
  • •Experience with infrastructure and understanding of model serving, training, orchestration, and GPU resource management.
  • •Experience with machine learning and distributed PyTorch.
  • •Strong critical thinking and analytical, quantitative problem-solving skills.
Experience:5+ yearsMachine learning
Education:PhD / Doctorate in Computer science or related field
Skills:Critical thinkingAnalytical problem-solvingQuantitative problem-solvingCommunicationTeamwork
Tech Stack:FireflyKubernetesPythonAWSGPUsPyTorchFSDPModel parallelismAutoMLKubeFlowMLflowRaySageMakerMPIMegatronHorovod

Company Brief

Adobe Systems
Provides creative, marketing, and document management software and cloud services, including Photoshop, Illustrator, Acrobat, and the Adobe Experience Cloud, serving creative professionals, enterprises, and governments worldwide.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1982
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor