Principal Machine Learning Systems Engineer (GenAI Products & Knowledge Innovations)

Atlassian
Washington DC, Mountain View, San Francisco
Workplace: RemoteFull timeFunction: IT Operations (Systems/Network Admin)Experience: 6+ yearsSkills: ["Collaboration","Problem-solving","Reliability focus","Rapid prototyping"]

Design and build scalable ML systems for training, fine-tuning, and serving large language models, embeddings, and enterprise knowledge-powered retrieval and RAG pipelines. Partner with applied scientists to evolve proof-of-concept prototypes into production-ready services, optimizing latency, throughput, and resource efficiency. Collaborate with ML engineers, backend developers, and product teams to ship end-to-end GenAI innovations, while advancing best practices in deployment, monitoring, and evaluation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Atlassian
Atlassian
1 day ago

Principal Machine Learning Systems Engineer (GenAI Products & Knowledge Innovations)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Design and build scalable ML systems for training, fine-tuning, and serving large language models, embeddings, and enterprise knowledge-powered retrieval and RAG pipelines. Partner with applied scientists to evolve proof-of-concept prototypes into production-ready services, optimizing latency, throughput, and resource efficiency. Collaborate with ML engineers, backend developers, and product teams to ship end-to-end GenAI innovations, while advancing best practices in deployment, monitoring, and evaluation.
Location: Washington DC, Mountain View, San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Architect and implement scalable systems for training, fine-tuning, and serving large language models and embeddings.
  • •Build retrieval, hybrid search, and RAG pipelines integrated with knowledge-grounded data.
  • •Develop tools and infrastructure to support rapid experimentation, evaluation, and deployment of prototypes.
  • •Partner with applied scientists to deliver robust, production-ready pipelines and evolve POCs into reliable services.
  • •Optimize latency, throughput, and resource efficiency for GenAI workloads and contribute to deployment, monitoring, and evaluation best practices.

Key Requirements

  • •6+ years in ML systems engineering, backend engineering, or infrastructure roles.
  • •Experience building and scaling ML-powered services in production.
  • •Experience with large-scale model training, inference pipelines, or search/retrieval systems.
  • •Proficiency in backend systems and ML frameworks including Python, PyTorch, TensorFlow, and Hugging Face.
  • •Strong coding skills and the ability to optimize systems for performance and reliability.
Experience:6+ years
Education:
Skills:CollaborationProblem-solvingReliability focusRapid prototyping
Tech Stack:PythonPyTorchTensorFlowHugging FaceWeaviatePineconeFAISSLangChainLlamaIndexAWSGCPAzureKubernetesDocker

Company Brief

Atlassian
Atlassian develops collaboration and workflow software—Jira, Confluence, Trello, Loom and more—used by teams for project management, software development, and IT service management across enterprises worldwide.
Industry: Developer Tools
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Sydney, Australia
Founded: 2002
Glassdoor
Glassdoor: 3.1
WebsiteLinkedInGlassdoor