Senior Machine Learning Systems Engineer

Atlassian
United States
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Skills: []

Design and optimize large-scale model serving systems end-to-end for Atlassian’s AI & ML Platform inference team. You’ll architect distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling), apply deep low-level GPU and inference optimizations (GPU kernels, quantization, speculative decoding), and ensure reliable, high-concurrency performance. You’ll also benchmark/accelerate inference engines and build CI/CD infrastructure for seamless deployment and updates.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Atlassian
Atlassian
1 day ago

Senior Machine Learning Systems Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design and optimize large-scale model serving systems end-to-end for Atlassian’s AI & ML Platform inference team. You’ll architect distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling), apply deep low-level GPU and inference optimizations (GPU kernels, quantization, speculative decoding), and ensure reliable, high-concurrency performance. You’ll also benchmark/accelerate inference engines and build CI/CD infrastructure for seamless deployment and updates.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
  • •Optimize latency and throughput of model inference for real production workloads.
  • •Build reliable, high-concurrency serving systems that can handle billions of requests.
  • •Benchmark, fine-tune, and accelerate inference engines.
  • •Create robust CI/CD infrastructure for model deployment and inference engine updates.

Pay and Benefits

Perks:Health Insurance

Key Requirements

  • •5+ years of software engineering experience and 2+ years of system performance optimization experience.
  • •Deep low-level systems programming in C/C++ or Rust.
  • •Experience with large-scale, high-concurrent production serving systems.
  • •Experience with GPU inference engines such as vLLM, SGLang, Triton, or TensorRT-LLM.
  • •Strong experience optimizing inference systems (batching, caching, load balancing, parallelism) and testing/benchmarking reliability.
Experience:Large-scale systemsGPU inference
Tech Stack:CC++RustVLLMSGLangTritonTensorRT-LLMGPU kernelsCI/CDQuantizationSpeculative decodingLoad balancingAuto-scalingKV cacheContinuous batching

Company Brief

Atlassian
Atlassian develops collaboration and workflow software—Jira, Confluence, Trello, Loom and more—used by teams for project management, software development, and IT service management across enterprises worldwide.
Industry: Developer Tools
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Sydney, Australia
Founded: 2002
Glassdoor
Glassdoor: 3.1
WebsiteLinkedInGlassdoor