Staff Engineer, Compute

DataDog
Dublin, Madrid, Paris
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical leadership","Influencing architecture","Systems mindset","Problem-solving","Cross-team collaboration"]

Lead technical direction for capacity management and workload placement of a Kubernetes platform scaling across 100,000+ VMs on AWS, Google Cloud, and Azure. Design and build systems for scheduling and regional deployment that balance reliability, performance, and capacity constraints. Partner with infrastructure teams to evolve multi-region, multi-cloud orchestration and write production Go software to improve platform automation, efficiency, and forecasting using operational capacity signals.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
DataDog
DataDog
1 day ago

Staff Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead technical direction for capacity management and workload placement of a Kubernetes platform scaling across 100,000+ VMs on AWS, Google Cloud, and Azure. Design and build systems for scheduling and regional deployment that balance reliability, performance, and capacity constraints. Partner with infrastructure teams to evolve multi-region, multi-cloud orchestration and write production Go software to improve platform automation, efficiency, and forecasting using operational capacity signals.
Location: Dublin, Madrid, Paris
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead technical direction for capacity management and workload placement for the Kubernetes platform across multiple cloud providers.
  • •Design and build systems that optimize workload scheduling and deployment across regions while balancing capacity constraints, reliability, and performance.
  • •Partner with infrastructure teams to evolve multi-region, multi-cloud capacity orchestration as the platform scales.
  • •Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency.
  • •Use capacity signals and operational data to influence infrastructure decisions, forecast growth, and improve workload placement strategies.

Pay and Benefits

Equity and Bonus:Equity
Perks:RsusEquity

Key Requirements

  • •Significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems.
  • •Strong software engineering skills, ideally with Go or a comparable systems programming language.
  • •Hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and distributed cloud infrastructure.
  • •A systems mindset with enjoyment for complex infrastructure, scheduling, or capacity management problem-solving.
  • •Comfort working with operational data, capacity forecasting, or analytical approaches that inform engineering decisions.
Experience:KubernetesCloud infrastructureDistributed systems
Skills:Technical leadershipInfluencing architectureSystems mindsetProblem-solvingCross-team collaboration
Languages:English
Tech Stack:KubernetesGoAWSGoogle CloudAzure

Company Brief

DataDog
Provides a cloud-native monitoring and observability platform that unifies metrics, traces, logs, and security signals to help engineering, operations, and security teams monitor and troubleshoot modern applications and infrastructure.
Industry: Developer Tools
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 2010
Glassdoor
Glassdoor: 4.1
WebsiteLinkedIn