Senior Systems Engineer, Workers AI

Cloudflare
Indiana, Austin, London
Workplace: HybridFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Mentorship","Technical leadership","Collaboration","Problem-solving","High-agency"]

Design and build core infrastructure for AI inference across Cloudflare’s global network, powering real-time voice, frontier open LLMs, and customer-deployed models on heterogeneous GPU/accelerator fleets. Drive performance and optimization across scheduling and request routing, improve reliability and resilience, expand observability (metrics, logging, tracing), and lead cross-functional technical projects from concept through deployment and operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cloudflare
Cloudflare
5 months ago

Senior Systems Engineer, Workers AI

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Design and build core infrastructure for AI inference across Cloudflare’s global network, powering real-time voice, frontier open LLMs, and customer-deployed models on heterogeneous GPU/accelerator fleets. Drive performance and optimization across scheduling and request routing, improve reliability and resilience, expand observability (metrics, logging, tracing), and lead cross-functional technical projects from concept through deployment and operations.
Location: Indiana, Austin, London
Workplace: Hybrid
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Develop and maintain core components of a serverless inference platform to ensure high availability and scalability.
  • •Optimize model scheduling to improve efficiency and utilization, and improve request routing to reduce end-user latency.
  • •Identify and mitigate systemic risks to drive measurable improvements in reliability and resilience.
  • •Expand and refine observability (metrics, logging, tracing) and fine-tune alerts to proactively resolve production issues.
  • •Lead complex, cross-functional technical projects from initial concept and design through deployment and operationalization.

Key Requirements

  • •Proven experience in systems engineering with a primary focus on distributed, high-performance systems.
  • •Expert proficiency in Rust, particularly in an asynchronous environment.
  • •Deep understanding and hands-on experience with networking and application protocols including TCP, HTTP, and WebSocket.
  • •Solid experience scaling and optimizing performance techniques such as load balancing and caching in distributed environments.
Experience:Distributed systemsHigh-performance computingAI inferenceLLMs
Skills:MentorshipTechnical leadershipCollaborationProblem-solvingHigh-agency
Languages:English
Tech Stack:RustTCPHTTPWebSocketKubernetesNomadMetricsLoggingTracingLoad balancingCachingKV cacheGPUsAccelerators

Company Brief

Cloudflare
Provides a global network and cloud platform that delivers security, performance, and reliability services for web applications, APIs, and Internet properties, including CDN, DDoS protection, DNS, and zero-trust security solutions.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn