AI Inference Core - Infrastructure SW Engineer

Cerebras
Sunnyvale
Workplace: HybridFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Independent problem solving","Systems thinking","Debugging","Problem-solving","Sound software engineering fundamentals"]

Design and build the core software that powers engineering infrastructure across machines and clusters. Work on Python frameworks, orchestration/control-plane systems, distributed execution, scheduling, test infrastructure, and developer tooling. Define APIs and module boundaries infrastructure teams depend on, and help handle concurrency, retries, idempotency, and partial failures. Collaborate with platform, CI, release, quality, ML systems, and product engineering teams to translate requirements into scalable designs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
1 month ago

AI Inference Core - Infrastructure SW Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Design and build the core software that powers engineering infrastructure across machines and clusters. Work on Python frameworks, orchestration/control-plane systems, distributed execution, scheduling, test infrastructure, and developer tooling. Define APIs and module boundaries infrastructure teams depend on, and help handle concurrency, retries, idempotency, and partial failures. Collaborate with platform, CI, release, quality, ML systems, and product engineering teams to translate requirements into scalable designs.
Location: Sunnyvale
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design, develop, test, and maintain Python frameworks and services used to orchestrate engineering workflows across machines and clusters.
  • •Build reusable abstractions for scheduling, distributed execution, resource management, test execution, workflow planning, and failure recovery.
  • •Define clear APIs, module boundaries, extension points, and data models to keep infrastructure maintainable as it evolves.
  • •Reason about concurrency and failure handling including asynchronous execution, multiprocessing, state management, retries, idempotency, cancellation, and partial failures.
  • •Debug complex issues across Python applications and distributed environments, and write high-quality automated tests and documentation.

Key Requirements

  • •3+ years of professional software-engineering experience.
  • •Strong proficiency in Python with solid understanding of runtime behavior.
  • •Experience designing maintainable software systems, libraries/frameworks, backend services, or developer-facing APIs.
  • •Understanding of concurrency (processes, threads, asynchronous execution, synchronization, shared state).
  • •Foundational understanding of distributed-systems concepts (retries, timeouts, idempotency, partial failure, coordination, eventual consistency).
Experience:3+ yearsAI infrastructureDistributed systemsDeveloper toolsCloud inference
Skills:Independent problem solvingSystems thinkingDebuggingProblem-solvingSound software engineering fundamentals
Tech Stack:PythonAsyncioMultiprocessingConcurrent futuresEvent-driven systemsPytestKubernetes

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn