Software Engineer, GPU Inference

Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 8+ yearsEducation: phdSkills: ["Technical guidance","Cross-functional collaboration","Problem-solving","Communication","Presentation"]

Design and implement APIs and machine-learning features that enable state-of-the-art generative AI models to run efficiently on Cerebras custom hardware. Own scalable, high-throughput, low-latency inference and serving backends for multimodal (image, audio, video) inputs, improving latency, throughput, memory usage, and compute efficiency. Lead technical guidance for engineers, drive observability and performance optimization, and build robust automated test suites.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
9 months ago

Software Engineer, GPU Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design and implement APIs and machine-learning features that enable state-of-the-art generative AI models to run efficiently on Cerebras custom hardware. Own scalable, high-throughput, low-latency inference and serving backends for multimodal (image, audio, video) inputs, improving latency, throughput, memory usage, and compute efficiency. Lead technical guidance for engineers, drive observability and performance optimization, and build robust automated test suites.
Location: United States, Canada
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Drive and provide technical guidance to a team of software engineers on complex machine learning integration projects.
  • •Design and implement ML features to improve generative AI model performance at inference time (e.g., structured outputs, biased sampling).
  • •Design and implement high-throughput, low-latency multimodal inference models for image, audio, and video inputs and outputs.
  • •Maintain and scale a serving backend for many concurrent requests per minute, including building detailed observability across the stack.
  • •Analyze and optimize latency, throughput, memory usage, and compute efficiency; build automated test suites to ensure software quality and reliability.

Key Requirements

  • •Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Mathematics, or a related field.
  • •8+ years of experience in large-scale software engineering with a focus on deep learning or related domains.
  • •Proficiency in Python for building and maintaining scalable systems.
  • •Advanced proficiency in C++ with emphasis on multi-threaded programming, performance optimization, and system-level development.
  • •Experience building and scaling large-scale inference systems for LLMs or multimodal models, including familiarity with LLM serving frameworks such as vLLM, SGLang, and TensorRT-LLM.
Experience:8+ yearsDeep learningLLMsMultimodalGenerative AI
Education:PhD / Doctorate in Computer Science, Computer Engineering, Mathematics, or a related field
Skills:Technical guidanceCross-functional collaborationProblem-solvingCommunicationPresentation
Tech Stack:PythonC++VLLMSGLangTensorRT-LLMPyTorchLLMsMultimodal inference

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn