Staff Inference ML Runtime Engineer
Cerebras
United States, Canada
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 8+ yearsEducation: phdSkills: ["Technical guidance","Cross-functional collaboration","Problem-solving","Communication","Presentation"]Design and implement APIs and machine-learning features that enable state-of-the-art generative AI models to run efficiently on Cerebras custom hardware. Own scalable, high-throughput, low-latency inference and serving backends for multimodal (image, audio, video) inputs, improving latency, throughput, memory usage, and compute efficiency. Lead technical guidance for engineers, drive observability and performance optimization, and build robust automated test suites.

