ML Research Engineer (Inference)
Cerebras
Bengaluru
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Experience: 1-3 yearsEducation: bachelorsSkills: ["Problem-solving","Debugging","Collaboration","Analysis"]Adapt state-of-the-art language and vision models for efficient inference on Cerebras’ flagship hardware. You’ll work with the Inference ML team to prototype, validate, and optimize models, focusing on speculative decoding, pruning/compression, sparse attention, and sparsity-driven techniques to achieve low latency and high throughput. Responsibilities include running experiments, debugging issues, profiling performance with internal tools, and collaborating across ML, software, and hardware teams to bring up and validate models at scale.

