Research Engineer, Infrastructure

Cognition
San Francisco
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Skills: ["Problem-solving","Communication","Collaboration","Autonomy"]

Join a small, highly skilled team building and operating the core distributed training infrastructure that powers large AI models. Own end-to-end systems from training job orchestration and data pipelines to fault-tolerant GPU clusters, enabling researchers to run thousands of GPUs efficiently and with minimal downtime.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cognition
Cognition
4 months ago

Research Engineer, Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Join a small, highly skilled team building and operating the core distributed training infrastructure that powers large AI models. Own end-to-end systems from training job orchestration and data pipelines to fault-tolerant GPU clusters, enabling researchers to run thousands of GPUs efficiently and with minimal downtime.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Own and operate distributed training infrastructure, including job launchers, checkpointing, recovery, fault tolerance, and monitoring for researchers across thousands of GPUs
  • •Design and maintain experiment orchestration and tooling to launch, track, and analyze research experiments, reducing friction in the research loop
  • •Develop high-throughput data pipelines for training and evaluation, ensuring data quality and reproducibility at scale
  • •Profile and optimize training throughput and compute efficiency, addressing bottlenecks in data loading, communication, and memory usage
  • •Anticipate researchers' needs and build scalable infrastructure ahead of time to avoid constraints

Key Requirements

  • •Proficiency in Python and C++; experience with PyTorch or equivalent deep learning frameworks at a systems level, not just API usage
  • •Strong experience building and operating distributed training systems for large models; end-to-end ownership from cluster level down to the training loop
  • •Hands-on experience with GPU performance profiling, memory optimization, and compute efficiency; able to diagnose why a training run is underperforming and fix it
  • •Experience implementing or optimizing parallelism strategies (data, tensor, pipeline, sequence) for large model training
  • •Ability to design and build tooling to accelerate research workflows and reduce friction in the research loop
Experience:AIDistributed systemsGPU
Skills:Problem-solvingCommunicationCollaborationAutonomy
Languages:English
Tech Stack:PythonC++PyTorchGPUDistributed systemsMemory optimizationProfiling

Company Brief

Cognition
Builds Devin, an autonomous AI software engineer and AI-native developer tools (Windsurf, DeepWiki) to automate software engineering workflows for enterprise customers.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2023
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor