Senior Software Engineer, Cosmos Infrastructure and End to End Performance

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 152,000 - 287,500 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: mastersSkills: ["Intellectual curiosity","Ability to dig in","Interpersonal skills"]

Build NVIDIA Cosmos’ foundational Physical AI platform by developing world foundation model capabilities and driving end-to-end performance analysis across data center infrastructure and edge deployments. Develop and automate infrastructure for data ingestion, curation, (pre/post) training, and export/quantization for deployment. Optimize robustness, fault tolerance, and inference performance while partnering with customers to ensure the models are easy to use and enable an ecosystem.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Software Engineer, Cosmos Infrastructure and End to End Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build NVIDIA Cosmos’ foundational Physical AI platform by developing world foundation model capabilities and driving end-to-end performance analysis across data center infrastructure and edge deployments. Develop and automate infrastructure for data ingestion, curation, (pre/post) training, and export/quantization for deployment. Optimize robustness, fault tolerance, and inference performance while partnering with customers to ensure the models are easy to use and enable an ecosystem.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build SoTA world foundation models such as Cosmos3.
  • •Drive end-to-end performance analysis and HW-SW codesign for both data center infrastructure and edge deployments.
  • •Develop infrastructure to automate data ingestion, curation, pre-training, post-training, and export/quantization and deployment on the edge.
  • •Design for robustness and fault tolerance in the platform.
  • •Engage with customers to ensure Cosmos models are easy to use and enable the ecosystem.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Equity and Bonus:Equity

Key Requirements

  • •5 years of relevant work experience.
  • •Masters in Computer Engineering, Computer Science, Electrical Engineering, or related STEM, or equivalent experience.
  • •Expertise in large-scale parallel and distributed accelerator-based systems and optimizing performance for AI workloads at scale, including performance modeling and benchmarking.
  • •Proficiency in Distributed PyTorch, Python, and C/C++, with strong foundations in computer architecture, networking, storage systems, and accelerators.
  • •Expertise in DNNs for emerging AI/ML applications and a deep understanding of World Foundation Models for Physical AI.
Experience:5+ yearsPhysical AIGenerative modelsDNNsMultimodal dataDistributed systems
Education:Master's
Skills:Intellectual curiosityAbility to dig inInterpersonal skills
Tech Stack:Distributed PyTorchPythonC/C++GCPAWSAzureOCITensorFlowJAXCosmosMegatron-LMTensorRT-LLMVLLMCUDAContainerizationExportQuantization

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor