Senior Software Engineer

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 10+ yearsSkills: ["Technical leadership","Mentorship","Influence without authority","Problem-solving","Establishing engineering standards"]

Design, build, and operate enterprise observability and AI-driven reliability platform capabilities at large scale. You’ll develop distributed telemetry and event-processing systems, reusable platform services/APIs/automation frameworks, and intelligent reliability features such as agentic workflows for anomaly detection, forecasting, root-cause analysis, and closed-loop remediation. Lead architecture across Storage, Compute, Network, and platform domains while improving scalability, security, performance, MTTR, operational toil, and engineering productivity.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Senior Software Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Design, build, and operate enterprise observability and AI-driven reliability platform capabilities at large scale. You’ll develop distributed telemetry and event-processing systems, reusable platform services/APIs/automation frameworks, and intelligent reliability features such as agentic workflows for anomaly detection, forecasting, root-cause analysis, and closed-loop remediation. Lead architecture across Storage, Compute, Network, and platform domains while improving scalability, security, performance, MTTR, operational toil, and engineering productivity.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate distributed software platforms for enterprise observability, telemetry, automation, and infrastructure reliability at scale.
  • •Develop reusable platform services, APIs, automation frameworks, and control planes to enable self-service and reduce operational toil.
  • •Build scalable telemetry and event-processing systems for metrics, logs, traces, events, topology, and alerts with high performance and efficiency.
  • •Implement AI-native reliability capabilities such as agentic workflows for anomaly detection, forecasting, root-cause analysis, automated debugging, and closed-loop remediation.
  • •Drive technical architecture and engineering direction across Storage, Compute, Network, and platform domains; improve reliability, MTTR, operational toil, productivity, and infrastructure efficiency.

Key Requirements

  • •10+ years of software engineering/SRE/infrastructure/distributed-systems experience with demonstrated technical leadership.
  • •Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.
  • •Strong Go or Python expertise building production-grade distributed systems, platform services, APIs, and automation.
  • •Experience owning complex platform initiatives across multiple teams and infrastructure domains, from architecture through measurable impact.
  • •Deep knowledge of distributed systems/event-driven architectures and high-throughput telemetry, including experience with Kafka/NATS/gRPC (or equivalents).
Experience:10+ yearsSREDistributed systemsObservabilityInfrastructureOn-premises
Education:
Skills:Technical leadershipMentorshipInfluence without authorityProblem-solvingEstablishing engineering standards
Tech Stack:GoPythonKafkaNATSGRPCOpenTelemetryPrometheusVictoriaMetricsVectorLokiGrafanaClickHouseKubernetesOpenShiftVMwareBare-metalTerraformAnsibleFastAPIMicroservices

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor