Software Engineering Technical Leader | AI Cluster Orchestrator & Automation Engineer | 15+ years

Cisco
Bengaluru, Pune
Workplace: OnsiteFull timeFunction: QA, Test & Release EngineeringExperience: 12+ yearsEducation: bachelorsSkills: ["Ownership","Collaboration","Communication","Troubleshooting"]

Design and implement end-to-end, repeatable automation for AI cluster bring-up and lifecycle management across compute, network, and storage. Build idempotent orchestration workflows for GPU/service nodes, network fabrics, and storage, automating PXE, BIOS/firmware, OS, Kubernetes/operators, and Slurm integration. Coordinate multi-plane dependencies, add health checks and drift detection with validation/rollback, and produce runbooks and APIs to support operations at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
4 days ago

Software Engineering Technical Leader | AI Cluster Orchestrator & Automation Engineer | 15+ years

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Design and implement end-to-end, repeatable automation for AI cluster bring-up and lifecycle management across compute, network, and storage. Build idempotent orchestration workflows for GPU/service nodes, network fabrics, and storage, automating PXE, BIOS/firmware, OS, Kubernetes/operators, and Slurm integration. Coordinate multi-plane dependencies, add health checks and drift detection with validation/rollback, and produce runbooks and APIs to support operations at scale.
Location: Bengaluru, Pune
Workplace: Onsite
Employment Type: Full time
Job Function: QA, Test & Release Engineering

Key Responsibilities

  • •Design and implement repeatable, end-to-end automation for AI cluster bring-up, configuration, validation, lifecycle management, and teardown.
  • •Build idempotent orchestration workflows for GPU nodes, service nodes, network fabrics, and storage.
  • •Automate PXE, NVIDIA BCM, DHCP, Redfish, BIOS, firmware, OS, Kubernetes/operators, and Slurm integration.
  • •Coordinate dependencies across compute, Cisco networking, storage, GPU platforms, and service nodes.
  • •Implement health checks, configuration drift detection, validation gates, rollback, failure recovery, and operational observability; document runbooks, APIs, interfaces, and support handoffs.

Key Requirements

  • •Bachelors + 12 years of related experience, or Masters + 8 years of related or equivalent work experience.
  • •Experience with Linux systems and AI/GPU cluster architecture.
  • •Coding experience using Python and automation/API development.
  • •Prior experience with PXE, DHCP, Kubernetes, Slurm, BIOS/firmware, networking, and storage integration.
  • •Experience troubleshooting distributed provisioning failures and system dependencies.
Experience:12+ yearsAIGPU clustersDistributed systemsInfrastructure automation
Education:Bachelor's
Skills:OwnershipCollaborationCommunicationTroubleshooting
Tech Stack:PythonLinuxPXEDHCPNVIDIA BCMRedfishBIOSFirmwareKubernetesSlurmRESTInfrastructure-as-codeGPU nodesNetwork fabricsService nodesStorage automationCI/CDConfiguration managementLoggingTelemetry

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor