Cloud & Customer Solutions Engineer - DC GPU

AMD
Bellevue
Workplace: HybridFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Debugging","Root-cause analysis","Customer communication","Performance tuning","Troubleshooting"]

Embed with AMD’s strategic AI customers to take AMD Instinct GPU clusters from deployment through sustained production excellence. Own end-to-end customer outcomes including cluster bring-up and certification, workload onboarding and performance validation, and Sev-1 incident response. Write production code, operate live GPU clusters, build observability/benchmarking tooling, and transfer operational capability so customers progress toward independent operation. Contribute field learnings upstream to ROCm and the serving ecosystem.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

Cloud & Customer Solutions Engineer - DC GPU

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Embed with AMD’s strategic AI customers to take AMD Instinct GPU clusters from deployment through sustained production excellence. Own end-to-end customer outcomes including cluster bring-up and certification, workload onboarding and performance validation, and Sev-1 incident response. Write production code, operate live GPU clusters, build observability/benchmarking tooling, and transfer operational capability so customers progress toward independent operation. Contribute field learnings upstream to ROCm and the serving ecosystem.
Location: Bellevue
Workplace: Hybrid
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Own customer deployments end-to-end, including cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets.
  • •Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) for customer-specific workloads and SLOs across cloud, NeoCloud, and bare metal.
  • •Lead root-cause analysis and resolution of production incidents on customer clusters, including Sev-1 response, and drive fixes to permanent closure.
  • •Deploy agentic AI solutions in customer environments in partnership with Agentic Data Engineers and own their production behavior within the engagement.
  • •Transfer operational capability to customer teams through documentation, runbooks, and hands-on enablement, progressing customers toward independent production ownership.

Key Requirements

  • •5+ years of production software or infrastructure engineering with experience operating or deploying systems in environments you did not build.
  • •Hands-on GPU compute experience at scale, including cluster deployment, distributed training or high-throughput inference, and performance debugging.
  • •Strong AI infrastructure knowledge: Kubernetes and/or Slurm, containerized GPU workloads, RCCL/NCCL, and high-performance networking (RoCE/InfiniBand) plus observability tooling (Prometheus, Grafana).
  • •Cloud platform depth across AWS, Azure, GCP, or NeoCloud, including hybrid and bare-metal deployment patterns.
  • •Proficiency in Python and at least one systems language, plus familiarity with LLM serving patterns (inference serving, RAG, agentic workflows).
Experience:5+ yearsAIGPU computingInfrastructure engineeringCloudLLM servingOpen-source
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience
Skills:DebuggingRoot-cause analysisCustomer communicationPerformance tuningTroubleshooting
Languages:English
Tech Stack:PythonROCmVLLMSGLangRCCLKubernetesSlurmNCCLAWSAzureGCPNeoCloudRoCEInfiniBandPrometheusGrafanaLLMRAGAgentic workflowsContainerized GPU workloads

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn