Infrastructure Production Engineer

Vultr
United States
Workplace: RemoteFull timeUSD 60,000 - 80,000 annuallyFunction: Manufacturing & Production OperationsSkills: ["Analytical skills","Attention to detail","Documentation","Communication","Collaboration"]

Build and maintain automated diagnostic and validation frameworks to ensure production readiness of NVIDIA and AMD GPU hardware across Vultr’s global environment. Engineer Python-based agents and services using Ansible automation to orchestrate provisioning, telemetry, health monitoring, and onboarding workflows. Analyze performance and diagnostic outputs, evolve testing and verification (including RMA processes), and create tooling that gates deployments and reduces infrastructure risk.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Vultr
Vultr
1 day ago

Infrastructure Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build and maintain automated diagnostic and validation frameworks to ensure production readiness of NVIDIA and AMD GPU hardware across Vultr’s global environment. Engineer Python-based agents and services using Ansible automation to orchestrate provisioning, telemetry, health monitoring, and onboarding workflows. Analyze performance and diagnostic outputs, evolve testing and verification (including RMA processes), and create tooling that gates deployments and reduces infrastructure risk.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Entry level

Key Responsibilities

  • •Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware using vendor tooling and infrastructure automation technologies.
  • •Engineer and support Python-based agents/services/APIs and Ansible automation to orchestrate hardware provisioning, telemetry, health monitoring, and production onboarding.
  • •Analyze workload performance, utilization, thermals, and diagnostic output to identify issues and improve validation and infrastructure readiness standards.
  • •Execute and evolve testing and verification processes for production onboarding and Return Material Authorization (RMA), improving reliability, scalability, and automation coverage.
  • •Create deployment-gating tooling, enforce quality standards, and document system designs and operational guidance while maintaining accurate Jira records.

Pay and Benefits

Salary: USD 60,000 - 80,000 annually
Perks:Health InsuranceDentalVision401kPaid LeaveLearning BudgetOffice StipendGym MembershipSabbatical

Key Requirements

  • •Strong analytical skills and attention to detail for evaluating and optimizing validation and testing methodologies.
  • •Strong Linux proficiency in command-line production environments.
  • •Design, write, and modify Python and shell scripts for infrastructure diagnostics, automation, and validation workflows.
  • •Familiarity with infrastructure automation or configuration management frameworks such as Ansible.
  • •Understand data center networking concepts and protocols such as DHCP, IPv6, and ICMP, and communicate/document work clearly.
Experience:Cloud infrastructureData centerGPUInfrastructure automation
Skills:Analytical skillsAttention to detailDocumentationCommunicationCollaboration
Tech Stack:LinuxPythonShell scriptsAnsibleAnsible playbooksAnsible rolesAnsible pipelinesNVIDIAAMDTelemetryJiraDHCPIPv6ICMP

Company Brief

Vultr
Provides cloud infrastructure services including VPS, dedicated instances, block storage, and bare metal across global data centers. Targets developers and businesses with simple, high-performance, and cost-effective cloud compute and networking solutions.
Industry: Cloud Computing
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: West Palm Beach, United States
Founded: 2014
WebsiteLinkedIn