Sr. AI Infrastructure Engineer, LLM/AI Platforms (Remote)

Crowdstrke
United States
Workplace: RemoteFull timeUSD 140,000 - 215,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsEducation: bachelorsSkills: ["Technical leadership","Mentorship","Engineering best practices","Stakeholder communication","Results delivery"]

Design, build, and deploy AI infrastructure that powers large-scale training, fine-tuning, and inference for LLM-based security products. Provision GPU clusters, optimize model-serving frameworks, and lead model lifecycle management (versioning, checkpointing, reproducibility). Architect data platforms for LLMs, RAG, and agentic systems, define robust evaluation and MLOps/DataOps practices, and mentor engineers to deliver production-ready, performance-focused systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crowdstrke
Crowdstrke
1 month ago

Sr. AI Infrastructure Engineer, LLM/AI Platforms (Remote)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 days agoStatus: Live

Job Summary

Design, build, and deploy AI infrastructure that powers large-scale training, fine-tuning, and inference for LLM-based security products. Provision GPU clusters, optimize model-serving frameworks, and lead model lifecycle management (versioning, checkpointing, reproducibility). Architect data platforms for LLMs, RAG, and agentic systems, define robust evaluation and MLOps/DataOps practices, and mentor engineers to deliver production-ready, performance-focused systems.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Provision and configure GPU clusters and compute resources for LLM training, fine-tuning, and inference workloads.
  • •Develop and optimize LLM model-serving infrastructure, including deployment and inference framework optimization.
  • •Lead model lifecycle management (versioning, checkpointing, reproducibility) across training and inference deployments.
  • •Architect and maintain data platforms and pipelines for LLMs, RAG, and AI agentic systems at scale.
  • •Define and enforce MLOps/DataOps best practices (monitoring, observability, zero-touch recovery) and mentor the team.

Pay and Benefits

Salary: USD 140,000 - 215,000 annually
Perks:Health Insurance401kPaid LeaveParental LeaveEquity

Key Requirements

  • •Bachelor’s degree in Computer Science, Data Engineering, or related STEM field; Master’s preferred.
  • •6+ years of experience in Infrastructure/Data Engineering, including at least 2 years building and maintaining LLM-supporting platforms and pipelines.
  • •Hands-on LLM infrastructure engineering experience: cluster provisioning, training workload optimization, and inference pipeline maintenance.
  • •Demonstrable ability to write clean, performant, well-tested code and deliver quickly without quality tradeoffs.
  • •Strong understanding of engineering practices (peer code reviews, resilient architecture) and technical leadership/mentorship.
Experience:6+ yearsAILLMsInfrastructure engineeringMLOpsDistributed systems
Education:Bachelor's
Skills:Technical leadershipMentorshipEngineering best practicesStakeholder communicationResults delivery
Tech Stack:MLflowSagemakerVertex AICUDANVIDIAGPUTPUVLLMTriton Inference ServerPyTorchRayMegatronJAXPythonDockerKubernetesSlurmAirflowTerraformAnsible

Company Brief

Crowdstrke
Provides cloud-native endpoint protection, threat intelligence, and security operations solutions that prevent breaches and stop sophisticated cyberattacks across endpoints, cloud workloads, identity, and APIs for enterprises worldwide.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Sunnyvale, United States
Founded: 2011
Glassdoor
Glassdoor: 4.4
WebsiteLinkedIn