Senior Solutions Architect – Large Scale AI Inference

NVIDIA
Switzerland, Germany, France, Poland, Spain, United Kingdom
Workplace: RemoteFull timePLN 292,500 - 507,000Function: Solutions Engineering & Sales EngineeringExperience: 5+ yearsEducation: mastersSkills: ["Technical leadership","Collaboration","Technical communication","Customer engagement"]

Guide EMEA AI Natives to deploy and optimize large-scale AI inference on multi-node GPU clusters. Architect inference pipelines for dense and MoE (dense/sparse/latent) models, improving performance across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large deployments. Partner with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) and lead community technical workshops, hackathons, and reference architectures to set scalable inference direction.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Solutions Architect – Large Scale AI Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Guide EMEA AI Natives to deploy and optimize large-scale AI inference on multi-node GPU clusters. Architect inference pipelines for dense and MoE (dense/sparse/latent) models, improving performance across quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for large deployments. Partner with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) and lead community technical workshops, hackathons, and reference architectures to set scalable inference direction.
Location: Switzerland, Germany, France, Poland, Spain, United Kingdom
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Guide EMEA customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters.
  • •Architect efficient inference pipelines for dense and sparse/latent MoE models across thousands of GPUs.
  • •Improve inference efficiency across quantization (INT4/FP8), speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP.
  • •Collaborate with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) to accelerate customer success.
  • •Animate the AI inference developer community in EMEA via technical workshops, hackathons, and reference architectures.

Pay and Benefits

Salary: PLN 292,500 - 507,000

Key Requirements

  • •MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience.
  • •5+ years of experience optimizing Neural Networks inference.
  • •Solid understanding of transformer inference optimization (quantization, disaggregated inference, speculative decoding, continuous batching, KV cache optimization).
  • •Practical experience performing MoE inference at scale (expert parallelism, WideEP, all-to-all communication, routing overhead, load balancing).
  • •Ability to engage effectively with ML engineers, researchers, and systems architects at a deep technical level.
Experience:5+ yearsAIMachine learningHigh-performance computing
Education:Master's
Skills:Technical leadershipCollaborationTechnical communicationCustomer engagement
Tech Stack:NVIDIA DynamoTensorRT-LLMNIXLMixture-of-Experts (MoE)INT4FP8Speculative decodingDisaggregated inferencePrefill/decodeKV cacheWideEPTransformersContinuous batchingNVLinkInfiniBandRDMAUCXDynamoAll-to-all communicationLoad balancing

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor