Senior Solutions Architect – Large Scale Neural Networks Inference

NVIDIA
France, United Kingdom, Spain, Poland, Switzerland
Workplace: RemoteFull timePLN 292,500 - 507,000Function: Solutions Engineering & Sales EngineeringExperience: 8+ yearsEducation: phdSkills: ["Technical leadership","Communication","Stakeholder alignment","Technical collaboration","Roadmap feedback"]

Define the technical direction for large-scale AI inference across EMEA by leading engagements from proof of concept to production. Drive inference strategy for customer portfolios by identifying bottlenecks (latency, efficiency, cost per token, memory use, networking) and architecting optimized pipelines using NVIDIA inference backends. Translate customer deployment insights into actionable product feedback for the NVIDIA stack, collaborating across research, engineering, and executives.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Solutions Architect – Large Scale Neural Networks Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Define the technical direction for large-scale AI inference across EMEA by leading engagements from proof of concept to production. Drive inference strategy for customer portfolios by identifying bottlenecks (latency, efficiency, cost per token, memory use, networking) and architecting optimized pipelines using NVIDIA inference backends. Translate customer deployment insights into actionable product feedback for the NVIDIA stack, collaborating across research, engineering, and executives.
Location: France, United Kingdom, Spain, Poland, Switzerland
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead inference strategy for a portfolio of EMEA AI Natives customers from proof of concept through production-scale deployments.
  • •Identify and analyze inference challenges across deployments, including latency, efficiency, cost per token, memory utilization, and low-latency networking.
  • •Architect and optimize high-performance inference pipelines to improve GPU utilization and AI cluster efficiency using NVIDIA Dynamo, TensorRT-LLM, vLLM, SGLang, and other backends.
  • •Convert customer insights and deployment patterns into actionable product feedback that informs roadmap development for the NVIDIA inference stack.
  • •Collaborate with NVIDIA and customer teams to influence strategic technology decisions for next-generation AI inference at scale.

Pay and Benefits

Salary: PLN 292,500 - 507,000

Key Requirements

  • •MS or PhD in Computer Science, Engineering, or equivalent experience.
  • •8+ years in AI/ML infrastructure with deep expertise in LLM/VLM inference optimization and production deployment at scale.
  • •Deep understanding of transformer inference acceleration, including quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and WideEP for MoE models.
  • •Understanding of GPU memory hierarchies and low-latency networking and how they impact inference performance.
  • •Proven track record to lead technical initiatives and communicate effectively with research scientists, infrastructure engineers, and executive stakeholders.
Experience:8+ yearsAI/MLLLMVLMInferenceKubernetesAI infrastructureProduction deployment
Education:PhD / Doctorate in Computer Science, Engineering
Skills:Technical leadershipCommunicationStakeholder alignmentTechnical collaborationRoadmap feedback
Tech Stack:NVIDIA DynamoTensorRT-LLMVLLMSGLangInference backendsINT4FP8Speculative decodingDisaggregated inferenceContinuous batchingKV cache optimizationWideEPMoE modelsGPU memory hierarchiesLow-latency networkingTriton Inference ServerKServeNIMKubernetesGPU orchestration

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor