Senior Consultant Specialist (Model Hosting/Inference Optimization) (Xi'an, SN, CN, -)

HSBC
Guangzhou, Xi'an
Workplace: OnsiteFull timeFunction: Consulting & AdvisoryExperience: 3+ yearsSkills: ["Collaboration","Problem-solving","AI-native mindset"]

Build and operate scalable model hosting platforms for LLMs, embeddings, and speech models, and drive inference optimization to improve latency, throughput, and cost. Design end-to-end fine-tuning pipelines to adapt foundation models using domain datasets, then integrate and validate results with researchers and engineers. Work across heterogeneous hardware, evaluate inference frameworks, and ensure production-grade reliability, monitoring, and performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
HSBC
HSBC
2 days ago

Senior Consultant Specialist (Model Hosting/Inference Optimization) (Xi'an, SN, CN, -)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and operate scalable model hosting platforms for LLMs, embeddings, and speech models, and drive inference optimization to improve latency, throughput, and cost. Design end-to-end fine-tuning pipelines to adapt foundation models using domain datasets, then integrate and validate results with researchers and engineers. Work across heterogeneous hardware, evaluate inference frameworks, and ensure production-grade reliability, monitoring, and performance.
Location: Guangzhou, Xi'an
Workplace: Onsite
Employment Type: Full time
Job Function: Consulting & Advisory
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate scalable, reliable model hosting platforms for LLMs, embeddings, and STT/TTS across heterogeneous hardware.
  • •Drive inference optimization for latency, throughput, and cost using techniques like quantisation, KV-cache optimization, and dynamic/continuous batching.
  • •Evaluate, integrate, and tailor inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to maximize performance on target hardware.
  • •Own inference health and performance monitoring (latency, throughput, TTFT, memory, availability) and troubleshoot bottlenecks and deployment issues.
  • •Build end-to-end fine-tuning pipelines and work with data scientists/domain experts to define objectives, validate results, and integrate fine-tuned models into hosting/inference systems.

Key Requirements

  • •Bachelor’s/Master’s/PhD in ML/NLP/CS/Data Science/Statistics (or related).
  • •3 years on AI platforms covering model hosting/inference optimization and fine-tuning pipelines; LLM experience strongly preferred.
  • •Strong engineering skills in Python and CUDA, with solid understanding of GPU/CPU architecture and HPC fundamentals.
  • •Deep inference expertise including KV-cache, batching, quantisation (INT4/FP8/GPTQ/AWQ), operator optimization, and framework integration (vLLM, TensorRT-LLM, SGLang); hands-on hosting with Docker/Kubernetes and AWS/GCP/Azure.
  • •End-to-end fine-tuning expertise: data prep, distributed training, hyperparameter tuning, HF/Accelerate/LoRA/QLoRA, plus benchmarking, monitoring, and troubleshooting.
Experience:3+ yearsLLMsAI platforms
Education:
Skills:CollaborationProblem-solvingAI-native mindset
Tech Stack:PythonCUDAGPUCPUHPCLLMsEmbeddingsSTTTTSKV-cacheBatchingQuantisationINT4FP8GPTQAWQVLLMTensorRT-LLMSGLangDocker

Company Brief

HSBC
Global banking and financial services organisation offering retail, commercial, corporate and investment banking, wealth management, and global markets services across Europe, Asia, the Americas and the Middle East.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: London, United Kingdom
Founded: 1865
Glassdoor
Glassdoor: 3.6
WebsiteLinkedIn