Senior Consultant Specialist (Model Hosting/Inference Optimization) (Guangzhou, GD, CN, 510620)
HSBC
Guangzhou
Workplace: OnsiteFull timeFunction: Consulting & AdvisoryExperience: 3+ yearsEducation: bachelorsSkills: ["Python","CUDA","GPU","HPC","Docker","Kubernetes","VLLM","TensorRT-LLM","SGLang","LLM","Embedding","STT","TTS","Quantisation","KV-cache","Batching","LoRA","QLoRA","Accelerate"]Lead the design and operation of scalable model hosting platforms for LLMs, embeddings, and STT/TTS, while building end-to-end fine-tuning pipelines. Collaborate with AI researchers and product teams to deliver production-grade inference optimized for latency, throughput, and cost, with strong emphasis on reliability, security, and cross-hardware deployment.

