Senior Consultant Specialist (Model Hosting/Inference Optimization) (Xi'an, SN, CN, -)
Guangzhou, Xi'an
Workplace: OnsiteFull timeFunction: Consulting & AdvisoryExperience: 3+ yearsSkills: ["Collaboration","Problem-solving","AI-native mindset"]Build and operate scalable model hosting platforms for LLMs, embeddings, and speech models, and drive inference optimization to improve latency, throughput, and cost. Design end-to-end fine-tuning pipelines to adapt foundation models using domain datasets, then integrate and validate results with researchers and engineers. Work across heterogeneous hardware, evaluate inference frameworks, and ensure production-grade reliability, monitoring, and performance.
Loading
Loading job details...
Preparing the role view and application actions.

