Staff AI Engineer, Model Post-Training and Alignment

OKX
Hong Kong, Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 8+ yearsEducation: bachelorsSkills: []

Hands-on ML engineer focused on post-training and alignment of large language models, leading end-to-end post-training pipelines (supervised fine-tuning, preference optimization, RL-based methods), and deploying low-latency inference. Collaborates across data strategy, reward modeling, and production deployment to improve model performance, controllability, and domain adaptation within an APAC-focused team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OKX
OKX
4 months ago

Staff AI Engineer, Model Post-Training and Alignment

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Hands-on ML engineer focused on post-training and alignment of large language models, leading end-to-end post-training pipelines (supervised fine-tuning, preference optimization, RL-based methods), and deploying low-latency inference. Collaborates across data strategy, reward modeling, and production deployment to improve model performance, controllability, and domain adaptation within an APAC-focused team.
Location: Hong Kong, Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Lead and execute the full post-training pipeline for large language models, including supervised fine-tuning, preference optimization, and reinforcement learning–based methods.
  • •Design and implement advanced training paradigms such as DPO and GRPO.
  • •Develop domain-specific data recipes, curation strategies, and augmentation pipelines to optimize task performance.
  • •Conduct post-training of specialized small models from scratch, including architecture selection, dataset construction, and optimization strategy.
  • •Build and refine Reward Models to support alignment and downstream optimization.

Key Requirements

  • •Bachelor's in Computer Science, AI, Machine Learning, or related fields
  • •At least 8 years of industry experience
  • •Hands-on experience across the full post-training pipeline for large models (supervised fine-tuning, preference optimization, RL-based methods)
  • •Deep familiarity with DPO (Direct Preference Optimization), GRPO (Generalized Reward Policy Optimization), and RL-based post-training methodologies
  • •Experience deploying models in low-latency production environments using frameworks such as vLLM and SGLang
Experience:8+ yearsAIMachine LearningLLMsNLP
Education:Bachelor's
Languages:English
Tech Stack:VLLMSGLangDirect Preference OptimizationGeneralized Reward Policy OptimizationReinforcement learningSupervised fine-tuningRLAIFLow-latency serving

Company Brief

OKX
Global cryptocurrency exchange and Web3 technology provider offering spot and derivatives trading, custody, wallet services, and blockchain infrastructure to retail and institutional users worldwide.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Headquarters: Victoria, Seychelles
Founded: 2017
WebsiteLinkedIn