Principal Machine Learning Engineer

BJAK
United States
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Technical leadership","Problem-solving","Independent execution","Pragmatic decision-making","Iteration and continuous improvement"]

Design and evolve mission-critical machine learning systems across training, inference, evaluation, and deployment. Architect scalable training pipelines on GPU infrastructure, build reliable low-latency inference systems, and maintain data systems for synthetic and real-world training. Own production deployment with GPU optimization, memory efficiency, and scaling policies, while building evaluation for robustness, safety, and bias. Collaborate to integrate ML into backend and client products.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
BJAK
BJAK
3 days ago

Principal Machine Learning Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 minutes agoStatus: Live

Job Summary

Design and evolve mission-critical machine learning systems across training, inference, evaluation, and deployment. Architect scalable training pipelines on GPU infrastructure, build reliable low-latency inference systems, and maintain data systems for synthetic and real-world training. Own production deployment with GPU optimization, memory efficiency, and scaling policies, while building evaluation for robustness, safety, and bias. Collaborate to integrate ML into backend and client products.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Sr. Manager level

Key Responsibilities

  • •Architect and build large-scale ML systems across data, training, evaluation, inference, and deployment.
  • •Design reproducible, high-performance training pipelines across GPU infrastructure.
  • •Architect inference systems balancing latency, throughput, cost, and reliability at scale.
  • •Implement evaluation pipelines for performance, robustness, safety, and bias in partnership with research leadership.
  • •Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies.

Key Requirements

  • •Strong background in deep learning and transformer-based architectures.
  • •Hands-on experience training, fine-tuning, or deploying large-scale ML models in production.
  • •Proficiency with at least one modern ML framework (e.g., PyTorch, JAX) and ability to learn others quickly.
  • •Experience with distributed training and inference frameworks (e.g., DeepSpeed, FSDP, Megatron, ZeRO, Ray).
  • •Experience with GPU optimization (memory efficiency, quantization, mixed precision) and writing production-grade systems.
Experience:Deep learningTransformersLLMsGPUDistributed trainingSynthetic dataOpen source
Skills:Technical leadershipProblem-solvingIndependent executionPragmatic decision-makingIteration and continuous improvement
Tech Stack:PyTorchJAXDeepSpeedFSDPMegatronZeRORayVLLMTensorRT-LLMFasterTransformerApache ArrowSparkRLHFPPODPOORPOQuantizationMixed precisionGPU optimization

Company Brief

BJAK
Bjak is a Malaysia-based insurtech that operates an online insurance comparison and distribution platform across SEA, simplifying purchase and servicing of motor and life insurance using digital tools and AI-enabled services.
Industry: InsurTech
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Selangor, Malaysia
Founded: 2019
Glassdoor
Glassdoor: 2.5
WebsiteLinkedInGlassdoor