Senior Technical Product Manager - AI Agents, Evals & Reliability

BJAK
London
Workplace: HybridFull timeFunction: Product ManagementSkills: ["Judgment","Decision-making","Independent execution","Prioritization"]

Own end-to-end requirements and execution for AI-agent features focused on long-running reliability, persistent context, and real-world task completion. Translate model behavior, data constraints, and evaluation results into clear system and product decisions. Define evaluation frameworks across offline metrics and online experiments, drive hard trade-offs across quality, latency, cost, and reliability, and partner closely with ML, backend, and mobile engineers to ship safely with strong feedback loops.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
BJAK
BJAK
1 day ago

Senior Technical Product Manager - AI Agents, Evals & Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Own end-to-end requirements and execution for AI-agent features focused on long-running reliability, persistent context, and real-world task completion. Translate model behavior, data constraints, and evaluation results into clear system and product decisions. Define evaluation frameworks across offline metrics and online experiments, drive hard trade-offs across quality, latency, cost, and reliability, and partner closely with ML, backend, and mobile engineers to ship safely with strong feedback loops.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: Product Management
Seniority: Mid level

Key Responsibilities

  • •Research and define end-to-end AI system requirements from capability to behavior to user impact.
  • •Translate model capabilities, data constraints, and evaluation results into product and system decisions.
  • •Make hard trade-offs across quality, latency, cost, reliability, and UX.
  • •Define and evolve evaluation frameworks across offline metrics, online experiments, and human feedback.
  • •Own product quality end-to-end (correctness, predictability, and user trust) and drive safe, reliable execution with strong feedback loops.

Key Requirements

  • •Strong computer science fundamentals, including algorithms, data structures, and system design.
  • •Solid understanding of ML fundamentals and how modern AI systems behave in production.
  • •Hands-on exposure to AI-powered products, including LLM-based systems, with model evaluation, prompt/pipeline iteration, and feedback loops.
  • •Strong intuition for model limitations, including hallucinations, bias, and drift.
  • •Significant experience owning complex technical products end-to-end with strong judgment in ambiguous, fast-moving environments.
Experience:LLM-based systemsAI-powered products
Skills:JudgmentDecision-makingIndependent executionPrioritization

Company Brief

BJAK
Bjak is a Malaysia-based insurtech that operates an online insurance comparison and distribution platform across SEA, simplifying purchase and servicing of motor and life insurance using digital tools and AI-enabled services.
Industry: InsurTech
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Selangor, Malaysia
Founded: 2019
Glassdoor
Glassdoor: 2.5
WebsiteLinkedInGlassdoor