AI Platform Engineer, Capabilities

Brain Co.
San Francisco
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["System design","Observability mindset","Collaboration","Problem-solving","Ownership"]

Build and operate shared platform backend services and data pipelines that power AI products, owning the full lifecycle from architecture through deployment and long-term maintenance. Design scalable systems for ML experiment tracking, artifact management, and automated training/evaluation pipelines. Create highly available, fault-tolerant services with strong observability to meet enterprise and government SLAs, partnering across engineering, product, and ML research to reduce time-to-ship.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Brain Co.
Brain Co.
3 months ago

AI Platform Engineer, Capabilities

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and operate shared platform backend services and data pipelines that power AI products, owning the full lifecycle from architecture through deployment and long-term maintenance. Design scalable systems for ML experiment tracking, artifact management, and automated training/evaluation pipelines. Create highly available, fault-tolerant services with strong observability to meet enterprise and government SLAs, partnering across engineering, product, and ML research to reduce time-to-ship.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate platform backend services and data pipelines powering AI products, owning the full lifecycle from architecture through maintenance.
  • •Build systems to accelerate AI product development, including scalable ML experiment tracking, artifact management, and automated training/evaluation pipelines.
  • •Engineer highly available, fault-tolerant systems with deep observability to meet uptime and latency SLAs.
  • •Create modular, scalable architectures and clean APIs (REST, gRPC), continuously optimizing for latency, throughput, and cloud compute costs.
  • •Partner with engineering, product, and ML research teams to build shared platform capabilities that remove bottlenecks and speed up delivery.

Pay and Benefits

Perks:Equity401kHealth InsuranceDentalVisionPaid LeaveMeal AllowanceCommuter Benefits

Key Requirements

  • •3+ years building and scaling production backend services or platforms, ideally with Python, Go, Rust, or C++.
  • •Deep understanding of consistency, availability, distributed failure modes, and idempotency.
  • •Build intuitive, well-documented, highly maintainable APIs and shared platforms.
  • •Own services with real uptime expectations; design for observability and be comfortable with on-call.
  • •Break down open-ended problems into clear system designs from first principles to production-ready systems.
Experience:3+ yearsAIMLDistributed systemsPlatform engineering
Skills:System designObservability mindsetCollaborationProblem-solvingOwnership
Tech Stack:PythonGoRustC++RESTGRPCWeights & BiasesMLflowClearMLRayFlyteKubeflowTemporalKafkaSparkFlinkSOC2FedRAMP

Company Brief

Brain Co.
Brain Co. builds agent-native operating systems for regulated institutions across government, health, insurance, finance, and enterprise. Its platform emphasizes secure deployment, workflow integration, and AI applications tailored to complex operational environments.
Industry: AI & Machine Learning
Growth: Growth Stage Startup
Headquarters: San Francisco, United States
Founded: 2026
WebsiteLinkedIn