Associate - AI Tooling Ops - Platform Reliability Engineer

Jefferies
Pune
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsEducation: bachelorsSkills: ["Analytical","Problem-solving","Communication","Stakeholder engagement","Self-motivated"]

Join the global Platform Reliability Engineering team as an Associate Platform Reliability Engineer (AI Tooling Ops). You’ll design, build, and maintain AI tooling infrastructure on AWS Kubernetes, monitor health and availability, triage incidents, and run post-incident reviews. Partner across engineering and business teams to improve resilience and operational visibility, reduce toil through automation, and strengthen deployment, monitoring, alerting, and observability using Grafana, Datadog, Prometheus, and OpenTelemetry.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Jefferies
Jefferies
1 day ago

Associate - AI Tooling Ops - Platform Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join the global Platform Reliability Engineering team as an Associate Platform Reliability Engineer (AI Tooling Ops). You’ll design, build, and maintain AI tooling infrastructure on AWS Kubernetes, monitor health and availability, triage incidents, and run post-incident reviews. Partner across engineering and business teams to improve resilience and operational visibility, reduce toil through automation, and strengthen deployment, monitoring, alerting, and observability using Grafana, Datadog, Prometheus, and OpenTelemetry.
Location: Pune
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain AI Tooling infrastructure running on AWS Kubernetes with a focus on stability and scalability.
  • •Monitor platform health, proactively identify risks, and improve system reliability, performance, and availability.
  • •Perform incident triage and troubleshooting, including communication and post-incident reviews to prevent recurrence.
  • •Develop automation to reduce operational toil and eliminate manual support activities.
  • •Build and enhance deployment, monitoring, alerting, and observability capabilities, including dashboards, alerts, and service health monitoring.

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • •3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.
  • •Strong programming and scripting experience in Python, Go, C#, Java, or C++.
  • •Solid understanding of software engineering principles, system design, and distributed applications in production.
  • •Hands-on observability and monitoring experience with platforms such as Grafana, Datadog, Prometheus, and OpenTelemetry.
Experience:3+ yearsSite reliability engineering (SRE)Platform reliability engineering (PRE)DevOpsProduction supportDistributed systemsObservability
Education:Bachelor's
Skills:AnalyticalProblem-solvingCommunicationStakeholder engagementSelf-motivated
Tech Stack:AWSKubernetesLinux/UnixWindows ServerPythonGoC#JavaC++GrafanaDatadogPrometheusOpenTelemetryLokiJaegerKafkaRedisMongoDBElasticsearchDocker

Company Brief

Jefferies
Global investment banking and capital markets firm providing advisory, underwriting, trading, research, and asset management services to corporations, governments, and institutional investors.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: New York, United States
Founded: 1962
WebsiteLinkedIn