Site Reliability Engineer

Schonfeld
New York
Full timeUSD 175,000 - 225,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Problem-solving","Communication","Root cause investigation","Translating technical concepts"]

Join the AI Technology team to ensure the stability and scalability of an internal Agentic AI platform. Define reliability standards with SLOs, error budgets, and incident response runbooks, and own observability, incident response, and scalability. Help keep agents, gateways, LLM proxies, and RAG pipelines highly available and efficient, support users through a help channel, and improve code quality and developer practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Schonfeld
Schonfeld
15 hours ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Join the AI Technology team to ensure the stability and scalability of an internal Agentic AI platform. Define reliability standards with SLOs, error budgets, and incident response runbooks, and own observability, incident response, and scalability. Help keep agents, gateways, LLM proxies, and RAG pipelines highly available and efficient, support users through a help channel, and improve code quality and developer practices.
Location: New York
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Set reliability standards for Enterprise AI by defining SLOs, error budgets, and incident response runbooks
  • •Own observability, incident response, reliability, and scalability of the AI platform
  • •Ensure agents, gateways, LLM proxies, and RAG pipelines operate with high availability, accuracy, and financial efficiency
  • •Support users via a help channel by investigating root causes, providing solutions, monitoring upstream dependencies, and communicating updates
  • •Improve code quality and best practices by identifying development lifecycle gaps and continuously enhancing the platform and developer workflows

Pay and Benefits

Salary: USD 175,000 - 225,000 annually

Key Requirements

  • •5+ years of professional cloud automation or site reliability engineering experience, or similar
  • •Familiarity with software development best practices and experience creating applications in Python
  • •Experience creating and integrating REST APIs, and event-driven and asynchronous architectures
  • •Experience with both relational (Postgres, MySQL) and NoSQL databases (DynamoDB, Elastic Search)
  • •Experience with AWS services, Kubernetes-driven deployments, GitHub Actions, and modern CI/CD pipelines
Experience:5+ yearsCloud automationSite reliability engineeringAIAgentic AIHedge fundMachine learning
Skills:Problem-solvingCommunicationRoot cause investigationTranslating technical concepts
Tech Stack:PythonREST APIsEvent-driven architectureAsynchronous architecturePostgresMySQLDynamoDBElastic SearchAWSS3OpenSearchKubernetesGitHub ActionsCI/CDObservabilityDataDogSLOsError budgetsRAG pipelines

Company Brief

Schonfeld
Schonfeld is a multi‑strategy investment firm that deploys capital across systematic and discretionary strategies across equities, macro, and quantitative trading, serving institutional and private investors globally.
Industry: Hedge Funds
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: New York, United States
Website