Software Engineer, Data Infrastructure

Cohere
New York, Toronto, San Francisco, Montreal
Workplace: RemoteFull timeFunction: Software EngineeringExperience: 4+ yearsSkills: ["Python","Kubernetes","S3","GCS","POSIX","Apache Beam","Spark","Flink","BigQuery","Airflow","Dbt"]

Build and maintain the high-performance data layer for AI training workloads, working on petabyte-scale storage infrastructure and collaborating with researchers and engineers to transform unstructured data into performant datasets across S3, GCS, and POSIX. Apply Python, Kubernetes storage primitives, and distributed processing frameworks to enable state-of-the-art AI models, with a remote-friendly, inclusive culture and strong focus on impact.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cohere
Cohere
4 months ago

Software Engineer, Data Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Build and maintain the high-performance data layer for AI training workloads, working on petabyte-scale storage infrastructure and collaborating with researchers and engineers to transform unstructured data into performant datasets across S3, GCS, and POSIX. Apply Python, Kubernetes storage primitives, and distributed processing frameworks to enable state-of-the-art AI models, with a remote-friendly, inclusive culture and strong focus on impact.
Location: New York, Toronto, San Francisco, Montreal
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and maintain petabyte-scale storage infrastructure and the networking/perf challenges it entails.
  • •Collaborate daily with researchers and engineers to design and optimize data storage systems.
  • •Transform unstructured data into performant datasets across multiple storage backends (S3, GCS, POSIX).
  • •Leverage distributed data processing frameworks to enable scalable data workflows (Beam, Spark, Flink).
  • •Contribute to the data infrastructure that powers AI training and evaluation workloads.

Pay and Benefits

Perks:Health InsuranceDentalRemote WorkMeal AllowanceParental LeavePaid Leave

Key Requirements

  • •4+ years of experience working on data storage infrastructure
  • •Strong command of Python
  • •Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.)
  • •The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX
  • •Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink
Experience:4+ yearsData storageStorage infrastructureCloud
Skills:PythonKubernetesS3GCSPOSIXApache BeamSparkFlinkBigQueryAirflowDbt
Languages:English
Tech Stack:PythonKubernetesS3GCSPOSIXApache BeamSparkFlinkBigQueryAirflowDbt

Company Brief

Cohere
Builds security-first foundation models and enterprise AI products (LLMs, retrieval, agent platforms) for regulated industries, enabling customizable, private deployments across cloud and on-premises for real-world business applications.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: Toronto, Canada
Founded: 2019
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor