Inference Software Engineer

Etched
San Jose
Workplace: OnsiteFull timeUSD 175,000 - 275,000 annuallyFunction: Software EngineeringSkills: ["Communication","Problem-solving"]

Join Etched to help port advanced transformer models to a cutting-edge architecture. You’ll build abstractions and testing for rapid model porting, scale multi-node inference, optimize routing, and use profiling tools to remove bottlenecks. Collaborate on a performance-focused runtime and work with C++/Rust and ML frameworks to push the limits of AI inference hardware.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Etched
Etched
1 year ago

Inference Software Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join Etched to help port advanced transformer models to a cutting-edge architecture. You’ll build abstractions and testing for rapid model porting, scale multi-node inference, optimize routing, and use profiling tools to remove bottlenecks. Collaborate on a performance-focused runtime and work with C++/Rust and ML frameworks to push the limits of AI inference hardware.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Port state-of-the-art models to Etched architecture and develop abstractions/testing for rapid porting.
  • •Build, enhance, and scale the runtime for multi-node inference, including state management and error handling.
  • •Optimize routing and communication layers using high-performance interconnects.
  • •Use profiling and debugging tools to identify bottlenecks and correctness issues.

Pay and Benefits

Salary: USD 175,000 - 275,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionHousing SubsidyRelocationWellness StipendMeal Allowance

Key Requirements

  • •Proficiency in C++ or Rust.
  • •Understanding of performance-sensitive or complex distributed software systems like Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects (e.g. NVLink, InfiniBand).
  • •Familiarity with PyTorch or JAX.
  • •Ported applications to non-standard accelerator hardware or hardware platforms.
Experience:AI hardwareDistributed systemsInference
Skills:CommunicationProblem-solving
Tech Stack:C++RustPyTorchJAXLinuxNVLinkInfiniBand

Company Brief

Etched
Designs transformer‑specialized AI inference ASICs (product: Sohu) to accelerate large‑language‑model workloads, working with TSMC for fabrication and targeting energy‑efficient inference performance.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Jose, United States
Founded: 2022
WebsiteLinkedInGlassdoor