Inference Intern

Etched
San Jose
Workplace: OnsiteInternshipFunction: Architecture & Urban PlanningSkills: ["Problem-solving","Teamwork","Analytical thinking"]

Architecture intern contributes to next-generation AI accelerators by developing compute architectures for transformer workloads, working on performance modeling, porting models, and optimizing runtime and communication. You’ll collaborate on HW instructions, model architecture operations, and high-performance software components, with mentorship and hands-on experience in a cutting-edge hardware-in-the-loop environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Etched
Etched
9 months ago

Inference Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Architecture intern contributes to next-generation AI accelerators by developing compute architectures for transformer workloads, working on performance modeling, porting models, and optimizing runtime and communication. You’ll collaborate on HW instructions, model architecture operations, and high-performance software components, with mentorship and hands-on experience in a cutting-edge hardware-in-the-loop environment.
Location: San Jose
Workplace: Onsite
Employment Type: Internship · 3 months
Job Function: Architecture & Urban Planning

Key Responsibilities

  • •Port state-of-the-art models to our architecture and help build abstractions and testing capabilities to rapidly iterate on model porting.
  • •Build, enhance, and scale the runtime, including multi-node inference, intra-node execution, state management, and robust error handling.
  • •Optimize routing and communication layers using Sohu’s collectives.
  • •Use performance profiling and debugging tools to identify bottlenecks and correctness issues.
  • •Co-design HW instructions and model architecture operations to maximize model performance; implement high-performance software components for the Model Toolkit.

Pay and Benefits

Perks:Paid InternshipHousing SupportMeal Allowance

Key Requirements

  • •Progress towards a Bachelor’s, Master’s, or PhD degree in computer science, computer engineering, or a related field
  • •Proficiency in C++ or Rust
  • •Understanding of performance-sensitive or complex distributed software systems, e.g. Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects
  • •Familiarity with PyTorch or JAX
  • •Ported applications to non-standard accelerator hardware or hardware platforms
  • •Deep knowledge of transformer model architectures and/or inference serving stacks (vLLM, SGLang, etc.)
Experience:TransformersDistributed systemsInference
Education:
Skills:Problem-solvingTeamworkAnalytical thinking
Tech Stack:C++RustPyTorchJAXLinuxGPUsTPUsNVLinkInfiniBand

Company Brief

Etched
Designs transformer‑specialized AI inference ASICs (product: Sohu) to accelerate large‑language‑model workloads, working with TSMC for fabrication and targeting energy‑efficient inference performance.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Jose, United States
Founded: 2022
WebsiteLinkedInGlassdoor