Performance Tools Intern

Etched
San Jose
Workplace: OnsiteInternshipFunction: Healthcare (Clinical, Medical, Wellness)Skills: ["Problem-solving","Curiosity","Quick learning","Passion for performance","Collaboration"]

Develop and extend performance analysis and profiling tooling for a custom ML accelerator, helping teams understand workload behavior and pinpoint bottlenecks. Build components for data collection using hardware counters, execution traces, and memory behavior; create low-overhead tracing and correlated timelines across CPU and accelerator activities. Partner across hardware, compiler, firmware, and inference engineering to deliver visualization and analysis tools that improve performance and developer productivity.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Etched
Etched
3 days ago

Performance Tools Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Develop and extend performance analysis and profiling tooling for a custom ML accelerator, helping teams understand workload behavior and pinpoint bottlenecks. Build components for data collection using hardware counters, execution traces, and memory behavior; create low-overhead tracing and correlated timelines across CPU and accelerator activities. Partner across hardware, compiler, firmware, and inference engineering to deliver visualization and analysis tools that improve performance and developer productivity.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Healthcare (Clinical, Medical, Wellness)
Seniority: Intern level

Key Responsibilities

  • •Build components of performance analysis and profiling infrastructure.
  • •Collect and analyze performance data from custom ML accelerators, including hardware counters, execution traces, and memory behavior.
  • •Develop tooling to trace host-side runtime activity, system behavior, and accelerator execution.
  • •Correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads.
  • •Create analysis and visualization tools that help engineers identify bottlenecks and optimize models.

Key Requirements

  • •Strong programming skills in C++ or Rust; Python experience is a plus.
  • •Solid understanding of computer architecture, including CPUs/GPUs/AI accelerators, memory hierarchies, and parallel programming.
  • •Experience or strong interest in low-level performance analysis, profiling, and performance optimization.
  • •Familiarity with performance analysis tools such as Nsight, VTune, Xprof, Perfetto, or similar tools is a plus.
  • •Experience or strong interest in operating systems, compilers, firmware, drivers, or other low-level systems software.
Experience:ML acceleratorsInferencePerformance profilingDistributed workloadsHardware counters
Skills:Problem-solvingCuriosityQuick learningPassion for performanceCollaboration
Tech Stack:C++RustPythonNsightVTuneXprofPerfettoPCIeLinuxWindowsGPUsTPUs

Company Brief

Etched
Designs transformer‑specialized AI inference ASICs (product: Sohu) to accelerate large‑language‑model workloads, working with TSMC for fabrication and targeting energy‑efficient inference performance.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series A
Headquarters: San Jose, United States
Founded: 2022
WebsiteLinkedInGlassdoor