Staff Modeling Architect

Neurophos
Austin, Sunnyvale
Workplace: OnsiteFull timeUSD 250,000 - 290,000 annuallyFunction: Software EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Analytical judgment","Modeling methodology ownership","Mentorship","Cross-team collaboration","Technical documentation"]

Build and validate production-grade inference workloads for a hardware/software co-design accelerator. You’ll bind Hugging Face and PyTorch models to a functional and programming/runtime model stack, then drive performance, energy, and PPA predictions through roofline/limiter analysis, design space exploration, and cycle-accurate simulation. Own workload methodology, mentor model engineers, and keep model outputs consistent with RTL simulation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Neurophos
Neurophos
2 days ago

Staff Modeling Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Build and validate production-grade inference workloads for a hardware/software co-design accelerator. You’ll bind Hugging Face and PyTorch models to a functional and programming/runtime model stack, then drive performance, energy, and PPA predictions through roofline/limiter analysis, design space exploration, and cycle-accurate simulation. Own workload methodology, mentor model engineers, and keep model outputs consistent with RTL simulation.
Location: Austin, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Bring up inference workloads as they ship, including dense/MoE transformers, attention/KV cache, expert routing, quantization, and hybrid/SSM models; cover retrieval, speech, vision, and recommendation workloads where applicable.
  • •Bind Hugging Face and PyTorch workloads to the programming model and runtime, run them on the functional model, and ensure software and architecture show matching behavior.
  • •Co-design tiling, scheduling, ISA, SRAM/HBM hierarchy, NoC traffic, and multi-chip mapping across pipeline/tensor/sequence parallelism and collectives.
  • •Develop Python energy/latency models and implement bit-accurate C++ functional models for optical GEMM, SRAM vector processors, dataflow engines, and HBM so bring-up can start before tape-out.
  • •Run roofline/limiter analysis and design space exploration, implement cycle-approximate/cycle-accurate performance/power/area (PPA) models, align with RTL via Verilator/SystemVerilog co-simulation, and set modeling methodology for a workload area.

Pay and Benefits

Salary: USD 250,000 - 290,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceHsaPaid Leave401kEquityDentalVisionLife InsuranceCritical IllnessAccident Insurance

Key Requirements

  • •BS, MS, or PhD in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
  • •8+ years in hardware/functional modeling, performance modeling, performance simulation, or accelerator performance analysis used by architects, RTL, compiler/runtime, or silicon teams.
  • •Track record shipping a model or study another team depended on.
  • •Strong grounding in computer architecture, microarchitecture, memory systems, and AI accelerators (GPU/TPU/NPU/custom SoC).
  • •Modern C++ (C++17+) plus Python modeling with NumPy, Pandas, and Matplotlib; experience with discrete-event, cycle-approximate, or cycle-accurate simulators (e.g., SystemC/gem5/SST or custom kernels).
Experience:8+ yearsSilicon photonicsAI hardwareHardware modelingFunctional modelingPerformance simulation
Education:Bachelor's
Skills:Analytical judgmentModeling methodology ownershipMentorshipCross-team collaborationTechnical documentation
Tech Stack:C++C++17PythonNumPyPandasMatplotlibPyTorchHugging FaceVerilatorSystemVerilogSystemCGem5SSTMLIRTVMXLAONNXMcPATCACTIHBM

Company Brief

Neurophos
Develops metamaterial-based photonic optical processing units (OPUs) to deliver high-performance, energy-efficient AI inference chips for datacenters, aiming to scale photonic compute to exaflop levels.
Industry: Deep Tech
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: Austin, United States
Founded: 2020
WebsiteLinkedIn