Senior Modeling Architect, Performance Benchmarking

Neurophos
Austin, Sunnyvale
Workplace: OnsiteFull timeUSD 210,000 - 250,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Ownership","Attention to detail","Documentation","Collaboration","Analytical thinking"]

Own the benchmarking numbers that drive architecture and product decisions for an optical inference accelerator. You will define and continuously update measurement methodologies across roofline/limiter models, architecture performance models, in-house RTL simulation, and end-to-end measured runs on competing GPUs/accelerators. Establish workload definitions, build reproducible benchmark harnesses, measure key latency/throughput/energy metrics, and keep results aligned as models and hardware generations evolve.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Neurophos
Neurophos
2 days ago

Senior Modeling Architect, Performance Benchmarking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own the benchmarking numbers that drive architecture and product decisions for an optical inference accelerator. You will define and continuously update measurement methodologies across roofline/limiter models, architecture performance models, in-house RTL simulation, and end-to-end measured runs on competing GPUs/accelerators. Establish workload definitions, build reproducible benchmark harnesses, measure key latency/throughput/energy metrics, and keep results aligned as models and hardware generations evolve.
Location: Austin, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own performance and energy metrics architecture, product, and leadership rely on, keeping them consistent across modeling fidelities and measured hardware.
  • •Produce comparable benchmark numbers for the same workloads across roofline/limiter models, architecture performance models, RTL simulation, and measured competitor hardware.
  • •Maintain workload definitions constant across fidelities (model/application, sequence length, batch, precision, prefill vs decode, and relevant parallelism settings).
  • •Bring up and measure inference workloads from Hugging Face, PyTorch, papers, and stacks such as vLLM, SGLang, TensorRT-LLM, and Triton.
  • •Measure and report TTFT, inter-token latency, tokens per second, tokens per second per watt, and energy; document disagreements and maintain a reviewed internal benchmark suite.

Pay and Benefits

Salary: USD 210,000 - 250,000 annually
Perks:Health InsuranceHsa ContributionsPaid Leave401kEquityDentalVisionLife InsuranceCritical IllnessAccident Insurance

Key Requirements

  • •BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
  • •5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement.
  • •Experience building benchmark harnesses that produce measured results on real GPUs/accelerators, including converting Hugging Face models/papers into runnable benchmarks.
  • •Hands-on skills in roofline analysis, limiter analysis, or analytical performance modeling.
  • •Proficiency in Python for harnesses/plots and comfort working in Linux, plus knowledge of LLM inference stacks (prefill vs decode, MoE, continuous batching) and cloud GPU operations (AWS/GCP/Azure).
Experience:5+ yearsAI hardwareGPU performanceHPCML systemsLLM inferenceAccelerator benchmarking
Education:
Skills:OwnershipAttention to detailDocumentationCollaborationAnalytical thinking
Tech Stack:Silicon photonicsProgrammable metasurfaceHugging FacePyTorchVLLMSGLangTensorRT-LLMTriton Inference ServerIn-house RTL simulationRoofline analysisLimiter analysisNvidia-smiDCGMPower cappingNvidia Nsight SystemsNvidia Nsight ComputeHBMFP16BF16FP8

Company Brief

Neurophos
Develops metamaterial-based photonic optical processing units (OPUs) to deliver high-performance, energy-efficient AI inference chips for datacenters, aiming to scale photonic compute to exaflop levels.
Industry: Deep Tech
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: Austin, United States
Founded: 2020
WebsiteLinkedIn