AI Computing Research Intern

Relixir
Mountain View, United States
Workplace: RemoteInternshipUSD 4,000 - 6,000 monthlyFunction: Research & Scientific (R&D)Skills: ["High agency","Move fast","Feedback-driven","Systems thinking","ML engineering"]

Build and ship research that makes running thousands of AI agents dramatically cheaper, faster, and more reliable. Optimize self-hosted model inference (quantization, batching, speculative decoding, KV-cache strategy, and tensor/pipeline parallelism) and design model routing for cost/performance. Benchmark and deploy across GPUs, edge, on-prem, and alternative accelerators, then push results into production along with agent-infrastructure improvements and evaluation-driven iteration.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Relixir
Relixir
2 months ago

AI Computing Research Intern

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Build and ship research that makes running thousands of AI agents dramatically cheaper, faster, and more reliable. Optimize self-hosted model inference (quantization, batching, speculative decoding, KV-cache strategy, and tensor/pipeline parallelism) and design model routing for cost/performance. Benchmark and deploy across GPUs, edge, on-prem, and alternative accelerators, then push results into production along with agent-infrastructure improvements and evaluation-driven iteration.
Location: Mountain View, United States
Workplace: Remote
Employment Type: Internship
Job Function: Research & Scientific (R&D)
Seniority: Intern level

Key Responsibilities

  • •Research and ship systems that reduce the cost, speed up, and improve reliability of running thousands of AI agents.
  • •Optimize local/self-hosted model inference using techniques like quantization, batching, speculative decoding, and KV-cache strategies.
  • •Build model routing to send requests to the cheapest model that can do the job, using frontier APIs when needed and local otherwise.
  • •Benchmark and deploy across hardware (GPUs, edge, on-prem, alternative accelerators) and translate results into deployment decisions.
  • •Advance agent infrastructure (orchestration, caching, context management, parallelization) and prototype recursive self-improvement loops.

Pay and Benefits

Salary: USD 4,000 - 6,000 monthly

Key Requirements

  • •Strong systems + ML engineering skills and comfort building in Python and PyTorch, including profiling, optimization, and shipping.
  • •Deep understanding of how transformers run in practice, including attention, KV cache, and throughput vs. latency tradeoffs.
  • •Have built and shipped something real (side project, OSS, hackathon win, research artifact, or prior internship).
  • •Comfort using LLMs both as tools and as objects of study, including API calls, prompts, and tool use.
  • •High agency and ability to move fast, take feedback, and push back when you’re right.
Education:
Skills:High agencyMove fastFeedback-drivenSystems thinkingML engineering
Tech Stack:PythonPyTorchTorchCUDALLMsQuantizationSpeculative decodingKV-cacheTensor parallelismPipeline parallelismVLLMTensorRT-LLMSGLangLlama.cppTriton

Eligibility

Nationality:US National

Company Brief

Relixir
Develops AI-powered autonomous agents and orchestration tools that enable companies to automate complex workflows, integrate data sources, and build intelligent assistants for enterprise applications using large language models and custom connectors.
Industry: AI & Machine Learning
Website