← All jobs

Applied Researcher – AI Expert

Designworks Talent · Bellevue

Open. First seen 2 October 2026.

APPLY on jobs.ashbyhq.com →

Description

Applied Researcher – AI Expert

Location: Hybrid | Bellevue, WA (downtown)

About the Opportunity

Our client is seeking an experienced AI Expert / Applied Researcher to help shape how a fast-growing technology organization understands and applies the rapidly evolving landscape of AI models, architectures, inference technologies, and accelerator systems.

This role sits at the intersection of AI research, systems engineering, and infrastructure strategy. You will evaluate emerging technologies, translate research into practical engineering implications, and help guide decisions around AI infrastructure, inference optimization, model serving, accelerator platforms, and the economics of delivering AI workloads at scale.

This is a highly strategic individual contributor role with broad technical influence. You will work closely with engineering, product, infrastructure, finance, and commercial teams to help determine what technologies to build, adopt, partner for, or invest in.

What You'll Do

  • Hold the company's view of where AI is going. Track lab and academic research, model releases, open-source projects, vendor roadmaps, and the startup and venture landscape across the model landscape, inference optimization, serving and routing, the kernel and compiler layer, post-training and adaptation, the agentic layer, evaluation, and the security and sovereignty constraints around all of it. Right now that means questions like what low-precision formats really cost in output quality, whether sparse attention solves long-context economics, how far open-weight models displace frontier APIs for serving volume, and what agent traffic does to caching and scheduling. Those specific questions will have changed within a quarter — holding the current version of them is the job.
  • Formulate and validate the product and engineering thesis. Turn that view into a defensible position on what we build, buy, or partner for across routing, serving, and the kernel and runtime layer, pressure-tested against measured cost per token and output quality — and say so plainly when the evidence does not hold up.
  • Own the company's technical position on inference: optimization across quantization, speculative decoding, KV-cache management, batching, prefill/decode disaggregation and long-context serving; the serving stack and what we adopt, extend, or build ourselves; and the intelligent routing logic that decides which model and which silicon serves each request.
  • Own multi-silicon portability: what it actually takes to run the same model well across NVIDIA, AMD and Cerebras, and where the compiler, kernel, and runtime layer is worth building versus buying or partnering for.
  • Own the token economics model — cost per million tokens by model, silicon, and traffic profile — and the evaluation and observability that keep it honest, with quality, latency, throughput, and cost measured continuously rather than benchmarked once. Finance, pricing, and sales will rely on this.
  • Make the work land commercially. Feed product and go-to-market with what the token factory can actually offer, support sales in demanding customer conversations about model fit, performance and cost, provide technical diligence on inference partners, white-label providers and prospective tuck-in or acquire hire targets, and advise on how IP developed elsewhere in the group is best leveraged here.

What We're Looking For

Required Qualifications

  • Significant hands-on experience with modern AI models,
  • Depth in inference rather than training alone,
  • Working fluency in one or more accelerator ecosystem, and
  • Hands-on depth at the compiler, kernel, or runtime layer (CUDA, Triton, ROCm/HIP, XLA, or similar).
  • Working fluency in more than one silicon ecosystem. CUDA plus ROCm, Cerebras, or another accelerator experience combined with a realistic view of what portability actually costs.

Preferred Qualifications

  • Fluency with the landscape you would be scanning: the frontier labs and open-weight model providers, the serving and inference startups, the silicon vendors, and the research groups doing the work that lasts — and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.
  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources — papers, model cards, vendor roadmaps, your own benchmarking — and producing a defensible position under genuine uncertainty. A PhD in a relevant field is one good route to this and is common among people with real depth here; sustained industry research, open-source contribution at depth, or a body of internal technical assessments that changed real decisions are equally valid. Either way, the role turns on the second half: translating that work for engineering, product, go-to-market, and finance, because it informs all four.
  • Significant experience working with AI, machine learning, or AI model technologies.
  • Strong understanding of AI model architectures and how models are developed.
  • Ability to understand both the research and engineering implications of emerging AI technologies.
  • Experience working with one or more major model types, such as language, vision, audio, or multimodal models.
  • Hands-on depth in inference rather than only training — serving, optimization, and the practical work of getting latency and cost down without giving up quality.
  • Working knowledge of a modern serving stack (vLLM, SGLang, TensorRT-LLM, or equivalent) and of quantization, batching, and KV-cache techniques in production.
  • Ability to reason quantitatively about cost to serve, and to build models of it that hold up to scrutiny from finance and commercial teams.
  • Strong foundational knowledge that allows you to quickly understand unfamiliar model architectures and research.
  • Ability to read technical papers and translate research concepts into practical engineering implications.
  • Strong analytical, communication, and influencing skills.
  • Ability to operate across engineering, research, and infrastructure organizations.
  • A track record of collaborating with researchers and engineers across groups and levels to shape long-term research directions and move research into practice.
  • Comfortable operating as an individual contributor with high ownership in a lean, early-stage team.

Nice to Have Qualifications

  • Depth in a specific kernel or runtime domain beyond general familiarity — attention kernels, collective libraries, memory allocators, or graph compilers.
  • Background at a hyperscaler, neocloud, or AI lab operating inference infrastructure at production scale.
  • Familiarity with distributed training frameworks (e.g., PyTorch Distributed, DeepSpeed, Megatron). Useful context, though training is not a target-state focus for us.
  • Experience with agentic frameworks, tool-use protocols, or multi-agent orchestration.
  • Publications, patents, or a recognized external presence in the AI research or systems community.
  • Spent your education and career working deeply in AI, but you have the foundational knowledge and intellectual curiosity to quickly understand new architectures and technologies as they emerge.

Location

  • Hybrid role based in downtown Bellevue, WA.
  • Approximately three days per week in the office.
  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.
  • U.S. work authorization is required. Visa sponsorship is not currently available.
  • Export control: this role involves technologies subject to U.S. export control regulations. Candidate eligibility may be subject to export control screening and, where applicable, licensing.
  • Travel: Willingness and ability to travel as needed internationally to data centers and co-locations (up to 25%)

Why Join?

  • High-impact technical role: Directly influence the technology direction of an organization building AI infrastructure at scale.
  • Ground-floor opportunity: Help establish technical strategy, architecture, processes, and culture within a growing organization.
  • High ownership: Operate as a senior individual contributor with substantial autonomy and direct access to senior technical leadership.
  • Cross-disciplinary exposure: Work across AI models, inference, accelerators, software systems, networking, infrastructure, and economics.
  • Cutting-edge technical problems: Work on multi-accelerator inference, intelligent routing, performance optimization, token economics, and compiler, kernel, and runtime technologies.
  • Research with practical impact: Turn emerging research and technology developments into decisions that directly affect engineering, product, commercial strategy, and investment.
  • Lean, senior environment: Work with a small group of highly experienced technical contributors rather than within a large management hierarchy.

Similar jobs

Is this your posting and you want it taken down? Email info@careerholo.com.