AI Benchmark Engineer, Inference Performance

Silicon Data
Silicon Data

Software Engineering, Data Science · Contractor,

United States · Remote

Posted on Sep 15, 2026

AI Benchmark Engineer, Inference Performance

Silicon Data · SiliconMark · Remote (8 AM – 1 PM ET overlap)

About Silicon Data

Silicon Data builds the data infrastructure for the global compute economy.

As demand for AI infrastructure accelerates, compute is becoming a tradable asset class. Silicon Data provides price transparency, benchmarking, and standardized market indices that let participants discover price, manage risk, and make data-driven decisions across the compute supply chain.

Our indices and data power decisions for hedge funds and financial institutions trading compute exposure, AI companies managing infrastructure cost, and data centers optimizing asset utilization. We are backed by leading global trading firms including DRW and Jump Trading, Valor Atreides AI Fund (VAAI), and we recently partnered with CME Group to launch the first-ever compute futures market built on Silicon Data’s indices.

About SiliconMark

SiliconMark is the measurement layer for transacting compute. Renting a GPU from a cloud provider comes with one standard spec sheet, but our research has found that actual delivered performance can vary by as much as 34% for the same GPU. That gap is a problem for buyers and sellers alike.

SiliconMark runs comprehensive, third-party performance benchmarks that let compute market participants settle on the real delivered value of what they are buying — not what the spec sheet claims.

The Role

We are hiring an AI Benchmark Engineer to build and run the inference performance side of SiliconMark.

Serving the same model on the same hardware can produce wildly different results depending on the inference engine, quantization scheme, batch policy, KV cache configuration, and a dozen other choices. Your job is to make those differences measurable, reproducible, and defensible — so that a throughput or latency number carrying the SiliconMark name holds up when a NeoCloud’s engineering team, a silicon vendor, or a hedge fund’s analyst tries to pick it apart.

This is a hands-on engineering role. You will write the harnesses, run the workloads, chase down the anomalies, and own the methodology that explains why the numbers are what they are.

What You’ll Own

  • Build and maintain the inference benchmarking harness. Automated, containerized, reproducible runs across inference engines (vLLM, TensorRT-LLM, SGLang, TGI), model families, and hardware targets.

  • Own the inference metrics that matter. Time-to-first-token, inter-token latency, sustained and peak throughput, latency percentiles under concurrency, goodput under SLO constraints, and their expression per dollar and per watt.

  • Design workload profiles that reflect real usage. Input and output length distributions, concurrency patterns, streaming versus batch, prefill-heavy versus decode-heavy — rather than synthetic best-case runs.

  • Capture and validate the configuration fingerprint. Engine version, model revision and quantization, driver and firmware, kernel and runtime versions, tensor and pipeline parallelism, scheduler settings. A result without its full configuration is not a result.

  • Establish run repeatability and statistical rigor. Define warmup and steady-state criteria, variance thresholds, minimum run counts, and outlier handling. Know when a run is invalid and be willing to throw it out.

  • Investigate performance anomalies. When a cohort underperforms, determine whether the cause is thermal, network, contention, misconfiguration, or a genuine silicon difference — and produce the evidence.

  • Partner with our data and GPU benchmarking engineers on cohort definition, cross-suite consistency, and how inference results align with hardware-level benchmarks.

  • Write the methodology. Public-facing write-ups, versioned methodology documents, and technical documentation that a skeptical customer can audit and reproduce.

  • Work with our pricing products so measured inference performance can be expressed in economic terms — cost of capability per measured token per second.

  • Track the field. New engines, new serving techniques, new model architectures, and emerging public benchmarks worth incorporating or explicitly rejecting.

Who Audits Your Work

SiliconMark results are read by people with strong incentives to challenge them. You should expect your numbers to be scrutinized by:

  • AI platform and infrastructure teams choosing a model and serving stack against latency SLOs.

  • Enterprise AI leaders deciding between self-hosting and API access.

  • FinOps, procurement, and CFO organizations using our data as negotiating leverage.

  • NeoClouds and inference providers seeking independent certification — or contesting an unflattering result.

  • Silicon challengers and accelerator startups validating against NVIDIA baselines.

  • Financial institutions and analysts trading compute exposure on our indices.

What We’re Looking For

Required

  • 4+ years of engineering experience, with substantial hands-on work in LLM inference, model serving, or performance engineering.

  • Strong Python. You write automation and tooling that other people can run and trust, not one-off scripts.

  • Direct experience with at least one production inference engine (vLLM, TensorRT-LLM, SGLang, TGI, or equivalent), including configuring and tuning it — not just calling its API.

  • Working knowledge of what actually drives inference performance: batching and scheduling, KV cache behavior, quantization, tensor and pipeline parallelism, memory bandwidth limits, and prefill versus decode characteristics.

  • Comfort with Linux, containers, and running workloads on GPU infrastructure.

  • Genuine measurement discipline. You are skeptical of your own results, you control variables, you distinguish signal from noise, and you can explain the uncertainty in a number.

  • Clear technical writing. You can document a methodology decision well enough that an outside engineer can reproduce it and a hostile reviewer cannot dismiss it.

  • Self-direction. The methodology here is not written yet. You will be defining the standard, not executing someone else’s spec.

Nice to Have

  • GPU profiling and telemetry experience — Nsight, DCGM, driver-level APIs, or equivalent.

  • CUDA or kernel-level familiarity.

  • Exposure to MLPerf Inference or comparable standardized benchmark suites.

  • Experience with distributed or multi-node inference.

  • Familiarity with non-NVIDIA accelerators (AMD, TPU, Cerebras, Groq, or custom silicon).

  • Experience building software that runs as an agent inside a customer environment, and the deployment, versioning, and trust problems that come with it.

  • Background in certification, attestation, provenance, or audit-grade data products.

  • Understanding of GPU hardware economics, cloud compute pricing, or AI infrastructure procurement.

  • Early-stage startup experience.

Engagement and Location

  • Remote. We are open to W-2 employment, onshore 1099 contract, or offshore corp-to-corp engagement — tell us which structure you are looking for and we will work with it.

  • Overlap requirement. You must be available during 8:00 AM – 1:00 PM Eastern Time, Monday through Friday. Outside that window, work when you work best.

  • Output over presence. We care about the quality and reproducibility of what you ship, not hours logged.

What You Can Expect From Us

  • Competitive compensation and, for W-2 hires, meaningful equity.

  • Ownership of the measurement standard for a market where no standard yet exists.

  • A small, senior team where your work has direct, visible impact and your opinions carry real weight.

  • Comprehensive health, dental, and vision benefits for W-2 employees.

  • A flexible, remote-friendly environment with a culture that values output over presence.

  • A founder-led culture built on merit — your growth here is limited only by the quality of your work.

Equal Opportunity

Silicon Data is proud to be an equal opportunity employer. We are committed to building a team that reflects a diversity of backgrounds, perspectives, and experiences. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, disability, age, or veteran status.