I would rather say: "benchmark with confidence intervals" or maybe "benchmark with comparisons".

It's very hard to say much about an absolute number. You need to compare against some alternative or control, and because of CPU load, throttling, GC, and a thousand other variables, you really should be comparing against that control _in the same run_, and importantly round robin across multiple runs to spread out the noise fairly across each implementation.

Then once you have a bunch of measurements you have a distribution and shouldn't just take a mean to compare, but should calculate something like the 95% confidence interval. If you see that those confidence intervals overlap, the you might not really know which is faster. If they don't overlap, then you probably do know which is faster.

If you have a good benchmark, then running it more times can narrow the confidence intervals and let you tease out very small improvements at the cost of longer runs. If confidence intervals don't narrow, then you hit the limits of signal-to-noise.

This is the only way I've been able to get reliable, actionable benchmarks outside of a very, very controlled hardware lab. It's what Google's Tachometer benchmark runner does, and I wish more runners did this: https://github.com/google/tachometer

Instead of giving users a confidence interval, ask the user to specify a p value. Then run something like Welch’s t test.

[deleted]

mine is: "benchmark in load and wall-insensitive witnesses"

it's tempting, but otherwise you're just attesting the limits of the current environment

Really depends on what you benchmark and how reliable you want the measurement and what domain you are benchmarking. For example, criterion can sometimes spend quite a bit of time because it needs to stabilize the measurements. The author’s claim is “well you don’t need that accuracy” but I’ve seen wasted time chasing ghosts or making claims on performance improvements that were either neutral or net negative due to this noise.

> Anything faster than, say, 10ms risks being skewed by fixed costs (e.g, interpreter startup).

Sounds like the author’s experience is strictly in Python. For example with Java you have to make sure the JIT has sufficient optimized your program.

Additionally there’s plenty of situations where it can take a really long time to generate a representative dataset worth benchmarking and it can take time to evaluate the performance (eg databases). Short and quick microbenchmarks can be useful as building points, but at some point you need to evaluate steady state performance of the full thing. Other domains this comes up with is game rendering performance where a 300ms sample tells you nothing about whether you have frame drops after minute 25 or have a memory leak.

Agree that the domain and time scales matter a lot. When working on systems where nanoseconds were quite important, time was given significant attention.

However in general - I agree with OP. Most of the time I care about milliseconds.

> Sounds like the author’s experience is strictly in Python. Er [1], no [2].

[1] https://github.com/rust-lang/rust-analyzer

[2] https://github.com/tigerbeetle/tigerbeetle

Some optimizations like hoisting still improve code readability even if the benchmark are inconclusive. And of course how a piece of code behaves in vivo and under test can vary quite a bit in both directions. Particularly with space/time tradeoffs, where you reduce or increase cache pressure with code running concurrently to your code.

A lot of my meditations on optimization date back to a profiler telling me that a redundant function call was responsible for 5% of the run time of a task. After removing it, run time decreased by 20%. Then I had to think about all the ways in which profilers can lie. It’s still a black art after all this time.

The tools tell you whether it might be worthwhile to look at a problem, but keeping your work is a completely different matter entirely. Unfortunately some people get Sunk Cost Fallacy, or worry about losing face, so once committed to a course will see it merged into the codebase whether it does anything or not. And they will push harder if they win the lottery and one test run says theirs is much faster. Nevermind that the next ten runs show the opposite.

> how a piece of code behaves in vivid and under test can vary quite a bit

I've never seen "in vivid" used this way. Are you thinking of "in vivo" which is from Latin meaning "in life" or "in living" distinguished against Latin "in vitro" meaning "in glass" referring to the glass petri dishes or beakers used to do science experiments in a laboratory?

That’s because Apple autocorrect is a long con to prove that humans are too stupid to be left in charge. We can’t even english good. I assure you I did not type “in vivid”.

In high volume systems, 10ms is kind of crazy. I’ve run systems with operation metrics in the ms scale and the server side latency was lower than 1ms (computing business logic, or heavily cached data). Client side was closer to 3-5ms. As you mentioned, this was Java.

There also needs to be care taken in how these measurements are aggregated. Averages will almost always tell you nothing. High percentiles (95%, 99%, 99.9%) under load may show you something completely different than the average or even median case.

You can still benchmark say 100 requests and time that. For JVM for example you'll have to consider the JIT effects and GC and stuff that only happens later. So 100x 1ms requests is still a very small benchmark that's probably unreliable

[deleted]

> human-perceptible range allows me to use my intuitive sense of time and speed

I didn’t intend to specialize in performance early in my career but it happened anyway. I tuned boring homework assignments to make them more interesting for myself. But I moved far away for my first gig out of school and when I showed up the UI painted so slow it looked like one of those videos of an artist drawing something by hand but sped up. I hid my panic at having my name associated with this stinking pile and as soon as I’d done a couple of challenging bug fixes to prove I wasn’t an idiot I got to work.

My first dozen changes needed no benchmarks, In part because I had the slowest machine in the office. I could literally count seconds in my head and tell that I’d taken >1/4 of a second off of an operation because I made it a syllable or two farther in the old version. Later on I used the stopwatch function on a handheld device, to catch 100ms differences. I was there for nearly two months before I needed to put console output of (end - start) into the code for the first time.

By the time I started running out of stuff I knew how to fix, the corpus of data had started showing a serious scalability problem in the data filtering operations, so we were back into classical architectural misdeeds. We had a 2n logn² intersection test that did two scans on different criteria and compared the results using a quadratic time comparison. This code was copy pasta’d in dozens and dozens of places around the project, with slight variations in variable names and parameter marshaling. I replaced all the copies with one function they did filter(filter(x)) over a long holiday weekend since I had nobody local. -500 lines of code and much much lower slope of call time on multi year data sets.

Benchmarking like that is often broken because of continuous CPU core clock speed adjustments, system interrupts, SMIs, etc.

I tried to fix it by switching hyperthreading off, playing with the scaling governor, boost, setting a CPU frequency to no avail. The jitter was too much and the results were not reproducible, so I just gave up.

Of course your mileage may vary; this was on an AMD Zen 3 CPU.

Intel published some guidance on doing precise benchmarks. Basically run your code inside the kernel, turn off interrupts, use the cpuid instruction to prevent out-of-order execution, use rdtscp instruction instead of rdtsc, etc.

> prevent out-of-order execution

Won't this make the results inaccurate?

The way you work around this (apart from doing what you can to make the system as predictable as possible) is to accept that the data is noisy, and then working around it by doing multiple trials and following up with a statistical analysis on the results.

You can use the desired confidence to inform the warm-up, number of trials, benchmark duration, and so on.

If you end up with a multimodal distribution it can be worth tracking percentiles.

I did warm-up rounds, to prime caches and so on. I fiddled with statistics: IIRC simple average seemed to give the best results overall. I thought keeping the lowest n results would be the best, but that turned out to be wrong.

The results were indeed multimodal.

All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.

Used to do this sort of thing for computations that needed to run in the 10 microsecond range (HFT stuff), circa 2008. Had very predictable results because:

a) language was not garbage collected (C++)

b) we avoided heap lock contentions in critical paths by pre-allocating object pools at startup

c) I/O operations were offloaded to separate threads, connected by mutex locked linked lists

d) processing thread was bound to its own CPU core

That's about as deterministic as we could get.

Did all of those trying to reduce the jitter. I think it was about the CPU clock not being stable.

Additionally I avoided core 0, because it was the noisiest and did some cgroups core pinning for the test workload.

There were zero page faults during the test runs and the CPU core was uncontested by other threads.

I think in 2008 CPUs were not so crazy about power and heat management.

The benchmarks for 'xc' - "my" compiler for heterogenous computing (gpu and cpu, all in the same language, compiler managing data-hazards between them) are all on the order of a second to a few seconds, for the long pole.

So if I'm comparing against (say) C++, Swift, ObjC, on a given (arm64 or x86_64) architecture, I want to see that sort of timescale, a bit less is fine, a bit more is fine.

Of course, when you (today [grin]) get auto-vectorisation of 2D matrix multiplies, and you have an SME/SME2 target on arm64 that most compilers don't pick up so you're 150x faster than clang/g++, you might have to run it a bit longer so you can get reasonable comparison numbers :)

1: https://compile-xc.org/compiler/performance/

One of the smartest guys I know said: The longer you let a benchmark run, the more certain you can be that it captures other unrelated activity.

His theory was you want to run a benchmark several times, for short periods of time, and take the fastest run as the benchmark.

I ended up doing a lot of benchmarking for the Python Need For Speed Sprint, and that advice seemed to work out well for us there.

This is basically the philosophy of https://github.com/c-blake/bu/blob/main/doc/tim.md , but it uses a somewhat fancier estimator of the minimum than the sample min of dt's (which in some sense is guaranteed to be "a little" high).

There are other benchmarking points in that document about measurement footguns that can arise from targeting a specific duration like the 200..400ms of TFA with an easy but careless problem scale-up (like the Ben Hoyt example of footnote 4).

With that tool, in user space with just CPU freq pinning, I routinely see CPU bound times that are stable to single digit microseconds month to month on the same machine and see 0.01% to 0.2% effects on 10ms scale activity (yes, 1..20 bps) and the reported uncertainty usually captures the variation all right, but the distribution is not really Gaussian/Normal and instead has more shape parameters and/or you would want a 95% CI or some such as mentioned elsethread here.

In creating https://github.com/bhouston/webgpu-bench, a microbenchmark for WebGPU, I found that if you run in browsers you do not control, you need to run longer than 10ms to get an accurate result. Modern browsers by default round to the nearest 1ms. [1] So I generally try to get 50 to 100ms of run time for browser benchmarks.

[1] https://developer.mozilla.org/en-US/docs/Web/API/Performance...

If you are benchmarking Java, there is Java Microbenchmark Harness (https://github.com/openjdk/jmh), which:

- Does several warm-up runs so the JIT-optimized code is benchmarked instead of the interpreted code/compilation.

- Create multiple forks of the JVM to eliminate JVM run variance.

- Provide utilities like `Blackhole` to prevent dead code elimination and `State` to do setup and prevent constant folding.

Don't roll your own criteria. Use well established frameworks to do measurements for you; https://benchmarkdotnet.org/ in the dotnet world for example.

Love it. Unfortunately, for some benchmarks, it can be bit difficult to get representative inputs which take hundreds of milliseconds. I'm curious as to why the author doesn't loop 1-10ms inputs to deal with variance? Which also deals with startup costs. Rust microbench harnesses were already good at this stuff-out of-the-box (run-to-run variance, etc) several years ago.

Looping works, though reusing the same small buffer mostly measures hot-cache performance.

This has to be read in terms of https://xkcd.com/2400/. If you're chasing low-hanging fruit, and searching for a 3x or 10x speedup on something that has never been optimized, this advice probably can be ok.

The harder you push, and the more you need to start finding smaller improvements, the more this advice becomes a rule of thumb you can't rely on.

I’ve only seen two projects brag about making a <2% performance improvement and that was crate and the guy working on pointer compression for v8.

I have mostly always known better and either report aggregate numbers, like 15% across six changes, or in milliseconds, like TTFB down from 600ms to 590ms. The tricky bit with the former is that you want to cluster changes to the same call tree or cross cutting concern so that the testing surface area of each additional change is only a small increment over the costs already incurred by the first change. Essentially amortizing the potential for uncaught regressions and delays on the release process across a handful of small changes over the same interactions.

Because if you make small changes scattered across the entire codebase, the testing cost and the risk of escape balloon, which is why people try to stop you from doing incremental improvements at all.

On the project where I figured this out, I delivered 30% improvements every milestone for 2 years before I started scraping the bottom of the barrel, moving from subsystem to subsystem and fixing everything I knew how to fix. Sometimes that was three changes, sometimes it was ten. The second time around was taking lessons learned back into areas I’d covered prior to learning them, but mostly looking for regressions that had arrived in new features and bug fixes.

Microseconds are a much more convenient unit when you are working with inter-thread communication concerns.

I also prefer microseconds when working with SQLite and AspNetCore.

Assuming the subject of the benchmark is a web/API request here, otherwise the advice does not really apply. Milliseconds would be too large for benchmarking GPU- or CPU-intensive work.

[deleted]

if you are doing benchmarks using interpreted languages but care about smaller timescales perhaps your problem is using interpreted languages.

[dead]

[flagged]

[deleted]