- name
- aspnet-benchmark
- description
- Run the TechEmpower ASP.NET Core benchmarks against a locally-built dotnet/runtime, either locally with an SDK overlay or on external infrastructure using Crank, to validate the end-to-end impact of a runtime change (e.g. sockets, io_uring, GC, JIT) under realistic HTTP load. Use this when asked to benchmark, load-test, or profile an ASP.NET Core / Kestrel scenario, or to compare an env-var-gated feature (e.g. `DOTNET_USE_IO_URING`) end-to-end.
# Running ASP.NET Core Benchmarks
Use the benchmark apps from the [`aspnet/Benchmarks`](https://github.com/aspnet/Benchmarks) repository to validate runtime changes under realistic HTTP server load. Microbenchmarks can isolate an operation's cost, but end-to-end ASP.NET benchmarks show how a change affects throughput, latency, and CPU usage under concurrent request/response processing.
These benchmarks are useful for changes to the socket/IO stack, GC, JIT, and other components whose impact depends on application behavior. Both workflows below run the benchmark against a locally-built runtime: start with local measurements to validate ideas, then use external infrastructure for final validation when access is available.
Use both of these implementations when assessing a runtime change:
- `PlatformBenchmarks` exercises a highly optimized raw Kestrel `HttpApplication`.
- `Minimal` exercises the ASP.NET Core Minimal APIs stack.
A final comparison must show that neither implementation regresses. An improvement in one does not compensate for a regression in the other.
## Common Prerequisites
Both workflows require a validated local dotnet/runtime build containing the change under test, targeting the benchmark machine's OS and architecture: the development machine for local runs, or the application agent for Crank runs. **Both the native runtime and managed libraries must be built in Release.**
Reuse existing validated artifacts when available; do not rebuild merely to stage binaries. Consult the `build-and-test` skill before building, and never build when the user requested documentation or explicitly prohibited builds. If a build is needed, build runtime and libraries in Release with `-c release`, not just `-rc release`: the latter only forces the runtime to Release and leaves libraries at their default (`Debug`), producing a mismatched testhost.
```bash
./build.sh clr+libs -c release
```
Confirm the artifacts correspond to the intended source revision. Switching Git branches does not update existing build outputs. Record source commits, any uncommitted changes, configuration, and binary hashes.
## Choosing a Benchmark
For both workflows, prefer `/json` for representative maximum-throughput comparisons. The standard TechEmpower `/plaintext` scenario uses HTTP pipelining (16 requests batched per connection), whereas `/json` sends one request per round trip and better represents normal request/response traffic. Their throughput numbers are not directly comparable.
Both `/json` and `/plaintext` can run locally without a database. Other benchmarks, such as fortunes, single/multiple queries, and updates, require setting up a database. These can also run locally, but database setup is outside the scope of this document.
## Benchmarking a Local Runtime Build Locally
**Always validate ideas locally first.** This workflow provides quick verification on the development machine and requires no VPN or access to external infrastructure. Drive the application with `wrk`, or [`bombardier`](https://github.com/codesenberg/bombardier) on Windows (because `wrk` is Linux/macOS-only).
Treat local results as an initial signal, not a definitive performance measurement: the application, load generator, and other processes share the same machine and compete for CPU, memory, and other resources. This contention can distort throughput and latency, so local results may not predict performance on dedicated benchmark machines.
### Prerequisites
- A load-generation tool: on Linux, `wrk` (`sudo apt install wrk` or build from [source](https://github.com/wg/wrk)). `wrk` doesn't build/run on Windows; use [`bombardier`](https://github.com/codesenberg/bombardier) instead (`go install github.com/codesenberg/bombardier@latest`, or download a prebuilt binary from its [releases page](https://github.com/codesenberg/bombardier/releases)) - the examples below use `wrk` syntax, with the `bombardier` equivalent noted alongside.
- The [`aspnet/Benchmarks`](https://github.com/aspnet/Benchmarks) repository, which contains the TechEmpower and other benchmark apps.
### Step 1: Clone the Benchmarks Repository
```bash
git clone https://github.com/aspnet/Benchmarks.git
```
The relevant apps are:
- `src/BenchmarksApps/TechEmpower/PlatformBenchmarks` — a raw Kestrel `HttpApplication` implementation of the TechEmpower benchmark suite (JSON serialization, plaintext, fortunes, single/multiple queries, updates).
- `src/BenchmarksApps/TechEmpower/Minimal` — an ASP.NET Core Minimal APIs implementation of the same TechEmpower benchmark suite.
### Step 2: Build the Apps Against the Repo's Own SDK
Build both apps with the dotnet/runtime repo's own SDK (`<runtime-repo>/.dotnet/dotnet`), not a system-wide SDK, so they target the same TFM as your local build:
```bash
TFM='<tfm>'
cd <benchmarks-repo>/src/BenchmarksApps/TechEmpower/PlatformBenchmarks
<runtime-repo>/.dotnet/dotnet build -c Release -p:TargetFrameworks="$TFM"
cd <benchmarks-repo>/src/BenchmarksApps/TechEmpower/Minimal
<runtime-repo>/.dotnet/dotnet build -c Release -p:TargetFramework="$TFM" -p:TargetFrameworks="$TFM"
```
This produces `bin/Release/<tfm>/PlatformBenchmarks.dll` and `bin/Release/<tfm>/Minimal.dll` under their respective project directories.
### Step 3: Make the Repo's SDK Run Against Your Local Runtime Build
`PlatformBenchmarks` needs `Microsoft.AspNetCore.App`, which the freshly-built `artifacts/bin/testhost` does **not** contain (it only has `Microsoft.NETCore.App`) — running it directly via `testhost`'s `corerun`/`dotnet` fails with a "no framework found" error. The repo SDK's own `dotnet` (under `<runtime-repo>/.dotnet`) already has `Microsoft.AspNetCore.App`, so the simplest way to combine "the ASP.NET Core framework" with "your locally-built `Microsoft.NETCore.App`" is to temporarily overlay the SDK's shared `Microsoft.NETCore.App/<version>` folder with the testhost's freshly-built one:
```bash
SDK_VERSION='<version>'
SDK_FX="<runtime-repo>/.dotnet/shared/Microsoft.NETCore.App/$SDK_VERSION"
TESTHOST_FX="<runtime-repo>/artifacts/bin/testhost/<tfm>-<os>-Release-<arch>/shared/Microsoft.NETCore.App/$SDK_VERSION"
# ALWAYS back up the original SDK framework files first, so they can be restored exactly.
# Use a fresh, uniquely-named backup directory every time: a fixed, reused path is unsafe
# if a previous run was interrupted before restoring — `mkdir -p` would silently succeed
# and the overlay step below would overwrite the only pristine backup with mutated files.
BACKUP_DIR="/tmp/netcoreapp_sdk_backup.$$"
mkdir "$BACKUP_DIR" # fails loudly (no -p) if this exact path somehow already exists
rsync -a --delete "$SDK_FX"/ "$BACKUP_DIR"/
# Overlay with the local build.
rsync -a --delete "$TESTHOST_FX"/ "$SDK_FX"/
```
**This mutates the repo's own SDK shared framework in place.** Never skip the backup step, and always restore it when done (Step 6) — leaving it mutated will silently break every other use of that SDK on the machine. Keep `$BACKUP_DIR` until Step 6 has verified the restored contents.
### Step 4: Run Each App and Verify It's Serving Requests
Run one app at a time because both commands below listen on port 5000. Start with `PlatformBenchmarks`:
```bash
cd <benchmarks-repo>/src/BenchmarksApps/TechEmpower/PlatformBenchmarks/bin/Release/"$TFM"
<runtime-repo>/.dotnet/dotnet exec PlatformBenchmarks.dll --urls http://127.0.0.1:5000 &
sleep 5
curl --fail-with-body -sS http://127.0.0.1:5000/json
```
After running its load tests, stop `PlatformBenchmarks`, then run `Minimal`:
```bash
cd <benchmarks-repo>/src/BenchmarksApps/TechEmpower/Minimal/bin/Release/"$TFM"
<runtime-repo>/.dotnet/dotnet exec Minimal.dll --urls http://127.0.0.1:5000 &
sleep 5
curl --fail-with-body -sS http://127.0.0.1:5000/json
```
Set whatever env var/`AppContext` switch you're comparing (e.g. `DOTNET_USE_IO_URING=1`/`=0`) *before* starting the process — it's read once at startup.
### Step 5: Run `wrk` (or `bombardier` on Windows) Against Each App
```bash
wrk -t12 -c256 -d15s --latency http://127.0.0.1:5000/json
wrk -t12 -c256 -d15s --latency http://127.0.0.1:5000/json
```
Run the first `/json` command while `PlatformBenchmarks` is running and the second while `Minimal` is running.
- `-t`: number of `wrk` threads — a good default is the machine's core count.
- `-c`: number of concurrent connections — 256 is a reasonable default for a many-concurrent-connections scenario.
- `-d15s`: run duration; 15s is normally enough to get a stable number.
- `--latency`: also reports the latency percentile breakdown, not just aggregate throughput.
On Windows, use `bombardier` instead, with equivalent options (run the first command against `PlatformBenchmarks` and the second against `Minimal`):
```powershell
bombardier -c 256 -d 15s -l http://127.0.0.1:5000/json
bombardier -c 256 -d 15s -l http://127.0.0.1:5000/json
```
- `-c`: concurrent connections (same meaning as `wrk -c`).
- `-d`: run duration (same meaning as `wrk -d`).
- `-l`: print the latency percentile breakdown (equivalent to `wrk --latency`).
- `bombardier` has no direct equivalent to `wrk -t`; it manages its own worker goroutines internally.
Run **at least two** runs per configuration (there's meaningful run-to-run variance) and report both, along with p50/p99 latency, not just a single throughput number. Present the results in a table.
Repeat Steps 4-5 for each configuration being compared (e.g. once with the env var on, once with it off) and for both apps, reusing the same overlay — only the running process needs to be restarted between configurations, not the overlay. Do not conclude that a change is beneficial unless neither app regresses.
### Step 6: Clean Up
1. Kill the server process — find the actual `dotnet exec ... PlatformBenchmarks.dll` or `dotnet exec ... Minimal.dll` PID (not the shell that launched it) with `pgrep -af PlatformBenchmarks` or `pgrep -af Minimal.dll`, then use `kill <pid>`.
2. **Restore the SDK's shared framework from the backup** and verify it with a recursive content comparison before considering the machine clean:
```bash
rsync -a --delete "$BACKUP_DIR"/ "$SDK_FX"/
diff --recursive --brief "$BACKUP_DIR" "$SDK_FX" # no output and exit code 0 means an exact content restore
```
3. Remove `$BACKUP_DIR` and any log files only after the recursive comparison succeeds.
### Profiling a Benchmark Run with `perfcollect`
If the load-test numbers show a difference (or don't, and you need to know why) and you need native + managed call stacks, use `perfcollect` rather than a raw `perf record` — it also captures LTTng CLR events and can resolve JIT-generated code, not just precompiled/R2R code. `perfcollect` itself is Linux-only (it relies on LTTng/perf), regardless of which platform you used to drive the load test.
- Tool and full instructions: <https://aka.ms/perfcollect>. Also see the repo's own [Linux Performance Tracing](../../../docs/project/linux-performance-tracing.md) doc.
- Basic flow (needs two terminals — one to drive `perfcollect`, one to run the server):
```bash
curl -OL https://aka.ms/perfcollect
chmod +x perfcollect
sudo ./perfcollect install # one-time, installs LTTng/perf prerequisites
# [App terminal] enable perf maps for JIT-generated code before starting the process:
export DOTNET_PerfMapEnabled=1
export DOTNET_EnableEventLog=1
<runtime-repo>/.dotnet/dotnet exec PlatformBenchmarks.dll --urls http://127.0.0.1:5000 &
# [Trace terminal] start collection, then drive load with wrk from a third terminal:
sudo ./perfcollect collect webapp_trace
# Ctrl+C to stop once wrk has run long enough (a few seconds under load is enough for a CPU investigation).
```
This produces `webapp_trace.trace.zip`, analyzable with PerfView (<https://aka.ms/perfview>) or `perf report`/`perf script` directly on the underlying `.perf.data`.
- Raising `kernel.perf_event_paranoid` may be required for an unprivileged `perf record`/`perfcollect` to work at all:
```bash
sudo sysctl -w kernel.perf_event_paranoid=-1 # temporary; restore the original value afterward
```
Always restore the original value once profiling is done — don't leave a relaxed `perf_event_paranoid` on a shared machine.
#### ⚠️ Symbol Resolution Warnings
- **JIT-generated (Tier0/Tier1) and precompiled R2R framework symbols require `DOTNET_PerfMapEnabled=1` to be set *before* the process starts.** This makes the runtime emit the perf map and JIT dump that `perfcollect` uses for symbol resolution; no separate `crossgen2` copy is required. If you forget this and only realize partway through a run, the trace from that run cannot be fixed after the fact — kill the process, re-export the env var, and restart it before collecting again.
- **Native runtime frames (`libcoreclr.so`, `libclrjit.so`, etc.) resolve out of the box, as long as you don't move or discard the build's `.dbg` files.** A Release native build ships each `.so` *stripped*, but the matching, un-stripped `libXyz.so.dbg` is written right alongside it in the same output directory (e.g. `artifacts/bin/coreclr/<rid>.<config>/`), linked via its `.gnu_debuglink` section and a matching Build ID. `perf`/`perfcollect` follow that link automatically as long as the `.dbg` file stays next to its `.so` — don't clean up or selectively copy only the `.so` files into a deployment layout, or you'll silently lose native symbolication.
- Even after doing all of the above, expect *some* residual unresolved frames from components you didn't build locally (e.g. the OS's own libraries, or a missing/mismatched `.dbg` file if you mixed binaries from different build configurations) — this doesn't mean the whole trace is useless; managed app-level, R2R, and JIT-emitted frames still resolve correctly via the generated symbol data.
- `DOTNET_PerfMapEnabled=1` has a real (if usually small) overhead of its own — don't leave it set for the throughput (`wrk`/`bombardier`) runs themselves, only for the dedicated profiling run.
### Common Pitfalls
- **Running via `artifacts/bin/testhost` directly fails** with "no framework found" — that layout only has `Microsoft.NETCore.App`, not `Microsoft.AspNetCore.App`. Use the SDK-overlay approach in Step 3 instead.
- **Forgetting to set the env var before starting the process** — most feature switches (like `DOTNET_USE_IO_URING`) are read once at startup, so changing it and re-`curl`-ing the same running process has no effect.
- **Comparing a single run per configuration** — throughput varies run to run; always do at least two runs per configuration.
- **Leaving the SDK shared framework overlaid** — always restore it (Step 6) and verify with a recursive content comparison; a stale/mismatched overlay silently breaks unrelated work on the same machine later.
## Benchmarking a Local Runtime Build on External Infrastructure with Crank
**Use external infrastructure for final validation after testing locally.** The .NET benchmarking infrastructure described here requires VPN access and is unavailable to external contributors. Runs take substantially longer than local checks, and the infrastructure is not always available; do not depend on it for the initial iteration loop.
Use [Crank](https://github.com/dotnet/crank) to deploy `PlatformBenchmarks` with selected locally-built runtime binaries to a remote application agent; a separate load agent drives HTTP requests without competing for the application's machine resources. These access restrictions apply to the shared infrastructure, not to Crank itself: external contributors can use Crank with their own agents.
References:
- [Crank documentation](https://github.com/dotnet/crank/blob/main/docs/README.md) and [getting started](https://github.com/dotnet/crank/blob/main/docs/getting_started.md): controller installation, agents, scenarios, and profiles.
- [Crank controller command-line arguments](https://github.com/dotnet/crank/blob/main/src/Microsoft.Crank.Controller/README.md): complete controller and per-job option reference.
- [Selecting .NET versions](https://github.com/dotnet/crank/blob/main/docs/dotnet_versions.md): framework, SDK, runtime, and ASP.NET version overrides.
- [PlatformBenchmarks configuration](https://github.com/aspnet/Benchmarks/blob/main/scenarios/platform.benchmarks.yml): `json` scenario, `application` and `load` jobs, and load variables.
- [Minimal APIs configuration](https://github.com/aspnet/Benchmarks/blob/main/src/BenchmarksApps/TechEmpower/Minimal/minimal.benchmarks.yml): `json`, `plaintext`, `fortunes`, and other scenarios for the same `Minimal` project used for local runs above. Unlike `PlatformBenchmarks`'s config, this one lives alongside the project rather than under the top-level `scenarios/` folder.
- [Linux performance tracing](../../../docs/project/linux-performance-tracing.md): native and managed symbol resolution.
### 1. Establish the Inputs
- Use an installed Crank controller and reachable application/load agents. Prefer separate machines for throughput measurements so load generation does not compete with the server for CPU.
- Select a PlatformBenchmarks configuration and an authorized machine profile. The public configuration includes profiles for particular infrastructure; choose or define one for the actual agents rather than assuming those machines are accessible. The getting-started guide shows how profiles assign `application` and `load` endpoints.
- Pin the benchmark source revision and exact SDK, runtime, and ASP.NET versions for comparisons. Do not allow `main`, `latest`, or `edge` to change dependencies between runs. Ensure the benchmark TFM and selected versions are compatible with the local runtime overlay.
Keep stock ASP.NET binaries unless the experiment explicitly includes a local ASP.NET build. A runtime-only comparison must not accidentally include a modified Kestrel transport.
### 2. Choose the Smallest Coherent Binary Subset
**Do not upload the entire testhost, SDK, or artifacts directory.** Start with the changed component, then add only dependencies required for its implementation, managed/native ABI, or runtime compatibility. This reduces upload time and makes the experiment attributable.
Inspect the source diff and relevant project outputs before choosing files:
| Change under test | Files to consider for a Linux overlay |
|---|---|
| Managed library only, such as `System.Net.Sockets` | The changed implementation `.dll` and matching `.pdb`; additional implementation dependencies only if needed. |
| Native PAL or interop contract | The affected native library, such as `libSystem.Native.so`, its matching `.so.dbg`, and any managed assembly changed with that ABI. |
| CoreLib, including ThreadPool | `System.Private.CoreLib.dll` and `.pdb`, with compatible VM/JIT binaries. When using a locally-built CoreLib against a different published runtime, use the matching `libcoreclr.so` and `libclrjit.so` plus their `.dbg` files rather than assuming internal contracts match. |
| JIT only | `libclrjit.so` and `.so.dbg` if compatible with the selected VM's JIT/EE interface; otherwise include the matching runtime components. |
| VM or GC | `libcoreclr.so` and `.so.dbg`, plus matching CoreLib/JIT when their internal contracts require it. |
Auf GitHub ansehen