At 5% packet loss, 16 of 30 llama.cpp RPC requests fell below half their baseline

Experiments
Distributed systems
Across 30 repeated requests, 5% one-way loss made 16 run below half their local baseline. Every request still exited successfully, and the logs named no degraded path.
Author

Mikhail

Published

6 October 2026

NoteWhat this note is

This is a local campaign on llama.cpp RPC: 90 requests, three network conditions, and a rule for “slow” written down before the first run. The result describes this setup, not every distributed inference system.

The six-run puzzle

An earlier loss test left an awkward result. At 5% packet loss, six repeats generated anywhere from 8.0 to 32.9 tokens per second. Some looked healthy. Others did not.

Six repeats showed variability, but they were too few to estimate how often slow requests occurred. Was the machine drifting? Did loss lower every request a little? Or did it create a slow tail that a small sample happened to catch?

The distinction matters. A uniform slowdown can be summarized by one average. A tail asks a different operational question: how often does a request fall into the bad part of the distribution while the system still calls it a success?

So the next campaign kept the software, model, prompt, and topology fixed. It changed only packet loss and ran 30 requests under each condition to estimate the distribution.

What ran

The system was llama.cpp RPC v0.4.1, build b1-b29c606 (llama.cpp contributors 2026). One llama-cli process used two ggml-rpc-server workers. All three ran in Docker on one CPU host with Qwen2.5-0.5B-Instruct Q4_K_M.

The client asked for up to 32 new tokens after the prompt The capital of France is. There was no concurrent load. Each request had a 180 s timeout.

Only packets leaving rpc-b passed through tc netem (Linux man-pages project 2024). Incoming packets and both directions of rpc-a stayed untouched.

flowchart LR
  C[llama-cli]
  A[rpc-a]
  B[rpc-b]
  Q[egress netem]
  C -->|RPC| A
  C -->|RPC| B
  A -->|reply| C
  B -->|reply| Q
  Q -->|0%, 1%, or 5% loss| C

The campaign used ten blocks. Every block contained the same three conditions: no loss, 1% loss, and 5% loss. Their order was shuffled in advance. Each condition ran three times per block, giving 30 requests at every point.

Baseline medians ranged from 37.1 to 48.8 tokens per second across the ten blocks. Normalizing each lossy request against the baseline from its own block reduces confounding from that drift.

The rule was fixed first

Let g_i be the generation rate of request i. Let \tilde g_{b,0} be the median baseline rate in the same block. The normalized rate was

r_i^g = \frac{g_i}{\tilde g_{b,0}}.

A request was slow when r_i^g < 0.5. In ordinary words: it generated at less than half the local baseline. A timeout or non-zero exit would also count as slow.

That threshold was in the specification before the campaign started. Choosing it after seeing the histogram would bias the reported slow-request count.

Five percent widened the lower tail

The full summary is below. p10 is the 10th percentile: about 10% of measured rates were at or below it. Latency includes model loading and Docker overhead, so generation rate remains the primary metric.

Egress loss Requests Median, tok/s p10, tok/s Range, tok/s Median latency, s Slow Failed
0% 30 46.80 39.01 35.20–49.10 1.996 0/30 0/30
1% 30 45.05 33.03 20.50–49.10 2.261 1/30 0/30
5% 30 20.30 12.05 3.20–49.00 3.761 16/30 0/30

At 1%, the median barely moved. One request crossed the pre-set slow threshold, and the lower end widened.

At 5%, 16 of 30 requests ran below half their block baseline. None did so without loss. The median fell from 46.8 to 20.3 tokens per second, while median latency rose from 1.996 to 3.761 s.

The range is the important nuisance. One 5% request still reached 49.0 tokens per second. Another managed only 3.2. Packet loss did not apply a neat, predictable tax to every request; it made the outcome much less dependable.

The latency tail moved too. Its 90th percentile rose from 2.227 s at baseline to 5.111 s at 5% loss. With 30 repeats per condition, these percentiles remain estimates rather than stable tail bounds.

More drops, lower rank

The harness saved tc -s qdisc immediately after every lossy request and before removing the queue. This gave each repeat a kernel counter for packets that netem decided to drop.

First, the counter needed a calibration. In a separate ICMP run, ping observed 105 losses under the 1% setting and 490 under 5%. The qdisc counter reported the same two numbers. For this calibration traffic, the qdisc drop count matched the loss observed by ping exactly.

Inside the 5% condition, the Spearman rank correlation between the measured qdisc drop fraction and normalized generation rate was

\rho = -0.564.

Requests with higher measured drop fractions tended to rank lower in generation speed. The preregistered criterion was \rho \le -0.40; the observed \rho = -0.564 met it.

It does not establish the mechanism. Offload and segmentation can change what the qdisc denominator means for TCP, and the campaign did not record retransmissions. Higher measured qdisc drop fractions were associated with lower normalized generation rates. “TCP retransmission caused the whole tail” remains a hypothesis.

The system reported success

All 90 requests exited with code 0. None reached the timeout. The answer arrived.

For the 16 slow requests at 5% loss, the client and worker logs contained no explicit line about loss, retry, timeout, an unhealthy path, or an RPC error. If monitoring looked only at completion and exit status, every run was green.

The harness had a different view. It knew that 5% loss was installed, read the qdisc counter, and measured a request below half its local baseline. That difference is the useful signal here.

Exit status marked every request successful. The token-rate threshold marked 16 as slow.

What changed in our understanding

The initial six repeats suggested that small packet loss might produce unstable performance. The larger campaign turned that suspicion into a bounded statement.

In this setup, under 5% one-way egress loss, 16 of 30 requests fell below half their block baseline, compared with 0 of 30 without loss. The effect was associated with the measured qdisc drop fraction. It remained invisible in the application’s explicit signals.

This is more specific than saying that “networks matter.” It gives an operator a failure mode to look for: requests can complete correctly even when generation falls below half the local baseline.

What this does not show

The experiment used one llama.cpp build, one 0.5B model, and CPU execution. Loss affected only one worker’s egress. A symmetric path may behave differently.

The request was short and alone. Its reported generation speed is not the throughput of a long generation, and it says nothing about continuous batching or a queue of users.

No packet trace or TCP retransmission counter was collected. The qdisc association narrows the next question; it does not answer it.

Nothing here establishes transfer to vLLM, GPUs, a large KV cache, or another machine. Transfer to those settings requires separate measurements.

How to repeat the comparison

Use the same llama.cpp image, model, prompt, and two-worker topology. For each of ten blocks:

  1. Choose a pre-shuffled order for baseline, 1%, and 5% egress loss.
  2. Run three requests at each point.
  3. Save client output, exit status, worker logs, wall time, and tc -s qdisc before clearing the queue.
  4. Normalize every successful generation rate by the median baseline rate from its block.
  5. Keep every valid repeat, including the inconvenient ones.

Then report the distribution, not only its average. The 49.0 and the 3.2 tokens per second came from the same configured loss point. Hiding either would hide the result.

Questions left by the tail

  1. Does the slow tail remain when every request is forced to generate the same number of tokens, rather than being allowed to stop early?
  2. Do the slow repeats coincide with TCP retransmission counts and timing, or is the qdisc association standing in for another mechanism?
  3. If the mechanism survives that test, does a second distributed inference runtime show the same pattern of successful requests running below half-baseline speed?
  4. Under concurrent load, does the lower tail grow faster than the configured network loss?
Back to top

References

Linux man-pages project. 2024. Tc-Netem(8) — Network Emulator. Https://man7.org/linux/man-pages/man8/tc-netem.8.html.
llama.cpp contributors. 2026. Llama.cpp. Https://github.com/ggml-org/llama.cpp.