Percentes run report: percentes-process-kill instrument commit: bce0628868ca6c4a231a5b398c169aaf586ac8a8 config sha256: 9ebb9245d33d3f89fe0a21bdb8cfd432a3718f31fcf1eac7b6fb2e64776750b5 CAVEAT: one vLLM replica on one GPU, killed and restarted in place; no two-replica or Kubernetes claim. N=5 runs on rented hardware; the acceptance criteria (§8) certified the instrument against the mock; injected-fault-versus-reality gaps are named in the report. == Conditional headline (appendix template) == Under process_kill fault injection (single-replica process kill, restart in place): 100.0% of in-flight requests failed and 0.0% timed out at 30 s (29 in flight on the only replica at fire, 29 of them indeterminate); of the 136 requests scheduled in the 42.2 s outage, 1 completed, 135 errored and 0 were censored; the replica served again 42.2 s after the kill (replica_ready); decomposed segments: container_start 0.74s, log_bringup 28.05s, engine_init 24.32s, weight_load 3.10s, torch_compile 0.82s, profile_kv_capture 5.73s, engine_ready 39.93s, server_ready 41.80s, replica_ready 42.19s, goodput_restored 40.96s; recovery to the pre-fault baseline: 41.0 s; goodput deficit 39.0 goodput-seconds vs pre-fault. One replica on one GPU; no two-replica or Kubernetes claim. == Run validity == valid: true client-validity gate: pass=true (skew p99=16us max=683us; undispatched=0; cpu worst 5s window=3.1%; gc pause p99 in [0.459, 0.524) ms, runtime bucket edges, gate on the upper edge) share gate: applicable=false pass=true shares=map[] injection timing: fire error +40.3ms (tolerance +-500ms); armed 17:00:58.531 == Environment pins (§6) == {"vllm":{"version":"0.29.0","image_digest":"vllm/vllm-openai@sha256:7ef5a35d1ef8ce2cf9d671dd91eec6e367c5849262e0362b4d3d4a26be0d87d2"},"model":{"name":"Qwen/Qwen2.5-7B-Instruct","revision":"a09a35458c702b33eeacc393d103063234e8bc28","quantization":"none"},"engine":{"kv_cache_gb":16,"max_num_seqs":256,"scheduler_settings":"max_num_seqs 256, max_num_batched_tokens 2048, chunked prefill on, policy fcfs, async scheduling on, cudagraph modes PIECEWISE and FULL (/server_info, 15 Sep 2026)","chunked_prefill":"on","cuda_graphs":"on","prefix_caching":"off","continuous_batching":"on"},"gpu":{"sku":"NVIDIA L40","driver":"570.195.03","cuda":"12.9 in the container (torch.version.cuda) on host driver 570.195.03 (CUDA 12.8 driver API)","cudnn":"9.20.0 (torch.backends.cudnn.version 92000)","nccl":"2.29.7","clock_power_policy":"persistence on, clocks at driver default, power limit 300 W of 300 W, not settable in guest"},"kubernetes":{"version":"none","cni":"none","dataplane_mode":"none","kube_proxy_mode":"none","node_monitor_grace_period_s":0},"readiness_probe":{"path":"none","period_s":0,"timeout_s":0,"failure_threshold":0},"storage":{"weights_medium":"root block volume vda (virtio, 100 GB), Hugging Face cache under /home/ubuntu/.cache/huggingface"},"container":{"runtime":"docker (server version recorded in run-N-fingerprint-*.txt)","restart_policy":"on-failure","name":"vllm","compile_cache":"container writable layer, /root/.cache/vllm/torch_compile_cache","hf_hub_offline":"unset"}} == Window "baseline" [60.0s, 330.0s) == scheduled=835 completed=835 errored=0 censored=0 error rate=0.0000 censored rate=0.0000 (first-class) TTFT conditional on completion: n=835 p50=123.9ms p95=260.2ms p99=289.0ms | p99.9=370.9ms max=400.1ms (descriptive-only at fault-window sample sizes, §7) p95-CI [238.9, 268.5]ms p99-CI [283.1, 355.0]ms e2e conditional on completion: n=835 p50=6643.7ms p95=7888.9ms p99=8355.8ms | p99.9=8527.9ms max=8577.0ms (descriptive-only at fault-window sample sizes, §7) p95-CI [7626.6, 7967.0]ms p99-CI [8232.9, 8459.3]ms ITL pooled (per-window), inter-chunk (§3; §10 check: 31 of 835 matched): n=202373 p50=20.4ms p95=81.4ms p99=102.7ms | p99.9=181.5ms max=3549.2ms (descriptive-only at fault-window sample sizes, §7) completion length: content events n=835 mean=243.4 p50=248 p95=254 max=256; completion tokens from usage n=835 mean=256.0 p50=256 p95=256 max=256 receive path (§2, not run-failing): client TTFT mean 146.2ms over 835; no server histogram named; loopback canary (626 streams completed, 32 tokens at 20ms TTFT, 10ms ITL): event lag p99 1.80ms max 4.33ms; TTFT deviation p50 1.25ms max 2.47ms, ITL deviation p50 0.16ms p99 0.85ms max 2.64ms throughput=3.09 rps goodput=3.09 rps goodput-frac=1.0000 goodput-versus-threshold sweep (§4): TTFT<= 800ms & e2e<=12000ms: goodput 1.0000 TTFT<= 800ms & e2e<=14000ms: goodput 1.0000 TTFT<= 800ms & e2e<=18000ms: goodput 1.0000 TTFT<=1000ms & e2e<=12000ms: goodput 1.0000 TTFT<=1000ms & e2e<=14000ms: goodput 1.0000 TTFT<=1000ms & e2e<=18000ms: goodput 1.0000 TTFT<=1500ms & e2e<=12000ms: goodput 1.0000 TTFT<=1500ms & e2e<=14000ms: goodput 1.0000 TTFT<=1500ms & e2e<=18000ms: goodput 1.0000 Completion incidence, Aalen-Johansen (n=835 over ALL scheduled; errors compete, timeouts censored; 0 outstanding at the horizon; horizon 30s): p50 completion at 6.644s p90 completion at 7.400s p95 completion at 7.889s p99 completion at 8.355s t= 5.360s incidence=0.0012 (at risk 835) t= 5.894s incidence=0.0850 (at risk 765) t= 6.066s incidence=0.1689 (at risk 695) t= 6.213s incidence=0.2527 (at risk 625) t= 6.364s incidence=0.3365 (at risk 555) t= 6.511s incidence=0.4204 (at risk 485) t= 6.650s incidence=0.5042 (at risk 415) t= 6.765s incidence=0.5880 (at risk 345) t= 6.910s incidence=0.6719 (at risk 275) t= 7.041s incidence=0.7557 (at risk 205) t= 7.194s incidence=0.8395 (at risk 135) t= 7.555s incidence=0.9234 (at risk 65) t= 8.577s incidence=1.0000 (at risk 1) == Window "guard" [330.0s, 360.0s) == PRE-FAULT GUARD WINDOW (§3): the last pinned client timeout before the fire anchor; the fault can terminate requests intended here, so this is not pre-fault degradation and it feeds no baseline-derived quantity. scheduled=97 completed=68 errored=29 censored=0 error rate=0.2990 censored rate=0.0000 (first-class) error classes: map[malformed_stream:29] TTFT conditional on completion: n=68 p50=120.2ms p95=199.9ms p99=225.2ms | p99.9=257.4ms max=257.4ms (descriptive-only at fault-window sample sizes, §7) p95-CI omitted (sample budget insufficient, §7) p99-CI omitted (sample budget insufficient, §7) e2e conditional on completion: n=68 p50=6471.7ms p95=7626.8ms p99=7757.8ms | p99.9=7757.8ms max=7757.8ms (descriptive-only at fault-window sample sizes, §7) p95-CI omitted (sample budget insufficient, §7) p99-CI omitted (sample budget insufficient, §7) ITL pooled (per-window), inter-chunk (§3; §10 check: 8 of 68 matched): n=16609 p50=20.5ms p95=81.6ms p99=103.1ms | p99.9=178.2ms max=287.0ms (descriptive-only at fault-window sample sizes, §7) completion length: content events n=68 mean=245.2 p50=248 p95=256 max=256; completion tokens from usage n=68 mean=256.0 p50=256 p95=256 max=256 receive path (§2, not run-failing): client TTFT mean 135.5ms over 68; no server histogram named; loopback canary (69 streams completed, 32 tokens at 20ms TTFT, 10ms ITL): event lag p99 1.83ms max 2.56ms; TTFT deviation p50 1.15ms max 2.56ms, ITL deviation p50 0.21ms p99 0.80ms max 1.06ms throughput=2.27 rps goodput=2.27 rps goodput-frac=0.7010 goodput-versus-threshold sweep (§4): TTFT<= 800ms & e2e<=12000ms: goodput 0.7010 TTFT<= 800ms & e2e<=14000ms: goodput 0.7010 TTFT<= 800ms & e2e<=18000ms: goodput 0.7010 TTFT<=1000ms & e2e<=12000ms: goodput 0.7010 TTFT<=1000ms & e2e<=14000ms: goodput 0.7010 TTFT<=1000ms & e2e<=18000ms: goodput 0.7010 TTFT<=1500ms & e2e<=12000ms: goodput 0.7010 TTFT<=1500ms & e2e<=14000ms: goodput 0.7010 TTFT<=1500ms & e2e<=18000ms: goodput 0.7010 CONDITIONAL CAVEAT: error+censored fraction exceeds 5%; read the completed-only percentiles against the completion-incidence curve below (§3). Completion incidence, Aalen-Johansen (n=97 over ALL scheduled; errors compete, timeouts censored; 0 outstanding at the horizon; horizon 30s): p50 completion at 6.945s p90 unattainable (final completion incidence 0.70; ceiling 0.70) p95 unattainable (final completion incidence 0.70; ceiling 0.70) p99 unattainable (final completion incidence 0.70; ceiling 0.70) t= 5.520s incidence=0.0103 (at risk 76) t= 5.770s incidence=0.0722 (at risk 69) t= 6.011s incidence=0.1340 (at risk 63) t= 6.251s incidence=0.1959 (at risk 57) t= 6.386s incidence=0.2577 (at risk 51) t= 6.447s incidence=0.3196 (at risk 45) t= 6.565s incidence=0.3814 (at risk 39) t= 6.807s incidence=0.4433 (at risk 32) t= 6.945s incidence=0.5052 (at risk 24) t= 7.369s incidence=0.5670 (at risk 14) t= 7.539s incidence=0.6289 (at risk 8) t= 7.754s incidence=0.6907 (at risk 2) t= 7.757s incidence=0.7010 (at risk 1) == Window "fault_degraded" [360.0s, 401.0s) == scheduled=134 completed=0 errored=134 censored=0 error rate=1.0000 censored rate=0.0000 (first-class) error classes: map[connect:134] TTFT conditional on completion: no completed samples e2e conditional on completion: no completed samples ITL pooled (per-window), inter-chunk (§3, no usage in the stream): no completed samples completion length: content events none; completion tokens from usage not verifiable (no completed request carried a usage object) receive path (§2, not run-failing): client TTFT mean 0.0ms over 0; no server histogram named; loopback canary (95 streams completed, 32 tokens at 20ms TTFT, 10ms ITL): event lag p99 1.83ms max 2.06ms; TTFT deviation p50 1.15ms max 1.72ms, ITL deviation p50 0.29ms p99 0.64ms max 1.07ms throughput=0.00 rps goodput=0.00 rps goodput-frac=0.0000 goodput-versus-threshold sweep (§4): TTFT<= 800ms & e2e<=12000ms: goodput 0.0000 TTFT<= 800ms & e2e<=14000ms: goodput 0.0000 TTFT<= 800ms & e2e<=18000ms: goodput 0.0000 TTFT<=1000ms & e2e<=12000ms: goodput 0.0000 TTFT<=1000ms & e2e<=14000ms: goodput 0.0000 TTFT<=1000ms & e2e<=18000ms: goodput 0.0000 TTFT<=1500ms & e2e<=12000ms: goodput 0.0000 TTFT<=1500ms & e2e<=14000ms: goodput 0.0000 TTFT<=1500ms & e2e<=18000ms: goodput 0.0000 CONDITIONAL CAVEAT: error+censored fraction exceeds 5%; read the completed-only percentiles against the completion-incidence curve below (§3). Completion incidence, Aalen-Johansen (n=134 over ALL scheduled; errors compete, timeouts censored; 0 outstanding at the horizon; horizon 30s): p50 unattainable (final completion incidence 0.00; ceiling 0.00) p90 unattainable (final completion incidence 0.00; ceiling 0.00) p95 unattainable (final completion incidence 0.00; ceiling 0.00) p99 unattainable (final completion incidence 0.00; ceiling 0.00) == Window "fault" [360.0s, 960.0s) == scheduled=1906 completed=1771 errored=135 censored=0 error rate=0.0708 censored rate=0.0000 (first-class) error classes: map[connect:135] TTFT conditional on completion: n=1771 p50=126.0ms p95=246.0ms p99=332.3ms | p99.9=429.1ms max=557.6ms (descriptive-only at fault-window sample sizes, §7) p95-CI [239.1, 262.0]ms p99-CI [295.1, 368.7]ms e2e conditional on completion: n=1771 p50=6660.1ms p95=7774.2ms p99=8269.8ms | p99.9=8871.9ms max=8871.9ms (descriptive-only at fault-window sample sizes, §7) p95-CI [7698.9, 7877.3]ms p99-CI [8181.4, 8558.7]ms ITL pooled (per-window), inter-chunk (§3; §10 check: 112 of 1771 matched): n=429573 p50=20.4ms p95=83.6ms p99=103.9ms | p99.9=182.7ms max=3524.6ms (descriptive-only at fault-window sample sizes, §7) completion length: content events n=1771 mean=243.6 p50=248 p95=256 max=256; completion tokens from usage n=1771 mean=256.0 p50=256 p95=256 max=256 receive path (§2, not run-failing): client TTFT mean 147.4ms over 1771; no server histogram named; loopback canary (1390 streams completed, 32 tokens at 20ms TTFT, 10ms ITL): event lag p99 1.81ms max 3.11ms; TTFT deviation p50 1.25ms max 2.09ms, ITL deviation p50 0.21ms p99 0.84ms max 1.87ms throughput=2.95 rps goodput=2.95 rps goodput-frac=0.9292 goodput-versus-threshold sweep (§4): TTFT<= 800ms & e2e<=12000ms: goodput 0.9292 TTFT<= 800ms & e2e<=14000ms: goodput 0.9292 TTFT<= 800ms & e2e<=18000ms: goodput 0.9292 TTFT<=1000ms & e2e<=12000ms: goodput 0.9292 TTFT<=1000ms & e2e<=14000ms: goodput 0.9292 TTFT<=1000ms & e2e<=18000ms: goodput 0.9292 TTFT<=1500ms & e2e<=12000ms: goodput 0.9292 TTFT<=1500ms & e2e<=14000ms: goodput 0.9292 TTFT<=1500ms & e2e<=18000ms: goodput 0.9292 CONDITIONAL CAVEAT: error+censored fraction exceeds 5%; read the completed-only percentiles against the completion-incidence curve below (§3). Completion incidence, Aalen-Johansen (n=1906 over ALL scheduled; errors compete, timeouts censored; 0 outstanding at the horizon; horizon 30s): p50 completion at 6.711s p90 completion at 7.921s p95 unattainable (final completion incidence 0.93; ceiling 0.93) p99 unattainable (final completion incidence 0.93; ceiling 0.93) t= 5.421s incidence=0.0005 (at risk 1771) t= 6.041s incidence=0.0782 (at risk 1623) t= 6.217s incidence=0.1558 (at risk 1475) t= 6.356s incidence=0.2340 (at risk 1326) t= 6.467s incidence=0.3116 (at risk 1178) t= 6.570s incidence=0.3893 (at risk 1030) t= 6.659s incidence=0.4669 (at risk 882) t= 6.794s incidence=0.5446 (at risk 734) t= 6.942s incidence=0.6222 (at risk 586) t= 7.100s incidence=0.6999 (at risk 438) t= 7.294s incidence=0.7775 (at risk 290) t= 7.606s incidence=0.8552 (at risk 142) t= 8.872s incidence=0.9292 (at risk 1) == Window "outage" [360.0s, 402.2s) == scheduled=136 completed=1 errored=135 censored=0 error rate=0.9926 censored rate=0.0000 (first-class) error classes: map[connect:135] TTFT conditional on completion: n=1 p50=113.3ms p95=113.3ms p99=113.3ms | p99.9=113.3ms max=113.3ms (descriptive-only at fault-window sample sizes, §7) p95-CI omitted (sample budget insufficient, §7) p99-CI omitted (sample budget insufficient, §7) e2e conditional on completion: n=1 p50=5640.2ms p95=5640.2ms p99=5640.2ms | p99.9=5640.2ms max=5640.2ms (descriptive-only at fault-window sample sizes, §7) p95-CI omitted (sample budget insufficient, §7) p99-CI omitted (sample budget insufficient, §7) ITL pooled (per-window), inter-chunk (§3; §10 check: 0 of 1 matched): n=253 p50=18.6ms p95=67.3ms p99=71.7ms | p99.9=77.8ms max=77.8ms (descriptive-only at fault-window sample sizes, §7) completion length: content events n=1 mean=254.0 p50=254 p95=254 max=254; completion tokens from usage n=1 mean=256.0 p50=256 p95=256 max=256 receive path (§2, not run-failing): client TTFT mean 113.2ms over 1; no server histogram named; loopback canary (98 streams completed, 32 tokens at 20ms TTFT, 10ms ITL): event lag p99 1.83ms max 2.06ms; TTFT deviation p50 1.15ms max 1.72ms, ITL deviation p50 0.29ms p99 0.65ms max 1.07ms throughput=0.02 rps goodput=0.02 rps goodput-frac=0.0074 goodput-versus-threshold sweep (§4): TTFT<= 800ms & e2e<=12000ms: goodput 0.0074 TTFT<= 800ms & e2e<=14000ms: goodput 0.0074 TTFT<= 800ms & e2e<=18000ms: goodput 0.0074 TTFT<=1000ms & e2e<=12000ms: goodput 0.0074 TTFT<=1000ms & e2e<=14000ms: goodput 0.0074 TTFT<=1000ms & e2e<=18000ms: goodput 0.0074 TTFT<=1500ms & e2e<=12000ms: goodput 0.0074 TTFT<=1500ms & e2e<=14000ms: goodput 0.0074 TTFT<=1500ms & e2e<=18000ms: goodput 0.0074 CONDITIONAL CAVEAT: error+censored fraction exceeds 5%; read the completed-only percentiles against the completion-incidence curve below (§3). Completion incidence, Aalen-Johansen (n=136 over ALL scheduled; errors compete, timeouts censored; 0 outstanding at the horizon; horizon 30s): p50 unattainable (final completion incidence 0.01; ceiling 0.01) p90 unattainable (final completion incidence 0.01; ceiling 0.01) p95 unattainable (final completion incidence 0.01; ceiling 0.01) p99 unattainable (final completion incidence 0.01; ceiling 0.01) t= 5.637s incidence=0.0074 (at risk 1) == Window "fault_recovered" [401.0s, 960.0s) == scheduled=1772 completed=1771 errored=1 censored=0 error rate=0.0006 censored rate=0.0000 (first-class) error classes: map[connect:1] TTFT conditional on completion: n=1771 p50=126.0ms p95=246.0ms p99=332.3ms | p99.9=429.1ms max=557.6ms (descriptive-only at fault-window sample sizes, §7) p95-CI [239.1, 262.0]ms p99-CI [295.1, 368.7]ms e2e conditional on completion: n=1771 p50=6660.1ms p95=7774.2ms p99=8269.8ms | p99.9=8871.9ms max=8871.9ms (descriptive-only at fault-window sample sizes, §7) p95-CI [7698.9, 7877.3]ms p99-CI [8181.4, 8558.7]ms ITL pooled (per-window), inter-chunk (§3; §10 check: 112 of 1771 matched): n=429573 p50=20.4ms p95=83.6ms p99=103.9ms | p99.9=182.7ms max=3524.6ms (descriptive-only at fault-window sample sizes, §7) completion length: content events n=1771 mean=243.6 p50=248 p95=256 max=256; completion tokens from usage n=1771 mean=256.0 p50=256 p95=256 max=256 receive path (§2, not run-failing): client TTFT mean 147.4ms over 1771; no server histogram named; loopback canary (1295 streams completed, 32 tokens at 20ms TTFT, 10ms ITL): event lag p99 1.81ms max 3.11ms; TTFT deviation p50 1.27ms max 2.09ms, ITL deviation p50 0.19ms p99 0.85ms max 1.87ms throughput=3.17 rps goodput=3.17 rps goodput-frac=0.9994 goodput-versus-threshold sweep (§4): TTFT<= 800ms & e2e<=12000ms: goodput 0.9994 TTFT<= 800ms & e2e<=14000ms: goodput 0.9994 TTFT<= 800ms & e2e<=18000ms: goodput 0.9994 TTFT<=1000ms & e2e<=12000ms: goodput 0.9994 TTFT<=1000ms & e2e<=14000ms: goodput 0.9994 TTFT<=1000ms & e2e<=18000ms: goodput 0.9994 TTFT<=1500ms & e2e<=12000ms: goodput 0.9994 TTFT<=1500ms & e2e<=14000ms: goodput 0.9994 TTFT<=1500ms & e2e<=18000ms: goodput 0.9994 Completion incidence, Aalen-Johansen (n=1772 over ALL scheduled; errors compete, timeouts censored; 0 outstanding at the horizon; horizon 30s): p50 completion at 6.658s p90 completion at 7.517s p95 completion at 7.775s p99 completion at 8.303s t= 5.421s incidence=0.0006 (at risk 1771) t= 6.041s incidence=0.0841 (at risk 1623) t= 6.217s incidence=0.1676 (at risk 1475) t= 6.356s incidence=0.2517 (at risk 1326) t= 6.467s incidence=0.3352 (at risk 1178) t= 6.570s incidence=0.4187 (at risk 1030) t= 6.659s incidence=0.5023 (at risk 882) t= 6.794s incidence=0.5858 (at risk 734) t= 6.942s incidence=0.6693 (at risk 586) t= 7.100s incidence=0.7528 (at risk 438) t= 7.294s incidence=0.8363 (at risk 290) t= 7.606s incidence=0.9199 (at risk 142) t= 8.872s incidence=0.9994 (at risk 1) == In-flight loss accounting at fire == total=29 completed=0 errored=29 censored=0 by-replica=map[:29] errored by class: malformed_stream=29 indeterminate (in flight, terminal time within 20.4ms after the fire: the fire uncertainty 12.3ms plus the delivery allowance): 29 determinate: total=0 completed=0 errored=0 censored=0 == Modal during-fault latency vs thresholds (§4) == baseline SD: TTFT 49.10ms, e2e 601.40ms; modal during-fault: TTFT 119.0ms, e2e 6624.9ms TTFT threshold distances (baseline SDs, signed modal-threshold): map[1000ms:-17.941442118363984 1500ms:-28.124004254817972 800ms:-13.868417263782385] e2e threshold distances (baseline SDs, signed modal-threshold): map[12000ms:-8.937584974917472 14000ms:-12.263136145962616 18000ms:-18.914238488052902] == Recovery (two baselines, hysteresis; §5) == pre-fault baseline goodput: 1.0000 single-replica equilibrium baseline: NOT ESTIMABLE (no service during the degraded plateau; single-replica equilibrium undefined for this run) TTR to pre-fault baseline: TTR 41.0s (baseline 1.0000, canceled entries 0, re-degradations 0) TTR to equilibrium baseline: not applicable under process kill, which has no survivor (§5) integrated goodput deficit: 39.00 (vs pre-fault), n/a (vs equilibrium) goodput-seconds component e2e_slo: TTR 41.0s (baseline 1.0000, canceled entries 0, re-degradations 0) component error_rate: TTR 41.0s (baseline 1.0000, canceled entries 0, re-degradations 0) component ttft_slo: TTR 41.0s (baseline 1.0000, canceled entries 0, re-degradations 0) backlog drain: measured=false (no failed-request backlog exists in the Phase 0 mock (no queue); reported N/A per §5) == Sensitivity table (X x R x H; §5) == entry R H | TTR->prefault TTR->equilibrium 85 5 15 | 41.0s n/a 85 5 30 | 41.0s n/a 85 5 60 | 41.0s n/a 85 10 15 | 41.0s n/a 85 10 30 | 41.0s n/a 85 10 60 | 41.0s n/a 85 20 15 | 39.0s n/a 85 20 30 | 39.0s n/a 85 20 60 | 39.0s n/a 90 5 15 | 41.0s n/a 90 5 30 | 41.0s n/a 90 5 60 | 41.0s n/a 90 10 15 | 41.0s n/a 90 10 30 | 41.0s n/a 90 10 60 | 41.0s n/a 90 20 15 | 40.0s n/a 90 20 30 | 40.0s n/a 90 20 60 | 40.0s n/a 95 5 15 | 42.0s n/a 95 5 30 | 42.0s n/a 95 5 60 | 42.0s n/a 95 10 15 | 41.0s n/a 95 10 30 | 41.0s n/a 95 10 60 | 41.0s n/a 95 20 15 | 41.0s n/a 95 20 30 | 41.0s n/a 95 20 60 | 41.0s n/a == Recovery decomposition (§5: only measured boundaries are claimed) == reschedule [api] N/A: no scheduler: the container runtime restarts in place container_start [api] measured: 0.74s log_bringup [log] measured: 28.05s engine_init [log] measured: 24.32s weight_download [log] N/A: no download line after the fire: weights served from the mounted cache weight_load [log] measured: 3.10s torch_compile [log] measured: 0.82s profile_kv_capture [log] measured: 5.73s engine_ready [log] measured: 39.93s server_ready [log] measured: 41.80s replica_ready [probe] measured: 42.19s traffic_restored [probe] N/A: one replica addressed directly: no Service routing_propagation [probe] N/A: one replica addressed directly: no Service goodput_restored [client] measured: 40.96s figures printed in the server log: graph_capture_gib=0.52 graph_capture_s=5 init_engine_compilation_s=0.18 init_engine_s=7.98 model_loaded_gib=14.29 model_loaded_s=4.049787 torch_compile_s=0.18 weights_loaded_s=2.57 == Run-validity gates (§10 G1-G7) == G1 per-replica share 45-55% pre-fault n/a single-replica target: share gate not applicable G2 client-validity gate clean pass skew p99=16us/max=683us; undispatched=0; cpu_measured=true worst=3.1%; gc pause p99 in [0.459, 0.524) ms, runtime bucket edges, gate on the upper edge G3 zero errored outcomes among victim-attributed in-flight requests n/a not the black-hole variant: client-silence assertion not applicable G4 endpoint-staleness window >= 20s with victim-bound traffic observed n/a not the black-hole variant: staleness assertion not applicable G5 GPU clock/power fingerprints equal across replicas and runs n/a fingerprints are not fed into the gate; percentes-campaign writes them per run under process kill; reported not applicable G6 baseline goodput >= 0.99 pass baseline goodput 1.0000 (pinned minimum 0.99, §10 G6; below it the load calibration is wrong and is redone) G7 baseline queue stability: per-replica waiting-queue mean <= 1.0 pass vllm:num_requests_waiting baseline means r0=0.000/n270, sampled every 1 s, 270 expected (pinned maximum 1.0, §10 G7; 41 scrape errors) all pass: true; node-loss-representative: false CAVEAT: one vLLM replica on one GPU, killed and restarted in place; no two-replica or Kubernetes claim. N=5 runs on rented hardware; the acceptance criteria (§8) certified the instrument against the mock; injected-fault-versus-reality gaps are named in the report.