Skip to content

Edge LLM Inference Under Real-World Constraints

How fast local inference can get, and how safe it stays at the edge. Independent research by Sahil Kadadekar.

  • 1.46M+ measurements
  • 55 reports
  • 6 whitepapers

Synthesis and technical reports

Key findings

Concrete results pulled from the published reports. Numbers, not narrative.

100% ASR

Q2_K collapses refusals on the worst-affected model — 100% attack success on qwen2.5-1.5b. Not uniform: effects vary by model.

p = 0.942

Alignment type does not predict batch-induced safety fragility (RLHF, SFT, DPO, distilled — none differ).

25pp

Backend migration moved safety 7–25pp, peaking at 23–25pp on Llama 3.2 1B. Chat template divergence, not the framework.

13.9×

Quality metrics are not safety proxies. Safety degraded 13.9× faster than quality on llama3.2-1b at Q3_K_S.

99.4%

Dual Ollama reached 99.4% coordination efficiency on the best config, and cut contention to near zero. Architectural fix, not code fix.

+74%

GPU memory bandwidth is the multi-agent bottleneck — not the serving stack. Overturned the TR130 conclusion.

2.25×

Continuous batching delivers 2.25× throughput at N=8 via 77-80% kernel reduction.

Q4_K_M

The safe GGUF default — established across 5 models, extended to 7 in v2. 30-67% cost savings.

NULL

FP8 KV-cache produces no Holm-significant safety effect across 24K paired records on 3 models. Not pre-approved, not pre-banned — workload-specific paired eval required.

κ = 0.69

Cross-LLM judge agreement is "triangulate" — single-judge labels are insufficient for safety classification. 68K judge rows over the TR145 safety subset. Plus: safety-specialist judges measure a different axis than general LLMs.

Conclusive reports and appendices

Dissertation-style synthesis documents consolidating findings across multiple technical reports.