Q2_K collapses refusals on the worst-affected model — 100% attack success on qwen2.5-1.5b. Not uniform: effects vary by model.
Edge LLM Inference Under Real-World Constraints
How fast local inference can get, and how safe it stays at the edge. Independent research by Sahil Kadadekar.
- 1.46M+ measurements
- 55 technical reports
- 6 synthesis whitepapers
- 47 TR numbers · 3 baselines · versions counted
Synthesis and technical reports
Key findings
Concrete results pulled from the published reports. Numbers, not narrative.
Alignment type does not predict batch-induced safety fragility (RLHF, SFT, DPO, distilled — none differ).
Backend migration moved safety 7–25pp, peaking at 23–25pp on Llama 3.2 1B. Chat template divergence, not the framework.
Quality metrics are not safety proxies. Safety degraded 13.9× faster than quality on llama3.2-1b at Q3_K_S.
Dual Ollama reached 99.4% coordination efficiency on the best config, and cut contention to near zero. Architectural fix, not code fix.
GPU memory bandwidth is the multi-agent bottleneck — not the serving stack. Overturned the TR130 conclusion.
Continuous batching delivers 2.25× throughput at N=8 via 77-80% kernel reduction.
The safe GGUF default — established across 5 models, extended to 7 in v2. 30-67% cost savings.
FP8 KV-cache produces no Holm-significant safety effect across 24K paired records on 3 models. Not pre-approved, not pre-banned — workload-specific paired eval required.
Cross-LLM judge agreement is "triangulate" — single-judge labels are insufficient for safety classification. 68K judge rows over the TR145 safety subset. Plus: safety-specialist judges measure a different axis than general LLMs.
Conclusive reports and appendices
Dissertation-style synthesis documents consolidating findings across multiple technical reports.
- Conclusive Report: Phase 1 — Foundation (TR108–TR116)
Dissertation-style synthesis — language, architecture, runtime, and model selection for multi-agent LLM systems.
- Phase 1 Extended Appendices
Supplemental material extracted from the Phase 1 conclusive report.
- Conclusive Report: Phase 2 — Benchmarking (TR117–TR122)
Dissertation-style synthesis — performance, cost, scaling, compiler behavior, and physical limits of consumer-GPU inference.
- Phase 2 Extended Appendices
Supplemental material extracted from the Phase 2 conclusive report.
- Conclusive Report: Phase 3 — Optimization (TR123–TR133)
Dissertation-style synthesis — economics, quantization, context scaling, serving stacks, and predictive modeling.
- Phase 3 Extended Appendices
Supplemental material extracted from the Phase 3 conclusive report.
- Conclusive Report: Phase 4 — Safety Pivot (TR134–TR137)
Dissertation-style synthesis — quantization-induced alignment erosion, concurrency invariance, backend-driven template divergence, and cross-axis safety taxonomy.
- Phase 4 Extended Appendices
Supplemental material for the safety-critical deployment synthesis.
- Conclusive Report: Phase 5 — Attack Surface (TR138–TR143)
Safety attack-surface synthesis — batch perturbation, multi-turn jailbreaks, long-context exploitation, cross-architecture fragility, quality-safety divergence, and cross-request composition across 306,996 evaluated samples and 18+ models. TR138 Study D batch-invariant-kernel ablation is published as a standalone addendum.
- Phase 5 Extended Appendices
Supplemental material for the safety attack-surface synthesis.
- Conclusive Report: Phase 6 — Serving-State Safety Certification (TR144–TR149+TR152)
Measurement-validity substrate (judge triangulation, KV-cache safety null, speculative-decoding safety-invariance screen, mechanistic probing, portability validation) plus the FP8 KV-cache standardized batteries and serving-state factorial. The inference-flag safety null line for optimized LLM serving.
- Phase 6 Extended Appendices
Per-report data tables, named-method definitions, and cross-TR ledgers for serving-state safety certification.