Skip to content

Papers

Independent research on inference optimization, constitutional AI architectures, and empirical safety evaluation.

  • 2 accepted
  • 2 public preprints
  • 1 under peer review
  • 16 papers total
  • 1.46M+ measurements

Author: Sahil Kadadekar · Independent research

Accepted & public

The ICML 2026 workshop paper was accepted 2026-05-22 and presented at the workshop — the first peer-reviewed paper from the program. The first paper on Chimera’s own thesis, from Banterpacks, was accepted at the NeurIPS 2026 Workshop on Foundation and Large Model Security. The speculative-decoding screen and the quantization safety-proxy study are public arXiv preprints.

Under peer review

Double-blind, so titles stay off this page until decisions land.

Submitted

1 submission, with its PDF, artifact manifest, and venue checklist complete.

In preparation

Synthesis papers and methodology work derived from the published technical report archive, plus papers out of review and being revised for resubmission.

  1. 01

    Inference Optimization Is Not Safety-Neutral

    Synthesis paper. Across two shared anchor models, quantization, backend, and concurrency account for 57%, 41%, and 2% of normalized safety-score changes: descriptive shares, not a causal decomposition. Chat template divergence can induce larger safety shifts than numerical precision.

    SynthesisTarget: TBD
  2. 02

    Empirical Capacity Planning for Local LLM Inference

    Capacity planning as a fitted systems problem. Backend choice, context length, and memory pressure all materially change the feasible operating regime. Planner quality should be judged by validation against explicit targets, not analytic elegance.

    SynthesisTarget: Systems venue
  3. 03

    Multi-Agent Runtime Architecture

    Recasts "which language wins" as "which system design preserves throughput." Python and Rust near-parity on throughput; architecture and concurrency strategy drive larger differences. Dual Ollama achieves 99.4% multi-agent efficiency.

    SynthesisTarget: Systems venue
  4. 04

    KV-Cache Quantization and Safety

    KV-cache quantization is a serving-layer perturbation that touches retained attention state. 5-phase paired study on FP16 vs FP8 across 24K records, 3 models. Headline result is a null: no Holm-significant safety effect detectable at α=0.05, 80% power. Operational rule: workload-specific paired eval, not pre-approval.

    In preparationTarget: Workshop submission
    Evidence
  5. 05

    Serving-Stack Physics: When Continuous Batching Stops Amortizing

    A predictive bandwidth prior for the static-batch knee: η(B)=(1+r)/(1+Br) with r=Ck/W (context × KV-bytes-per-token over weight bytes). The parameter-free Ck/W value orders the amortization knee at Spearman ρ=0.84 across three 7-8B models and two datacenter GPUs, validated cross-backend (vLLM/SGLang) and against a served SGLang knee. Revising for resubmission.

    In preparationTarget: Systems venue
  6. 06

    Compile-Stack Attribution

    Independent upstream bugs in PyTorch and Triton jointly produce the torch.compile decode crash. Triton minor-version ablation on the same GPU flips the conclusion. Benchmark identity is a 5-tuple (GPU, Triton, PyTorch, cache, compile mode). Companion to upstream PR #175562 (merged to PyTorch main). Withdrawn from venue review.

    In preparationTarget: Revising for resubmission
    Evidence
  7. 07

    Many-Shot Jailbreak Under Quantization

    Q2_K is the recurring vulnerability threshold for many-shot and long-context attacks. Message-array vs faux-dialogue prompt formatting (92% vs 0% ASR) across 4 model families. Format mediates effect more strongly than quantization alone.

    In preparationTarget: Revising for resubmission
    Evidence
  8. 08

    Multi-Turn Jailbreak × Quantization

    8 attack strategies × 4 models × 6 quantization levels: 10,600 conversations, 37,825 judge labels. Threshold-specific shift in risk rather than universal multi-turn amplification.

    In preparationTarget: Revising for resubmission
    Evidence

3 workshop papers revising after decisions; titles withheld.

The first paper was presented at the ICML 2026 Workshop on Hypothesis Testing, a second was accepted at the NeurIPS 2026 Workshop on Foundation and Large Model Security, and 2 more studies are public on arXiv; 1 is under double-blind review, with 11 in preparation. Each is backed by reproducible technical reports and artifact-level provenance from a 1,463,000+ measurement program.