Skip to content
CLILLM deployment planner

Chimeraforge

Turns "which model, quantization, GPU, and backend — how many, will it fit, will it hit my SLO, what will it cost" into a fast, measured answer from your shell, your Python, or your AI assistant.

  • v0.34.0
  • MIT
  • Python 3.10 – 3.14
pip install chimeraforge
PyPISourceChangelog
uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
How plan decidesEvery candidate passes five gates in order; a plan that finds nothing names the gate that stopped it.
  1. Candidates

    model × quantization × backend × replicas × batch

  2. VRAM

    weights and KV-cache from the real architecture

  3. Quality

    measured, estimated, or unknown

  4. Safety

    opt-in: the TR134/TR142 refusal-rate lookup

  5. Latency

    TTFT and TPOT against your SLO

  6. Budget

    GPU $/hr × fleet size

  7. Plan

    the cheapest config that meets your SLO

The trust principle

Every number is labeled measured, extrapolated, derived, estimated, or unknown — and the tool refuses to fake the ones it cannot stand behind. VRAM and KV-cache are derived: exact arithmetic over the model’s real architecture, not a measurement. Throughput is a measured lookup only on the rig the corpus was measured on; on any other GPU it is scaled by memory bandwidth and labeled extrapolated, otherwise an explicit roofline estimate, never dressed up as data. Quality below the bundled corpus reports unknown rather than an invented score, and a zero-result plan names the exact gate that rejected every candidate.

Features

  • 01

    Model-agnostic planning

    Plan any registry name, Ollama tag, or HuggingFace repo — not just a bundled list. Tensor-parallel splitting for models too big for one card, and a KV-cache trade-off menu across cost, latency, and quality.

  • 02

    An MCP server

    chimeraforge mcp exposes the same planner to Claude, Cursor, and other MCP clients, so an assistant answers capacity questions from measured numbers instead of guessing.

  • 03

    22 GPU profiles

    Consumer Ampere, Ada, and Blackwell (RTX 30/40/50-series), datacenter (A100 40/80GB, H100, H200, B200, L4, T4), and AMD MI300X — each with VRAM, bandwidth, FP16 TFLOPS, TDP, and interconnect.

Commands

13 commands. Full flags and output samples live in the README.

  • plan

    predictive capacity planner

  • suggest

    discover and rank models that fit

  • measure

    benchmark live, plan on real numbers

  • workload

    derive plan inputs from real traffic

  • validate

    audit predictions against measurements

  • catalog

    local model catalog

  • safety

    live refusal screen

  • bench

    live inference benchmarking

  • eval

    quality evaluation

  • compare

    diff benchmark runs

  • refit

    update planner coefficients

  • report

    generate reports

  • mcp

    serve the planner to AI assistants

Evidence

Each capability traces to the measurements behind it. These are the published reports, not a summary of them.

CapabilityWhat it rests onReports
Capacity planningThe predictive planner and its coefficients come out of the optimization phase — KV-cache tuning, context scaling, and the capacity-planner report itself.TR133
Throughput and scalingBackend parity, scaling laws, and the inference-physics work behind the roofline estimates and the continuous-batching curve.TR120TR122
The opt-in safety gateThe refusal-rate lookup that powers plan --safety-target is a measured table from the safety-pivot reports, not a model fitted after the fact.TR134TR142
KV-cache precisionThe standardized FP16-versus-FP8 KV-cache battery behind the quantized-cache guidance.TR149

What it does not do

The limits the tool states about itself.

  • Speculative decoding is not modeled yet. For MoE, active-versus-total parameters are modeled, but expert parallelism and routing load imbalance are not.
  • Quantization coverage for vLLM, TGI, and SGLang is FP16, FP8, and AWQ/GPTQ; FP8 and W4A16 quality are estimated, not measured, because the quality corpus covers GGUF k-quants only.
  • The bundled quality corpus is 20 items, which resolves nothing smaller than about 21 percentage points, so every measured quant delta in it reports as indistinguishable from its FP16 baseline.
  • Heterogeneous fleets assume a capability-aware request router that no serving engine ships yet.
  • Tensor- and pipeline-parallel throughput are comms-modeled estimates, not measured, and cannot be combined in a single plan.
  • The bundled corpus is fit primarily on one rig (RTX 4080 12GB); other GPUs scale from bandwidth and compute until you run measure.