Chimeraforge
Turns "which model, quantization, GPU, and backend — how many, will it fit, will it hit my SLO, what will it cost" into a fast, measured answer from your shell, your Python, or your AI assistant.
uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"Candidates
model × quantization × backend × replicas × batch
VRAM
weights and KV-cache from the real architecture
Quality
measured, estimated, or unknown
Safety
opt-in: the TR134/TR142 refusal-rate lookup
Latency
TTFT and TPOT against your SLO
Budget
GPU $/hr × fleet size
Plan
the cheapest config that meets your SLO
The trust principle
Every number is labeled measured, extrapolated, derived, estimated, or unknown — and the tool refuses to fake the ones it cannot stand behind. VRAM and KV-cache are derived: exact arithmetic over the model’s real architecture, not a measurement. Throughput is a measured lookup only on the rig the corpus was measured on; on any other GPU it is scaled by memory bandwidth and labeled extrapolated, otherwise an explicit roofline estimate, never dressed up as data. Quality below the bundled corpus reports unknown rather than an invented score, and a zero-result plan names the exact gate that rejected every candidate.
Features
- 01
Model-agnostic planning
Plan any registry name, Ollama tag, or HuggingFace repo — not just a bundled list. Tensor-parallel splitting for models too big for one card, and a KV-cache trade-off menu across cost, latency, and quality.
- 02
An MCP server
chimeraforge mcp exposes the same planner to Claude, Cursor, and other MCP clients, so an assistant answers capacity questions from measured numbers instead of guessing.
- 03
22 GPU profiles
Consumer Ampere, Ada, and Blackwell (RTX 30/40/50-series), datacenter (A100 40/80GB, H100, H200, B200, L4, T4), and AMD MI300X — each with VRAM, bandwidth, FP16 TFLOPS, TDP, and interconnect.
Commands
13 commands. Full flags and output samples live in the README.
planpredictive capacity planner
suggestdiscover and rank models that fit
measurebenchmark live, plan on real numbers
workloadderive plan inputs from real traffic
validateaudit predictions against measurements
cataloglocal model catalog
safetylive refusal screen
benchlive inference benchmarking
evalquality evaluation
comparediff benchmark runs
refitupdate planner coefficients
reportgenerate reports
mcpserve the planner to AI assistants
Evidence
Each capability traces to the measurements behind it. These are the published reports, not a summary of them.
| Capability | What it rests on | Reports |
|---|---|---|
| Capacity planning | The predictive planner and its coefficients come out of the optimization phase — KV-cache tuning, context scaling, and the capacity-planner report itself. | TR133 |
| Throughput and scaling | Backend parity, scaling laws, and the inference-physics work behind the roofline estimates and the continuous-batching curve. | TR120TR122 |
| The opt-in safety gate | The refusal-rate lookup that powers plan --safety-target is a measured table from the safety-pivot reports, not a model fitted after the fact. | TR134TR142 |
| KV-cache precision | The standardized FP16-versus-FP8 KV-cache battery behind the quantized-cache guidance. | TR149 |
What it does not do
The limits the tool states about itself.
- Speculative decoding is not modeled yet. For MoE, active-versus-total parameters are modeled, but expert parallelism and routing load imbalance are not.
- Quantization coverage for vLLM, TGI, and SGLang is FP16, FP8, and AWQ/GPTQ; FP8 and W4A16 quality are estimated, not measured, because the quality corpus covers GGUF k-quants only.
- The bundled quality corpus is 20 items, which resolves nothing smaller than about 21 percentage points, so every measured quant delta in it reports as indistinguishable from its FP16 baseline.
- Heterogeneous fleets assume a capability-aware request router that no serving engine ships yet.
- Tensor- and pipeline-parallel throughput are comms-modeled estimates, not measured, and cannot be combined in a single plan.
- The bundled corpus is fit primarily on one rig (RTX 4080 12GB); other GPUs scale from bandwidth and compute until you run measure.