Extension Benchmark Study
36.6% Lower Aggregate AI Model Cost
Across six controlled extension benchmarks, candidate configurations used $13.7132 versus $21.6297 for their baselines. Each extension was evaluated separately; this cost-weighted portfolio result does not measure all extensions enabled together.
Extension Results
Quality is the candidate aggregate pass rate. Lower is better for cost, token, and duration measures.
| Extension | Quality vs baseline | Cost vs baseline | Secondary result |
|---|---|---|---|
| Subagents | 100% | 23.7% lower* | — |
| MarkItDown | 100% | 31.0% lower | 13.8% fewerPrompt tokens |
| Concise Output | 100% | 14.5% lower | 16.4% fewerOutput tokens |
| Prewalk | 100% | 66.0% lower | — |
| RTK | 83.3% | 26.6% lower | 40.6% fewerPrompt tokens |
| Codebase Memory | 100% | 5.2% lower | 3.0% shorterAgent-step duration |