← Benchmarks

Inherited-model session compaction

felan-inherit-compaction vs felan-no-session-compaction

Cost · minimize · Ratio of reduced sums · lower is better

+45.2%primary outcome · Ratio of reduced sums·+45.2% to +45.2%Case range (macro mean)
Costprimary · minimize · Ratio of reduced sums+45.2%Aggregate values: baseline $0.4955; candidate $0.2713 · case mean +45.2%; Case range (macro mean) +45.2% to +45.2%
Prompt tokenssecondary · minimize · Ratio of reduced sums+52.1%Aggregate values: baseline 83,304 tokens; candidate 39,891 tokens · case mean +52.1%; Case range (macro mean) +52.1% to +52.1%

Metrics

Arm values remain macro means across cases. The objective headline uses the ratio of reduced sums after per-case median trial reduction.

Benchmark metrics by arm
Metricfelan-no-session-compactionbaseline · 3/3 runs · 2/3 passed; 1 failed · verifierfelan-inherit-compactioncandidate · 3/3 runs · 3/3 passedDelta
Cost$0.4955$0.2713-$0.2242
Prompt tokens83,304 tokens39,891 tokens-43,413 tokens
quality.passRate110
Duration6209559720-2375
duration.stepsMs6079258731-2061
Uncached input tokens83,261 tokens39,856 tokens-43,405 tokens
Cache-read input tokens0 tokens0 tokens0 tokens
Cache-write input tokens0 tokens0 tokens0 tokens
Output tokens2,633 tokens2,401 tokens-232 tokens
Requests220

Tests

TestBaseline (Median trials; Ratio of reduced sums)Candidate (Median trials; Ratio of reduced sums)vs baseline
session-compaction-continuation$0.49552/3 passed; 1 failed · verifier$0.27133/3 passed+45.2%
Attempt 1$0.4951Failed · verifier$0.2708Passed+45.3%
Attempt 2$0.4955Passed$0.2713Passed+45.2%
Attempt 3$0.5020Passed$0.2807Passed+44.1%