PaperSubmitted 18 August 2026Page reviewed 28 September 2026

    Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

    Edward Xi Yang (Ertas AI)

    TL;DR

    QMSum ships without a scorer, so we rescored 15 systems under one implementation. A 406M specialist retrained on the same retrieved transcript spans as our 1.2B system scored 36.33 ROUGE-1 against 35.41, a difference QMSum cannot statistically resolve, while using about a third of the parameters and under half the peak inference memory.

    Key findings

    36.33 vs 35.41

    ROUGE-1 on the full QMSum test split

    QMSum cannot separate the span-trained 406M specialist from our 1.2B system: the meeting-cluster 95% interval for the difference is [-0.27, +2.22], and no reported metric separates them.

    2.8x

    fewer total parameters, 439M against 1.2B + 33M

    Both counts include the shared 33M locator. The smaller system also peaks at 2.66 GB of inference memory against 5.73 GB, and trained in about 45 minutes against 4.7 to 5.6 hours.

    +5.29

    ROUGE-1 from span-regime fine-tuning

    Within one fixed 1.2B base model, fine-tuning on retrieved spans adds 5.29 ROUGE-1 [+4.02, +6.56]. Swapping the first 4,500 transcript words for 2,000 retrieved words adds 1.55 on test and 0.29 on validation, so most of the gain comes from the training regime.

    -6.30

    ROUGE-1 when a specialist's input changes

    Moved from capped long input to 2,000-word retrieved spans through our inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1. Fine-tuning it on the span regime recovers the loss.

    6.2+

    ROUGE-1 lead over five hosted proprietary models

    Under one concise prompt and a reference-overlap scorer, a released 406M specialist leads every proprietary hosted model tested by at least 6.2 ROUGE-1. The hosted models' verbose outputs and the absence of human or factuality evaluation limit this ordering.

    Figures and tables

    Figure 1. ROUGE-1 for 15 summarization systems on the QMSum test split

    Horizontal bar chart of ROUGE-1 on the QMSum test split for 15 systems under one scorer. Span-trained SegEnc (406M) 36.33; Ertas LFM2.5-1.2B fine-tune 35.41 promoted and 33.39 protocol-exact; Socratic SegEnc port (406M) 35.30; five hosted proprietary models run zero-shot, 27.89 to 32.43; base LFM2.5-1.2B 30.12 on located spans and 28.57 truncated; four community checkpoints, 9.41 to 28.65.

    The runnable-protocol scale. Every bar uses predictions generated through our evaluation path on the official test split (n=281) and one frozen scorer. The Socratic SegEnc bar is our 35.30 port result, matching the separately labelled port row in Table 1. Systems are sorted by ROUGE-1, with parameter counts written on the bars. Hosted proprietary models ran zero-shot under one concise prompt, and their longer outputs contribute to the gap on this reference-overlap metric; no human or factuality evaluation was run. Full-size image

    Figure 2. ROUGE-1 differences between the promoted 1.2B system and seven others, with 95% intervals

    Chart of ROUGE-1 differences on the QMSum test split between the promoted Ertas LFM2.5-1.2B system and seven others, each with a 95% interval: +7.52 against claude-sonnet-5, +6.76 against distilbart, +6.74 against claude-haiku-4-5, +6.54 against claude-opus-5, +4.49 against gpt-5.6-sol, +2.98 against gpt-5.6-luna, and -3.19 against the stock Socratic SegEnc's released predictions. A shaded band marks the 1.59-point reseeding range around zero.

    Every margin against our promoted system, with both sources of uncertainty. Bars are 95% paired bootstrap intervals over the 281 test queries, and a positive value means our system scores higher. The shaded band is the 1.59-point reseeding range, which the paired bootstrap does not capture. The hosted models' longer outputs contribute to their margins on this reference-overlap metric, and no human or factuality evaluation was run. Full-size image

    Figure 3. What retrieval and fine-tuning each add within one 1.2B base model

    Bar chart of ROUGE-1 on the QMSum test split for one LFM2.5-1.2B base model under three conditions: 28.57 reading the first 4,500 transcript words; 30.12 reading 2,000 words of located spans, a gain of 1.55 (interval +0.58 to +2.51 on test, -0.64 to +1.19 on validation); and 35.41 after Ertas fine-tuned it on those same spans, a further gain of 5.29 (interval +4.02 to +6.56, with validation agreeing at +6.39).

    The two levers, de-confounded. One base model, one adapter, three conditions: only the input regime changes between the first two bars, and only the fine-tune between the second and third. Bracketed ranges are 95% paired bootstrap intervals. Full-size image

    Table 1. Every system on every reported metric, QMSum test split

    SystemParametersROUGE-1ROUGE-2ROUGE-LROUGE-LsumBERTScore
    Socratic-SegEnc (released outputs)406M38.6013.9124.9833.710.8737
    Span-trained SegEnc406M36.3312.7223.6932.170.8710
    Ertas LFM2.5-1.2B fine-tune (promoted)1.2B + 33M35.4112.2824.6331.360.8733
    Socratic-SegEnc (our port)406M35.3011.8723.0030.560.8695
    Ertas LFM2.5-1.2B fine-tune (protocol-exact)1.2B + 22.7M33.3910.6522.8329.300.8680
    GPT-5.6 Luna (zero-shot)undisclosed32.437.8919.0527.490.8608
    GPT-5.6 Sol (zero-shot)undisclosed30.926.9418.3326.100.8583
    LFM2.5-1.2B (zero-shot, spans)1.2B + 33M30.126.7019.2925.040.8687
    Claude Opus 5 (zero-shot)undisclosed28.878.4316.9725.010.8522
    Claude Haiku 4.5 (zero-shot)undisclosed28.668.4217.2624.190.8431
    DistilBART306M28.656.5417.9225.500.8546
    LFM2.5-1.2B (zero-shot, truncated)1.2B28.575.5517.7924.550.8607
    Claude Sonnet 5 (zero-shot)undisclosed27.897.5316.1123.980.8478
    BART-large-CNN406M27.325.6517.4124.080.8513
    PEGASUS570M20.094.4514.9217.230.8344
    LED-base162M9.412.377.957.960.7818

    The single-protocol scale. Full official test split (n=281), one frozen scorer, at most one test touch per system. Every row generated and scored by us, except Socratic-SegEnc (released outputs), scored from the authors' predictions.

    Socratic-SegEnc appears as the authors' released outputs and as our inference port; the port is a calibration, not a reproduction. Hosted zero-shot rows use the full transcript, and their longer outputs contribute to the gap on these reference-overlap metrics. All BERTScores use the same scorer. Parameter counts are measured with num_parameters(); the community checkpoints were fine-tuned by mikeadimech. Highlighted rows are systems trained for this paper.

    The span-trained 406M specialist against the promoted 1.2B system, test split

    MetricErtas LFM2.5-1.2B (promoted)Span-trained SegEnc (406M)Difference95% interval
    ROUGE-135.4136.33+0.93[-0.42, +2.24]
    ROUGE-212.2812.72+0.43[-0.77, +1.64]
    ROUGE-L24.6323.69-0.94[-2.06, +0.18]
    ROUGE-Lsum31.3632.17+0.81[-0.47, +2.05]
    BERTScore0.87330.8710-0.0022[-0.0046, +0.0001]

    The span-trained checkpoint was fixed on validation and taken to the test split once. Differences are span-trained SegEnc minus Ertas LFM2.5-1.2B, with 95% paired bootstrap intervals over 281 queries. Every interval crosses zero, and the meeting-cluster ROUGE-1 interval is [-0.27, +2.22].

    The paper prints this table without a number, in Section 7.5.

    Table 2. Resource use of the 406M and 1.2B systems at an unresolved test difference

    MeasureSpan-trained SegEncErtas LFM2.5-1.2B (promoted)Ratio
    Parameters, including the shared 33M locator439M1.2B + 33M2.8x
    Peak inference VRAM, max over the validation split2.66 GB5.73 GB2.16x
    Peak inference VRAM, median2.66 GB4.71 GB1.77x
    Mean latency per query, two validation runs each, locate step included1.85s and 2.00s2.35s and 2.76s1.3x to 1.4x
    Training to this checkpointAbout 45 minutes (eight matched-chunk epochs; the first four measured at 22 minutes)4.7 to 5.6 hoursLarge
    Test ROUGE-136.3335.41Difference unresolved

    Resource comparison at an unresolved test difference. Both systems include the shared locator; ratios divide our measured value by the span-trained SegEnc value. Crossing intervals do not establish formal equivalence.

    Measured on one NVIDIA RTX 5070 Ti (16 GB). Memory is the summarizer's peak allocated memory over the full validation split.

    Why it matters

    The largest lever this study isolated was training the summarizer on the same retrieved-span inputs it reads at inference, worth 5.29 ROUGE-1 within one base model. A 406M specialist trained that way scored within QMSum's resolution of our 1.2B system while needing under half the peak inference memory.

    For summarization run at volume, that footprint sets serving cost, so a smaller specialist is worth measuring before choosing a larger model.

    Limitations

    All conclusions are limited to QMSum and automatic metrics (ROUGE and BERTScore); no human or factuality evaluation was run. The 406M and 1.2B systems are not separated on any reported metric, and crossing intervals do not establish formal equivalence.

    The strongest released specialist's full-input row is scored from its authors' released predictions, our port of that system is not a reproduction, and training cost was measured on one machine at the edge of its memory.

    Abstract

    QMSum provides no scorer, making query-focused meeting summarization results difficult to compare. We rescore or generate 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1 when moved from capped long input to 2,000-word retrieved spans. Fine-tuning it on this span regime recovers the loss. On test it scores 36.33 ROUGE-1 versus 35.41 for our 1.2B system; the meeting-cluster 95% interval for the difference is [-0.27, +2.22], so QMSum does not statistically separate them. The smaller system uses about one-third as many total parameters and less than half the peak inference memory. Within the fixed 1.2B base, span-regime fine-tuning adds 5.29 [+4.02, +6.56], while replacing the first 4,500 transcript words with 2,000 retrieved words adds 1.55 on test and 0.29 on validation. Separately, under one concise prompt and reference-overlap scorer, a released 406M specialist exceeds five proprietary hosted models by at least 6.2 ROUGE-1, but output length and absent human or factuality evaluation limit this ordering. Conclusions are limited to QMSum and automatic metrics.

    Abstract, arXiv:2609.25028v1

    Paper, code and models

    Also on: Hugging Face paper page · DOI 10.48550/arXiv.2609.25028

    Common questions

    What is retrieved-span training?

    Retrieved-span training fine-tunes a summarizer on the inputs it will see at inference. For each query, a retrieval step scores fixed-width transcript windows and packs the highest-scoring ones, up to 2,000 words, and the model learns to summarize from those spans instead of from the transcript's opening words.

    Is the 406M model as good as the 1.2B model on QMSum?

    On the QMSum test split the 406M model scored 36.33 ROUGE-1 and the 1.2B system 35.41, and the benchmark cannot statistically separate them: the meeting-cluster 95% interval for the difference is [-0.27, +2.22]. Crossing intervals do not establish formal equivalence, so the paper reports the difference as unresolved.

    How does a released 406M specialist compare with hosted proprietary models on QMSum?

    Under one concise prompt and a reference-overlap scorer, a released 406M specialist scored at least 6.2 ROUGE-1 above each of the five proprietary hosted models tested. The paper limits that ordering: the hosted models produced verbose outputs, and no human or factuality evaluation was run.

    Why does QMSum need a common scorer?

    QMSum was released without an official scorer, so each paper generates and scores its own results with different ROUGE packages and generation protocols. The paper rescores or generates 15 systems under one implementation and one frozen protocol on the full 281-query test split.

    Is this state of the art on QMSum?

    No. The strongest published systems, which do not split the transcript into retrieved spans, report higher scores in their own papers, up to 38.82 ROUGE-1, and under our scorer our 1.2B system sits 3.2 ROUGE-1 below the strongest released specialist. The contribution is a common scoring protocol and a cost result, not a new top score.

    Are the code and models available?

    Yes. The code, data build and evaluation scripts are on GitHub, and the four trained components, two summarizers and two retrieval locators, are on Hugging Face under the ErtasAI account.

    About the authors

    Edward Xi Yang, Founder & CEO, Ertas AI

    Edward Xi Yang is the founder and CEO of Ertas AI, which builds custom small language models that run on-device, in customers' own cloud or on Ertas-managed infrastructure. He leads the company's research into training small specialist models for long-document tasks.

    Cite this work

    @misc{yang2026retrievedspan,
      title         = {Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on {QMSum}},
      author        = {Yang, Edward Xi},
      year          = {2026},
      eprint        = {2609.25028},
      archivePrefix = {arXiv},
      primaryClass  = {cs.CL},
      doi           = {10.48550/arXiv.2609.25028},
      url           = {https://arxiv.org/abs/2609.25028}
    }

    References

    Ertas AI builds custom small models trained on your data, deployed on-device, in your own cloud or on Ertas-managed infrastructure. See how Ertas AI works.