Gallery inside!
Research

AI Summaries of Annual Reports: Preserve the Evidence Before Polishing the Prose

Research on annual-report summarization explains sentence selection, evidence traceability and the limits of overlap scores for financial disclosures.

6

Select a figure to open it at full size.

The difficult part of summarizing an annual report is deciding what must survive compression. A fluent paragraph can still omit a qualification, separate a number from its reporting period, or repeat management’s optimism while leaving out the risk that makes it meaningful.

For finance and investor-relations teams, a useful starting point is a summary whose sentences remain traceable to the report. Research from Lancaster University offers a concrete example: combine two ways of ranking sentences, remove repetition, and restore the selected sentences to their original order.

Source update: An earlier version relied on Bloated Disclosures: Can ChatGPT Help Investors Process Information? That paper was withdrawn after an independent replication did not support its findings. This article now examines the separate HTAC annual-report summarization study. It does not retain the earlier claims about improved investor confidence or market outcomes.

The Core Insight: Importance and Representativeness Are Different

The Hybrid TF-IDF and Clustering system, HTAC, is an extractive summarizer. It selects existing sentences instead of asking a language model to write new ones. Its design combines a measure of distinctive vocabulary with a measure of how closely a sentence represents a recurring topic.

The first component builds word statistics from the training reports. Words that help distinguish content receive useful weight, while ubiquitous words contribute less. The system adds word scores to produce sentence scores. This is a vocabulary-based ranking, so the later checks for repetition and short fragments remain important.

The second component converts words into numerical representations using Word2Vec, combines these into sentence representations and groups sentences into nine clusters. A sentence close to a cluster’s center is treated as representative of that group. This supplies a different signal from distinctive vocabulary: whether a sentence expresses something central to a recurring theme.

Neither signal is financial materiality. A rare but consequential covenant qualification may be far from a topic center. That distinction is essential when applying the method to disclosures.

How the Research System Builds a Summary

The paper’s method first puts the two scoring systems on a comparable scale and reverses the distance score, so closer sentences rank higher. It combines the scores with 40% weight on TF-IDF and 60% on clustering. The authors selected these settings through comparisons; they are not universal constants for every reporting style.

The system works down the combined ranking, applying two checks before adding a sentence. Very short fragments are excluded. Sentences with substantial word overlap with an already selected sentence are also excluded, reducing the chance that repeated descriptions consume the budget.

It stops within a 1,000-word limit. Finally, it sorts the selected sentences by their positions in the original report. That last step matters: selection determines what enters the summary, while original ordering helps preserve the document’s progression. Ranking order alone might jump from outlook to historic results to a definition with no sensible connection.

The resulting pipeline is inspectable. If a disclosure disappears, an analyst can ask whether its sentence was ranked too low, rejected as repetitive, or squeezed out by the word budget. That is a more actionable diagnosis than simply asking a generative model to make its answer “more accurate.”

What the Research Actually Shows

HTAC was evaluated in the Financial Narrative Summarisation 2022 shared task on English, Spanish and Greek annual reports. The paper reports validation and test results using several ROUGE measures, which compare word or sequence overlap with reference summaries.

HTAC's original test-results table reports ROUGE scores for English, Spanish and Greek annual-report summaries.
Original Table 2, Ogden and El-Haj, FNS 2022. These are text-overlap scores, not percentages of financial facts correctly retained. Original test results. Select for full size.

For example, test ROUGE-2 F-scores were 0.143 for English, 0.134 for Spanish and 0.131 for Greek. These compare pairs of words against the reference. They show the system’s measured performance under that task, but do not answer whether the summary preserves every important risk or supports better investment decisions.

The shared-task organizers explain that the reference summaries were sections identified within the annual reports, rather than fresh summaries commissioned from independent analysts. Matching them therefore measures resemblance to those selected sections. It does not independently establish a balanced view of the company.

The paper also acknowledges that other systems performed better on the English dataset. Its contribution is a specific hybrid extraction method and multilingual evaluation, not proof that this approach is the best summarizer or that generation is unnecessary.

Where Extraction Helps—and Where It Can Mislead

Copying a sentence avoids inventing its wording, but it does not guarantee that the summary is faithful. A statement may depend on a preceding definition, a following exception or a table elsewhere. Sentence extraction can preserve the words while losing the conditions.

Consider an analyst reviewing a change in operating performance. The summary needs to keep the measure, reporting period, comparison basis and material qualification together. A revenue sentence without the explanation that a business was acquired could invite the wrong interpretation. This is a practical review scenario, not an outcome measured by HTAC.

A useful workflow keeps each selected sentence linked to its page and neighboring passage. Reviewers can then check omissions and context without searching the whole report. Generation, if added later, should rewrite that evidence set while retaining its links. This complements research on extracting signals from earnings calls, where polished interpretation also needs a verifiable source.

Implementation Frameworks

The paper uses Gensim’s Word2Vec for word representations and scikit-learn’s KMeans for sentence grouping. These are components for implementing the extraction method, not financial verification tools. A simpler TF-IDF-only extractor is a useful baseline before adding the clustering stage.

Start with one report format and preserve page and sentence identifiers during text extraction. Compare the simple baseline, the hybrid extractor and any proposed generative summary at the same length. Have a finance reviewer mark retained facts, missing qualifications, unsupported statements and time needed to verify the output.

Include unusual but material disclosures in the test set; otherwise a system that mostly selects central themes may look better than it is. Choose scoring weights on separate development reports. A working implementation must explicitly define normalization, sentence boundaries and duplicate handling; the short paper is not a complete production specification.

The acceptance condition is a summary that helps a reviewer find the necessary evidence faster without increasing material omissions. Better ROUGE alone is insufficient. If the output needs extensive reconstruction, retaining the original report with a targeted evidence index may be the better product.

TechClarity’s View

For financial disclosures, traceability should come before elegance. HTAC shows how selection can be decomposed into understandable choices and tested. The strongest practical adaptation is to keep that evidence trail even when a language model improves the presentation. A readable summary earns trust through what it preserves and lets the reader verify.

Original Research

Andrew Ogden and Mahmoud El-Haj, Financial Narrative Summarisation Using a Hybrid TF-IDF and Clustering Summariser: AO-Lancs System at FNS 2022, June 24, 2022, pp. 79–82. Dataset and reference-summary context: FNS 2022 shared task.

Related research

Tags:
Author
TechClarity Analyst Team
September 27, 2026