Abstract
In this paper we present an approach for an extractive query focused multi-document summarization which stands on an enhanced knowledge-based short text semantic similarity measures. We incorporate WordNet Taxonomy with Categorial Variation Database (CatVar) and Morphosemantic Links to determine query similarity with sentences and intra-sentences similarities. Besides, we enrich WordNet-derived similarity with named entity semantic relatedness inferred from Wikipedia and underpinned by Normalized Google Distance. We show that our summarizer built primarily on such an improved semantic similarity measure to model relevance, centrality and diversity factors outperforms the best-performing relevant DUC systems and recent closely related studies in at least one or more of the investigated ROUGE metrics. An anti-redundancy mechanism is augmented with the proposed summarizer design using Maximum Marginal Relevance algorithm-MMR.
Original language | English |
---|---|
Title of host publication | Proceedings - 9th IEEE International Conference on Big Data Science and Engineering, BigDataSE 2015 |
Publisher | IEEE |
Pages | 80-87 |
Number of pages | 8 |
Volume | 2 |
ISBN (Electronic) | 9781467379519 |
DOIs | |
Publication status | Published - 2 Dec 2015 |
Event | 14th IEEE International Conference on Trust, Security and Privacy in Computing and Communications, TrustCom 2015 - Helsinki, Finland Duration: 20 Aug 2015 → 22 Aug 2015 |
Conference
Conference | 14th IEEE International Conference on Trust, Security and Privacy in Computing and Communications, TrustCom 2015 |
---|---|
Country/Territory | Finland |
City | Helsinki |
Period | 20/08/15 → 22/08/15 |
Keywords
- knowledge-enriched similarity
- named entity relatedness
- query-based summarization
- Word category subsumption