Concept Extraction Using Pointer-Generator Networks
Concept extraction is crucial for a number of downstream applications. However, surprisingly enough, straightforward single token/nominal chunk-concept alignment or dictionary lookup techniques such as DBpedia Spotlight still prevail. We propose a generic open-domain OOV-oriented extractive model th...
Main Authors: | , |
---|---|
Format: | Text |
Language: | unknown |
Published: |
2020
|
Subjects: | |
Online Access: | http://arxiv.org/abs/2008.11295 |
id |
ftarxivpreprints:oai:arXiv.org:2008.11295 |
---|---|
record_format |
openpolar |
spelling |
ftarxivpreprints:oai:arXiv.org:2008.11295 2023-09-05T13:23:43+02:00 Concept Extraction Using Pointer-Generator Networks Shvets, Alexander Wanner, Leo 2020-08-25 http://arxiv.org/abs/2008.11295 unknown http://arxiv.org/abs/2008.11295 Computer Science - Computation and Language text 2020 ftarxivpreprints 2023-08-16T16:03:36Z Concept extraction is crucial for a number of downstream applications. However, surprisingly enough, straightforward single token/nominal chunk-concept alignment or dictionary lookup techniques such as DBpedia Spotlight still prevail. We propose a generic open-domain OOV-oriented extractive model that is based on distant supervision of a pointer-generator network leveraging bidirectional LSTMs and a copy mechanism. The model has been trained on a large annotated corpus compiled specifically for this task from 250K Wikipedia pages, and tested on regular pages, where the pointers to other pages are considered as ground truth concepts. The outcome of the experiments shows that our model significantly outperforms standard techniques and, when used on top of DBpedia Spotlight, further improves its performance. The experiments furthermore show that the model can be readily ported to other datasets on which it equally achieves a state-of-the-art performance. Comment: Contribution to the Proceedings of the 22nd International Conference on Knowledge Engineering and Knowledge Management (EKAW 2020). A link to the final authenticated publication will be added once it is available online. Keywords: Open-domain discourse texts, Concept extraction, Pointer-generator neural network, Distant supervision Text The Pointers ArXiv.org (Cornell University Library) |
institution |
Open Polar |
collection |
ArXiv.org (Cornell University Library) |
op_collection_id |
ftarxivpreprints |
language |
unknown |
topic |
Computer Science - Computation and Language |
spellingShingle |
Computer Science - Computation and Language Shvets, Alexander Wanner, Leo Concept Extraction Using Pointer-Generator Networks |
topic_facet |
Computer Science - Computation and Language |
description |
Concept extraction is crucial for a number of downstream applications. However, surprisingly enough, straightforward single token/nominal chunk-concept alignment or dictionary lookup techniques such as DBpedia Spotlight still prevail. We propose a generic open-domain OOV-oriented extractive model that is based on distant supervision of a pointer-generator network leveraging bidirectional LSTMs and a copy mechanism. The model has been trained on a large annotated corpus compiled specifically for this task from 250K Wikipedia pages, and tested on regular pages, where the pointers to other pages are considered as ground truth concepts. The outcome of the experiments shows that our model significantly outperforms standard techniques and, when used on top of DBpedia Spotlight, further improves its performance. The experiments furthermore show that the model can be readily ported to other datasets on which it equally achieves a state-of-the-art performance. Comment: Contribution to the Proceedings of the 22nd International Conference on Knowledge Engineering and Knowledge Management (EKAW 2020). A link to the final authenticated publication will be added once it is available online. Keywords: Open-domain discourse texts, Concept extraction, Pointer-generator neural network, Distant supervision |
format |
Text |
author |
Shvets, Alexander Wanner, Leo |
author_facet |
Shvets, Alexander Wanner, Leo |
author_sort |
Shvets, Alexander |
title |
Concept Extraction Using Pointer-Generator Networks |
title_short |
Concept Extraction Using Pointer-Generator Networks |
title_full |
Concept Extraction Using Pointer-Generator Networks |
title_fullStr |
Concept Extraction Using Pointer-Generator Networks |
title_full_unstemmed |
Concept Extraction Using Pointer-Generator Networks |
title_sort |
concept extraction using pointer-generator networks |
publishDate |
2020 |
url |
http://arxiv.org/abs/2008.11295 |
genre |
The Pointers |
genre_facet |
The Pointers |
op_relation |
http://arxiv.org/abs/2008.11295 |
_version_ |
1776204309220818944 |