Improving the Identification of the Discourse Function of News Article Paragraphs

Deya Banisakher, W. Victor Yarlott, Mohammed Aldawsari, Naphtali Rishe, Mark Finlayson


Abstract
Identifying the discourse structure of documents is an important task in understanding written text. Building on prior work, we demonstrate an improved approach to automatically identifying the discourse function of paragraphs in news articles. We start with the hierarchical theory of news discourse developed by van Dijk (1988) which proposes how paragraphs function within news articles. This discourse information is a level intermediate between phrase- or sentence-sized discourse segments and document genre, characterizing how individual paragraphs convey information about the events in the storyline of the article. Specifically, the theory categorizes the relationships between narrated events and (1) the overall storyline (such as Main Events, Background, or Consequences) as well as (2) commentary (such as Verbal Reactions and Evaluations). We trained and tested a linear chain conditional random field (CRF) with new features to model van Dijk’s labels and compared it against several machine learning models presented in previous work. Our model significantly outperformed all baselines and prior approaches, achieving an average of 0.71 F1 score which represents a 31.5% improvement over the previously best-performing support vector machine model.
Anthology ID:
2020.nuse-1.3
Volume:
Proceedings of the First Joint Workshop on Narrative Understanding, Storylines, and Events
Month:
July
Year:
2020
Address:
Online
Editors:
Claire Bonial, Tommaso Caselli, Snigdha Chaturvedi, Elizabeth Clark, Ruihong Huang, Mohit Iyyer, Alejandro Jaimes, Heng Ji, Lara J. Martin, Ben Miller, Teruko Mitamura, Nanyun Peng, Joel Tetreault
Venues:
NUSE | WNU
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
17–25
Language:
URL:
https://aclanthology.org/2020.nuse-1.3
DOI:
10.18653/v1/2020.nuse-1.3
Bibkey:
Cite (ACL):
Deya Banisakher, W. Victor Yarlott, Mohammed Aldawsari, Naphtali Rishe, and Mark Finlayson. 2020. Improving the Identification of the Discourse Function of News Article Paragraphs. In Proceedings of the First Joint Workshop on Narrative Understanding, Storylines, and Events, pages 17–25, Online. Association for Computational Linguistics.
Cite (Informal):
Improving the Identification of the Discourse Function of News Article Paragraphs (Banisakher et al., NUSE-WNU 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.nuse-1.3.pdf
Video:
 http://slideslive.com/38929742