Facebook AI's WMT20 News Translation Task Submission

This paper describes Facebook AI's submission to WMT20 shared news translation task. We focus on the low resource setting and participate in two language pairs, Tamil <-> English and Inuktitut <-> English, where there are limited out-of-domain bitext and monolingual data. We approac...

Full description

Bibliographic Details
Main Authors: Chen, Peng-Jen, Lee, Ann, Wang, Changhan, Goyal, Naman, Fan, Angela, Williamson, Mary, Gu, Jiatao
Format: Text
Language:unknown
Published: 2020
Subjects:
Online Access:http://arxiv.org/abs/2011.08298
id ftarxivpreprints:oai:arXiv.org:2011.08298
record_format openpolar
spelling ftarxivpreprints:oai:arXiv.org:2011.08298 2023-09-05T13:20:41+02:00 Facebook AI's WMT20 News Translation Task Submission Chen, Peng-Jen Lee, Ann Wang, Changhan Goyal, Naman Fan, Angela Williamson, Mary Gu, Jiatao 2020-11-16 http://arxiv.org/abs/2011.08298 unknown http://arxiv.org/abs/2011.08298 Computer Science - Computation and Language text 2020 ftarxivpreprints 2023-08-16T16:12:11Z This paper describes Facebook AI's submission to WMT20 shared news translation task. We focus on the low resource setting and participate in two language pairs, Tamil <-> English and Inuktitut <-> English, where there are limited out-of-domain bitext and monolingual data. We approach the low resource problem using two main strategies, leveraging all available data and adapting the system to the target news domain. We explore techniques that leverage bitext and monolingual data from all languages, such as self-supervised model pretraining, multilingual models, data augmentation, and reranking. To better adapt the translation system to the test domain, we explore dataset tagging and fine-tuning on in-domain data. We observe that different techniques provide varied improvements based on the available data of the language pair. Based on the finding, we integrate these techniques into one training pipeline. For En->Ta, we explore an unconstrained setup with additional Tamil bitext and monolingual data and show that further improvement can be obtained. On the test set, our best submitted systems achieve 21.5 and 13.7 BLEU for Ta->En and En->Ta respectively, and 27.9 and 13.0 for Iu->En and En->Iu respectively. Text inuktitut ArXiv.org (Cornell University Library)
institution Open Polar
collection ArXiv.org (Cornell University Library)
op_collection_id ftarxivpreprints
language unknown
topic Computer Science - Computation and Language
spellingShingle Computer Science - Computation and Language
Chen, Peng-Jen
Lee, Ann
Wang, Changhan
Goyal, Naman
Fan, Angela
Williamson, Mary
Gu, Jiatao
Facebook AI's WMT20 News Translation Task Submission
topic_facet Computer Science - Computation and Language
description This paper describes Facebook AI's submission to WMT20 shared news translation task. We focus on the low resource setting and participate in two language pairs, Tamil <-> English and Inuktitut <-> English, where there are limited out-of-domain bitext and monolingual data. We approach the low resource problem using two main strategies, leveraging all available data and adapting the system to the target news domain. We explore techniques that leverage bitext and monolingual data from all languages, such as self-supervised model pretraining, multilingual models, data augmentation, and reranking. To better adapt the translation system to the test domain, we explore dataset tagging and fine-tuning on in-domain data. We observe that different techniques provide varied improvements based on the available data of the language pair. Based on the finding, we integrate these techniques into one training pipeline. For En->Ta, we explore an unconstrained setup with additional Tamil bitext and monolingual data and show that further improvement can be obtained. On the test set, our best submitted systems achieve 21.5 and 13.7 BLEU for Ta->En and En->Ta respectively, and 27.9 and 13.0 for Iu->En and En->Iu respectively.
format Text
author Chen, Peng-Jen
Lee, Ann
Wang, Changhan
Goyal, Naman
Fan, Angela
Williamson, Mary
Gu, Jiatao
author_facet Chen, Peng-Jen
Lee, Ann
Wang, Changhan
Goyal, Naman
Fan, Angela
Williamson, Mary
Gu, Jiatao
author_sort Chen, Peng-Jen
title Facebook AI's WMT20 News Translation Task Submission
title_short Facebook AI's WMT20 News Translation Task Submission
title_full Facebook AI's WMT20 News Translation Task Submission
title_fullStr Facebook AI's WMT20 News Translation Task Submission
title_full_unstemmed Facebook AI's WMT20 News Translation Task Submission
title_sort facebook ai's wmt20 news translation task submission
publishDate 2020
url http://arxiv.org/abs/2011.08298
genre inuktitut
genre_facet inuktitut
op_relation http://arxiv.org/abs/2011.08298
_version_ 1776201329667997696