ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data

Large pretrained language models have been performing increasingly well in a variety of downstream tasks via prompting. However, it remains unclear from where the model learns the task-specific knowledge, especially in a zero-shot setup. In this work, we want to find evidence of the model's tas...

Full description

Bibliographic Details
Main Authors:	Han, Xiaochuang, Tsvetkov, Yulia
Format:	Text
Language:	unknown
Published:	2022
Subjects:	Computer Science - Computation and Language Computer Science - Machine Learning Orca
Online Access:	http://arxiv.org/abs/2205.12600

id	ftarxivpreprints:oai:arXiv.org:2205.12600
record_format	openpolar
spelling	ftarxivpreprints:oai:arXiv.org:2205.12600 2023-09-05T13:22:21+02:00 ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data Han, Xiaochuang Tsvetkov, Yulia 2022-05-25 http://arxiv.org/abs/2205.12600 unknown http://arxiv.org/abs/2205.12600 Computer Science - Computation and Language Computer Science - Machine Learning text 2022 ftarxivpreprints 2023-08-16T17:06:03Z Large pretrained language models have been performing increasingly well in a variety of downstream tasks via prompting. However, it remains unclear from where the model learns the task-specific knowledge, especially in a zero-shot setup. In this work, we want to find evidence of the model's task-specific competence from pretraining and are specifically interested in locating a very small subset of pretraining data that directly supports the model in the task. We call such a subset supporting data evidence and propose a novel method ORCA to effectively identify it, by iteratively using gradient information related to the downstream task. This supporting data evidence offers interesting insights about the prompted language models: in the tasks of sentiment analysis and textual entailment, BERT shows a substantial reliance on BookCorpus, the smaller corpus of BERT's two pretraining corpora, as well as on pretraining examples that mask out synonyms to the task verbalizers. Text Orca ArXiv.org (Cornell University Library)
institution	Open Polar
collection	ArXiv.org (Cornell University Library)
op_collection_id	ftarxivpreprints
language	unknown
topic	Computer Science - Computation and Language Computer Science - Machine Learning
spellingShingle	Computer Science - Computation and Language Computer Science - Machine Learning Han, Xiaochuang Tsvetkov, Yulia ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data
topic_facet	Computer Science - Computation and Language Computer Science - Machine Learning
description	Large pretrained language models have been performing increasingly well in a variety of downstream tasks via prompting. However, it remains unclear from where the model learns the task-specific knowledge, especially in a zero-shot setup. In this work, we want to find evidence of the model's task-specific competence from pretraining and are specifically interested in locating a very small subset of pretraining data that directly supports the model in the task. We call such a subset supporting data evidence and propose a novel method ORCA to effectively identify it, by iteratively using gradient information related to the downstream task. This supporting data evidence offers interesting insights about the prompted language models: in the tasks of sentiment analysis and textual entailment, BERT shows a substantial reliance on BookCorpus, the smaller corpus of BERT's two pretraining corpora, as well as on pretraining examples that mask out synonyms to the task verbalizers.
format	Text
author	Han, Xiaochuang Tsvetkov, Yulia
author_facet	Han, Xiaochuang Tsvetkov, Yulia
author_sort	Han, Xiaochuang
title	ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data
title_short	ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data
title_full	ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data
title_fullStr	ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data
title_full_unstemmed	ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data
title_sort	orca: interpreting prompted language models via locating supporting data evidence in the ocean of pretraining data
publishDate	2022
url	http://arxiv.org/abs/2205.12600
genre	Orca
genre_facet	Orca
op_relation	http://arxiv.org/abs/2205.12600
_version_	1776202868722761728

ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data

Similar Items