Home Catalogue search

eng

Refine your search:

Search in the Catalogues and Directories






	Sort by
Simple Search

Hits 1 – 6 of 6

1	Modeling the Unigram Distribution ...
	The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing 2021; Blasi, Damián; Cotterell, Ryan; Nikkarinen, Irene; Pimentel, Tiago. - : Underline Science Inc., 2021
	Abstract: Read paper: https://www.aclanthology.org/2021.findings-acl.326 Abstract: The unigram distribution is the non-contextual probability of finding a specific word form in a corpus. While of central importance to the study of language, it is commonly approximated by each word's sample frequency in the corpus. This approach, being highly dependent on sample size, assigns zero probability to any out-of-vocabulary (oov) word form. As a result, it produces negatively biased probabilities for any oov word form, while positively biased probabilities to in-corpus words. In this work, we argue in favor of properly modeling the unigram distribution---claiming it should be a central task in natural language processing. With this in mind, we present a novel model for estimating it in a language (a neuralization of Goldwater et al.'s (2011) model) and show it produces much better estimates across a diverse set of 7 languages than the naïive use of neural character-level language models. ...
	URL: https://dx.doi.org/10.48448/fx5z-4a29 https://underline.io/lecture/26417-modeling-the-unigram-distribution
	BASE
	Hide details

2	Modeling the Unigram Distribution ...
	Nikkarinen, Irene; Pimentel, Tiago; Blasi, Damián. - : ETH Zurich, 2021
	BASE
	Show details

3	Modeling the Unigram Distribution
	Blasi, Damián; Pimentel, Tiago; Nikkarinen, Irene...
	In: Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 (2021)
	BASE
	Show details

4	How (Non-)Optimal is the Lexicon?
	Cotterell, Ryan; Pimentel, Tiago; Nikkarinen, Irene...
	In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2021)
	BASE
	Show details

5	How (Non-)Optimal is the Lexicon? ...
	NAACL 2021 2021; Blasi, Damián; Cotterell, Ryan. - : Underline Science Inc., 2021
	BASE
	Show details

6	How (Non-)Optimal is the Lexicon? ...
	Pimentel, Tiago; Nikkarinen, Irene; Mahowald, Kyle. - : ETH Zurich, 2021
	BASE
	Show details

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern