Home Catalogue search

eng

Refine your search:
- Keyword:
- Creator / Publisher
- Year
- Medium
- Type:
  - Miscellaneous (6)
  - Article (2)
- BLLDB-Access:
  - free (8)
  - subject to license (0)

Search in the Catalogues and Directories






	Sort by
Simple Search

Hits 1 – 8 of 8

1	Universal Segmentations 1.0 (UniSegments 1.0)
	Žabokrtský, Zdeněk; Bafna, Nyati; Bodnár, Jan. - : Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2022
	BASE
	Show details

2	Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0
	Kuvač Kraljević, Jelena; Hržica, Gordana; Štefanec, Vanja. - : Jožef Stefan Institute, 2021. : Faculty of Education and Rehabilitation, University of Zagreb, 2021
	BASE
	Show details

3	The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.1
	Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
	BASE
	Show details

4	The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.0
	Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
	BASE
	Show details

5	The CLASSLA-StanfordNLP model for lemmatisation of non-standard Serbian 1.0
	Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
	BASE
	Show details

6	The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Serbian 1.0
	Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
	Abstract: This model for morphosyntactic annotation of non-standard Serbian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the SETimes.SR training corpus (http://hdl.handle.net/11356/1200), the ReLDI-NormTagNER-sr corpus (http://hdl.handle.net/11356/1240), the ReLDI-NormTagNER-hr corpus (http://hdl.handle.net/11356/1241), the hr500k training corpus (http://hdl.handle.net/11356/1210) and the RAPUT corpus (https://www.aclweb.org/anthology/L16-1513/), using the CLARIN.SI-embed.sr word embeddings (http://hdl.handle.net/11356/1206). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.91.
	Keyword: computer-mediated communication; language model; part-of-speech tagging
	URL: http://hdl.handle.net/11356/1332
	BASE
	Hide details

7	The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0
	Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
	BASE
	Show details

8	The CLASSLA-StanfordNLP model for lemmatisation of non-standard Croatian 1.1
	Ljubešić, Nikola; Štefanec, Vanja. - : Jožef Stefan Institute, 2020
	BASE
	Show details

© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern