DE eng

Search in the Catalogues and Directories

Hits 1 – 12 of 12

1
How OCR Performance can Impact on the Automatic Extraction of Dictionary Content Structures
In: 19th annual Conference and Members’ Meeting of the Text Encoding Initiative Consortium (TEI) -What is text, really? TEI and beyond ; https://hal.archives-ouvertes.fr/hal-02263276 ; 19th annual Conference and Members’ Meeting of the Text Encoding Initiative Consortium (TEI) -What is text, really? TEI and beyond, Sep 2019, Graz, Austria (2019)
BASE
Show details
2
LMF Reloaded
In: AsiaLex 2019: Past, Present and Future ; https://hal.inria.fr/hal-02118319 ; AsiaLex 2019: Past, Present and Future, Jun 2019, Istanbul, Turkey (2019)
BASE
Show details
3
Preparing the Dictionnaire Universel for Automatic Enrichment
In: 10th International Conference on Historical Lexicography and Lexicology (ICHLL) ; https://hal.inria.fr/hal-02131598 ; 10th International Conference on Historical Lexicography and Lexicology (ICHLL), Jun 2019, Leeuwarden, Netherlands ; https://easychair.org/smart-program/ICHLL-10/ (2019)
Abstract: International audience ; The Dictionnaire Universel (DU) is an encyclopaedic dictionary originally written by Antoine Furetière around 1676-78, later revised and improved by the Protestant jurist Henri Basnage de Beauval who expanded, corrected and included terms of arts, crafts and sciences, into the Dictionnaire.The aim of the BASNUM project is to digitize the DU in its second edition rewritten by Basnage de Beauval, to analyse it with computational methods in order to better assess the importance of this work for the evolution of sciences and mentalities in the 18th century, and to contribute to the contemporary movement for creating innovative and data-driven computational methods for text digitization, encoding and analysis.Based on the experience acquired within the research group, an enrichment workflow based upon a series of Natural Language Processing processes is being set up to be applied to Basnage's work. This includes, among others, automatic identification of the dictionary structure (macro-, meso- and microstructure), named-entity recognition (in particular persons and locations), classification of dictionary entries, detection and study of polysemy markers, tracking and classification of quotation use (bibliographic references), scoring semantic similarity between the DU and other dictionaries. The main challenges being the lack of available annotated data in order to train machine learning models, decreased accuracy when using modern pre-trained models due to the differences between present-day and 18th century French, and even unreliable or low quality OCRisation. The paper describes methods that are useful to tackle these issues in order to prepare the the DU for automatic enrichment going beyond what current available tools like Grobid-dictionaries can do, thanks to the advent of deep learning NLP models. The paper also describes how these methods could be applied to other dictionaries or even other types of ancient texts.
Keyword: [INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]; [INFO.INFO-TT]Computer Science [cs]/Document and Text Processing; [SHS.LANGUE]Humanities and Social Sciences/Linguistics; [SHS.MUSEO]Humanities and Social Sciences/Cultural heritage and museology
URL: https://hal.inria.fr/hal-02131598/document
https://hal.inria.fr/hal-02131598
https://hal.inria.fr/hal-02131598/file/ICHLL_10_Slides.pdf
BASE
Hide details
4
Enhancing Usability for Automatically Structuring Digitised Dictionaries
In: GLOBALEX workshop at LREC 2018 ; https://hal.archives-ouvertes.fr/hal-01708137 ; GLOBALEX workshop at LREC 2018, May 2018, Miyazaki, Japan (2018)
BASE
Show details
5
Retro-digitizing and Automatically Structuring a Large Bibliography Collection
In: European Association for Digital Humanities (EADH) Conference ; https://hal.archives-ouvertes.fr/hal-01941534 ; European Association for Digital Humanities (EADH) Conference, EADH, Dec 2018, Galway, Ireland (2018)
BASE
Show details
6
Automatically Encoding Encyclopedic-like Resources in TEI
In: The annual TEI Conference and Members Meeting ; https://hal.inria.fr/hal-01819505 ; The annual TEI Conference and Members Meeting, Sep 2018, Tokyo, Japan ; https://tei2018.dhii.asia/ (2018)
BASE
Show details
7
Presenting the Nénufar Project: a Diachronic Digital Edition of the Petit Larousse Illustré
In: GLOBALEX 2018 - Globalex workshop at LREC2018 ; https://hal.archives-ouvertes.fr/hal-01728328 ; GLOBALEX 2018 - Globalex workshop at LREC2018, May 2018, Miyazaki, Japan. pp.1-6 ; https://globalex.link/globalex2018/ (2018)
BASE
Show details
8
Automatic Extraction of TEI Structures in Digitized Lexical Resources using Conditional Random Fields
In: electronic lexicography, eLex 2017 ; https://hal.archives-ouvertes.fr/hal-01508868 ; electronic lexicography, eLex 2017, Sep 2017, Leiden, Netherlands (2017)
BASE
Show details
9
TermITH-Eval: a French Standard-Based Resource for Keyphrase Extraction Evaluation
In: LREC - Language Resources and Evaluation Conference ; https://hal.archives-ouvertes.fr/hal-01693805 ; LREC - Language Resources and Evaluation Conference, May 2016, Potoroz, Slovenia (2016)
BASE
Show details
10
A Lexicalized Tree-Adjoining Grammar for Vietnamese
In: International Conference on Language Resources and Evaluation - LREC 2006 ; https://hal.inria.fr/inria-00106159 ; International Conference on Language Resources and Evaluation - LREC 2006, May 2006, Gène/Italie, Italy (2006)
BASE
Show details
11
The relevance of standards for research infrastructures
In: International Conference on Language Resources and Evaluation - LREC 2006 ; https://hal.inria.fr/inria-00121474 ; International Conference on Language Resources and Evaluation - LREC 2006, elra, 2006, Gênes/Italie (2006)
BASE
Show details
12
Standards going concrete: from LMF to Morphalou
In: The 20th International Conference on Computational Linguistics - COLING 2004 ; https://hal.inria.fr/inria-00121489 ; The 20th International Conference on Computational Linguistics - COLING 2004, coling, 2004, Genève/Switzerland (2004)
BASE
Show details

Catalogues
0
0
0
0
0
0
0
Bibliographies
0
0
0
0
0
0
0
0
0
Linked Open Data catalogues
0
Online resources
0
0
0
0
Open access documents
12
0
0
0
0
© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern