« In this article we present the design and implementation of the Logoscope, the first tool especially developed to detect new words of the French language, to document them and allow a public access through a web interface. (…) »
« TELMA –Traitement électronique des manuscrits et des archives– est une collection de l’IRHT, dédiée aux éditions critiques électroniques et à la publication de répertoires numériques de sources médiévales. Née en 2005 comme un centre de ressources numériques du CNRS en collaboration avec l’École nationale…
« HathiTrust has reached a tremendous milestone in the history of HathiTrust and the HathiTrust Research Center’s services.
Since 2011, HTRC has been developing services and tools to allow researchers to employ text and data mining methodologies using the HathiTrust collection. To date, this service has been available only on the…
« Le projet interdisciplinaire TERRE-ISTEX a pour objectif d’identifier l’évolution des fronts de recherche en relation avec les territoires d’études, les croisements disciplinaires ainsi que les modalités concrètes de recherche à partir des contenus numériques hétérogènes disponibles dans les corpus scientifiques. Le projet se décompose en trois actions principales~: (1) identifier…
« The OpenMinTeD platform aims to bring full text Open Access scholarly content from a wide range of providers together with Text and Data Mining (TDM) tools from various Natural Language Processing frameworks and TDM developers in an integrated environment. In this way, it supports users who want to mine scientific…
« Since the first LREC held in Granada in 1998, LREC has become the major event on Language Resources (LRs) and Evaluation for Language Technologies (LT). The aim of LREC is to provide an overview of the state-of-the-art, explore new R&D directions and emerging trends, exchange information regarding LRs and their…
« Présentation du projet VisaTM
La création d’une offre de service en fouille de texte et de données – TDM (Text and Data Mining) – à destination des scientifiques se pose dans un contexte évolutif sur le plan légal, organisationnel et scientifique. Les progrès récents des méthodes d’analyse textuelle ouvrent…
« À la suite des ateliers « Décrire, transcrire et diffuser un corpus documentaire hétérogène : méthodes, formats, outils », « Géolocalisation et spatialisation de documents patrimoniaux » et « Explorer des corpus d’images. L’IA au service du patrimoine », une quatrième demi-journée d’étude a été…
« La BNF organisait le 10 juillet 2018 un atelier « Données liées et données à lier : quels outils pour quels alignements ?« , avec plein de bonnes choses dedans (…) »