Show simple item record

dc.contributor.authorReese, Samuel
dc.contributor.authorBoleda Torrent, Gemma
dc.contributor.authorCuadros Oller, Montserrat
dc.contributor.authorPadró, Lluís
dc.contributor.authorRigau Claramunt, German
dc.contributor.otherUniversitat Politècnica de Catalunya. Departament de Llenguatges i Sistemes Informàtics
dc.date.accessioned2010-06-07T11:47:22Z
dc.date.available2010-06-07T11:47:22Z
dc.date.created2010-05
dc.date.issued2010-05
dc.identifier.citationReese, S. [et al.]. Word-sense disambiguated multilingual Wikipedia corpus. A: International Conference on Language Resources and Evaluation. "7th International Conference on Language Resources and Evaluation". La Valetta: 2010.
dc.identifier.urihttp://hdl.handle.net/2117/7551
dc.description.abstractThis article presents a new freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the Wikipedia and has been automatically enriched with linguistic information. To our knowledge, this is the largest such corpus that is freely available to the community: In its present version, it contains over 750 million words. The corpora have been annotated with lemma and part of speech information using the open source library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assignsWordNet senses, andWordNet has been aligned across languages via the InterLingual Index, this sort of annotation opens the way to massive explorations in lexical semantics that were not possible before. We present a first attempt at creating a trilingual lexical resource from the sense-tagged Wikipedia corpora, namely, WikiNet. Moreover, we present two by-products of the project that are of use for the NLP community: An open source Java-based parser for Wikipedia pages developed for the construction of the corpus, and the integration of the WSD algorithm UKB in FreeLing.
dc.format.extent1 p.
dc.language.isoeng
dc.subject.lcshNatural language processing (Computer science)
dc.subject.lcshWikipedia
dc.titleWord-sense disambiguated multilingual Wikipedia corpus
dc.typeConference report
dc.subject.lemacProcessament de la parla
dc.subject.lemacWikipedia
dc.contributor.groupUniversitat Politècnica de Catalunya. GPLN - Grup de Processament del Llenguatge Natural
dc.description.peerreviewedPeer Reviewed
dc.relation.publisherversionhttp://www.lrec-conf.org/proceedings/lrec2010/pdf/222_Paper.pdf
dc.rights.accessOpen Access
local.identifier.drac2544129
dc.description.versionPostprint (published version)
local.citation.authorReese, S.; Boleda, G.; Cuadros, M.; Padró, L.; Rigau, G.
local.citation.contributorInternational Conference on Language Resources and Evaluation
local.citation.pubplaceLa Valetta
local.citation.publicationName7th International Conference on Language Resources and Evaluation


Files in this item

Thumbnail

This item appears in the following Collection(s)

Show simple item record