Mostra el registre d'ítem simple

dc.contributor.authorRuiz Costa-Jussà, Marta
dc.contributor.authorBanchs, Rafael E.
dc.contributor.authorGrivolla, Jens
dc.contributor.authorCodina, Joan
dc.contributor.otherUniversitat Politècnica de Catalunya. Departament de Teoria del Senyal i Comunicacions
dc.date.accessioned2017-03-13T17:38:48Z
dc.date.available2017-03-13T17:38:48Z
dc.date.issued2010
dc.identifier.citationRuiz, M., Banchs, R., Grivolla, J., Codina, J. Plagiarism detection using information retrieval and similarity measures based on image processing techniques. A: Conference on Multilingual and Multimodal Information Access Evaluation. "Notebook Papers of CLEF 2010 Labs and Workshops, 22-23 September, Padua, Italy, September 2010". 2010.
dc.identifier.isbn978-88-904810-2-4
dc.identifier.urihttp://hdl.handle.net/2117/102408
dc.description.abstractThis paper describes the Barcelona Media Innovation Center participation in the 2nd International Competition on Plagiarism Detection. Particularly, our system focused on the external plagiarism detection task, which assumes the source documents are available. We present a two-step a approach. In the first step of our method, we build an information retrieval system based on Solr/Lucene, segmenting both suspicious and source documents into smaller texts.We perform a search based on bag-of-words which provides a first selection of potentially plagiarized texts. In the second step, each promising pair is further investigated. We implemented a sliding window approach that computes cosine distances between overlapping text segments from both the source and suspicious documents on a pair wise basis. As a result, a similarity matrix between text segments is obtained, which is smoothed by means of low-pass 2-D filtering. From the smoothed similarity matrix, plagiarized segments are identified by using image processing techniques. Our results were placed in the middle of the official ranking, which considered together two types of plagiarism: intrinsic and external.
dc.language.isoeng
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/3.0/es/
dc.subjectÀrees temàtiques de la UPC::Informàtica
dc.subject.lcshPlagiarism
dc.subject.otherPlagiarism detection
dc.subject.otherInformation retrieval
dc.titlePlagiarism detection using information retrieval and similarity measures based on image processing techniques
dc.typeConference report
dc.subject.lemacPlagi
dc.contributor.groupUniversitat Politècnica de Catalunya. VEU - Grup de Tractament de la Parla
dc.rights.accessOpen Access
local.identifier.drac19719602
dc.description.versionPostprint (published version)
local.citation.authorRuiz, M.; Banchs, R.; Grivolla, J.; Codina, J.
local.citation.contributorConference on Multilingual and Multimodal Information Access Evaluation
local.citation.publicationNameNotebook Papers of CLEF 2010 Labs and Workshops, 22-23 September, Padua, Italy, September 2010


Fitxers d'aquest items

Thumbnail

Aquest ítem apareix a les col·leccions següents

Mostra el registre d'ítem simple