Plagiarism detection using information retrieval and similarity measures based on image processing techniques
Document typeConference report
Rights accessOpen Access
This paper describes the Barcelona Media Innovation Center participation in the 2nd International Competition on Plagiarism Detection. Particularly, our system focused on the external plagiarism detection task, which assumes the source documents are available. We present a two-step a approach. In the first step of our method, we build an information retrieval system based on Solr/Lucene, segmenting both suspicious and source documents into smaller texts.We perform a search based on bag-of-words which provides a first selection of potentially plagiarized texts. In the second step, each promising pair is further investigated. We implemented a sliding window approach that computes cosine distances between overlapping text segments from both the source and suspicious documents on a pair wise basis. As a result, a similarity matrix between text segments is obtained, which is smoothed by means of low-pass 2-D filtering. From the smoothed similarity matrix, plagiarized segments are identified by using image processing techniques. Our results were placed in the middle of the official ranking, which considered together two types of plagiarism: intrinsic and external.
CitationRuiz, M., Banchs, R., Grivolla, J., Codina, J. Plagiarism detection using information retrieval and similarity measures based on image processing techniques. A: Conference on Multilingual and Multimodal Information Access Evaluation. "Notebook Papers of CLEF 2010 Labs and Workshops, 22-23 September, Padua, Italy, September 2010". 2010.