An Efficient Machine Learning Approach for Plagiarism Detection in Text Documents
Clicks: 3
ID: 313249
2023
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
Emerging Content
0.6
/100
3 views
2 readers
AI Quality Assessment
Not analyzed
Readership in this journal
EmergingRanked #557 of 705 articles by views in Journal of Computing & Biomedical Informatics
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 705 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Plagiarism is when you use someone else's words or ideas as your own. On every subject, the internet is a reliable source of information. People can therefore simply copy data and use various techniques to cover up plagiarism. Extrinsic and intrinsic methods can both be used to detect plagiarism. Extrinsic plagiarism involves comparing the source and the allegedly plagiarised texts to one another in order to obtain precise similarity metrics like Jaccard and Cosine. Source materials are not necessary for intrinsic plagiarism recognition, though. Plagiarism can be detected by an author's writing style and other notable actions. Cross-Lingual Plagiarism (CLP) is a kind of plagiarism in which the author steals content by translating text from one language to another like Urdu-English. It is hard to identify CLP because the source and suspicious documents are in two different languages. In this regard, various approaches to tackling the problem of CPD in text documents were presented. We need to apply ML way to deal with the problems of CLPD. For PD task curpus is used to evaluate the performance of PD, so we use Urdu English language pair Corpus CLPD UE 19 [1]. The source text in the language-pair corpus (CLPD-UE-19) is written in Urdu, whereas the suspicious text is supplied in English. To construct a dataset that can be understood by machine learning tools and extract optimized features from a corpus using Python NLP techniques. Our created dataset is in the CSV format in which there are distinctive features of source, and suspected content is mentioned like Jaccard similarity and Cosine similarity. We have used one gram and tri-gram of the preprocessed text to get comparability measures. five ML classifiers, such as KNN, Naïve Bayes, SVM, Decision Tree, and Random Forest, are utilized to build models. Python language is used on PyCharm tool to build models from various classifiers. We use two methods to examine the models' accuracy (cross-validation and percentage split) in the python language. The trial shows that KNN, RF and have produce better results as compared to other models.
| Reference Key |
imported_1777060022_69ebc8b62a2c6
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Naeem Aslam |
| Journal | Journal of Computing & Biomedical Informatics |
| Year | 2023 |
| DOI |
DOI not found
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.