Translate this page into:
The hidden cost of plagiarism and artificial intelligence detection tools in academic writing
*Corresponding author: Deep Pankaj Shah, Department of Community Medicine, Gujarat Medical Education and Research Society Medical College, Navsari 396406, Gujarat, India. deeppankajshah@gmail.com
-
Received: ,
Accepted: ,
How to cite this article: Shah DP. The hidden cost of plagiarism and artificial intelligence detection tools in academic writing. RMC Glob J. 2026;2:97-101. doi: 10.25259/RMCGJ_10_2026
Abstract
Plagiarism-detection and AI-detection tools are now widely used in academic publishing. These systems were introduced to support research integrity by helping journals identify potential plagiarism, inappropriate text reuse, and concerns related to undisclosed use of artificial intelligence (AI). When used appropriately, they can serve as useful screening tools and support editorial decision-making. However, their growing use has also created new challenges. Similarity scores are often interpreted as direct measures of plagiarism, even though they only indicate matching text and require contextual evaluation. Likewise, AI-detection tools can produce uncertain or incorrect classifications, yet their output may influence perceptions of authorship and manuscript quality. As a result, researchers may spend considerable time reducing similarity scores or worrying about AI-detection reports, even when the underlying writing is appropriate. These challenges may be particularly relevant for early-career researchers and authors writing in a second language. This article discusses the benefits and limitations of both similarity-detection and AI-detection systems and argues that their output should be viewed as screening indicators rather than definitive judgments. Human interpretation should remain central to the evaluation of originality, authorship, and research quality.
Keywords
Academic writing
AI detection
Plagiarism detection
Research evaluation
Similarity tools
Academic writing in the age of automated evaluation
Academic writing is central to how research is shared, evaluated, and preserved. Across disciplines, it is expected to be clear, original, and transparent so that scientific knowledge can be communicated effectively and trusted by the wider community. Maintaining these qualities has always been an important part of scholarly publishing. In recent years, however, academic writing and publishing have become increasingly supported by digital tools.1 Journals now commonly use software systems to assist with manuscript screening, particularly tools that check text similarity and, more recently, tools that attempt to detect the use of artificial intelligence (AI) in writing.2 These systems are now widely integrated into editorial workflows and are often used at the early stages of manuscript evaluation.
The main reason for adopting these tools is to support research integrity. With the growing volume of submissions to journals and the increasing concern about plagiarism, duplicate publication, and unethical writing practices, similarity-detection software offers a practical way to screen manuscripts efficiently. In a similar way, the rise of generative AI has led to new questions about transparency in authorship,3 which has encouraged some publishers to explore AI-detection tools as an additional safeguard. While these systems can be useful for identifying manuscripts that may require closer attention, challenges arise when their outputs are treated as final judgments rather than preliminary indicators. Similarity scores and AI-detection results are often seen as objective measures, even though they require careful interpretation and contextual understanding. This creates a tension between the supportive role of these tools and the way their results are sometimes used in practice.
This article examines that tension. It discusses the role of similarity-detection and AI-detection systems in academic publishing, highlights their strengths and limitations, and argues that their outputs should be treated as screening signals rather than definitive assessments of originality or authorship. Human judgment, supported but not replaced by software, must remain central in evaluating academic work.
Challenges associated with similarity-detection systems
Similarity-detection software is now a commonly used tool in academic publishing. It works by comparing a submitted manuscript against large databases of published literature and online sources to identify overlapping text. Its main purpose is to help detect possible plagiarism and ensure proper attribution of sources.4 In this sense, it plays an important role in supporting research integrity. However, a key limitation lies in how its output is interpreted. A similarity score reflects textual overlap, not academic misconduct.5 Whether overlap represents plagiarism depends on context, including citation practices, disciplinary norms, and the type of content being assessed. This distinction is sometimes overlooked when numerical thresholds are used as decision-making tools.
Another important issue is the nature of scientific writing itself. Academic texts, especially in methods sections, often contain standardized and repeated expressions. Phrases describing study design, data collection, or statistical procedures may appear across many publications simply because they are conventionally used and clearly understood. As a result, similar tools may highlight text that is not problematic, which can lead to unnecessary concern or revision. Figure 1 provides examples of how similarity-detection tools may flag common methodological statements that are routinely used in scientific writing. As shown in both reports, standard methods text can contribute to similarity scores even when the overlap reflects conventional reporting language rather than inappropriate text reuse.

The interpretation of similarity reports is also affected by variability between systems. Different platforms use different databases and algorithms, which means that the same manuscript may produce different similarity scores depending on the tool used.6 This lack of consistency can make it difficult for authors to understand what level of similarity is actually meaningful. The practical consequence is that researchers may spend time modifying well-written and scientifically appropriate text simply to reduce similarity percentages. While some rewriting may improve clarity, excessive focus on numerical scores can shift attention away from the primary goal of academic writing, which is clear and accurate communication of research.
Challenges associated with AI-detection systems
AI-detection tools represent a different approach to manuscript screening. Instead of comparing text with existing sources, they attempt to estimate whether a piece of writing was produced by artificial intelligence. This is done using statistical and linguistic patterns rather than direct text matching.7 This difference makes AI detection fundamentally more uncertain than similarity checking. The output is not based on identifiable sources but on probability-based classification. As a result, AI-detection scores cannot be interpreted as direct evidence of AI use. They represent estimates that may vary depending on writing style, structure, and training data used by the detection model.
The interpretation of AI-detection results is also complicated by variability between systems. Because different tools rely on different algorithms, training data, and detection methods, the same manuscript may be classified differently across AI-detection platforms. This variability is illustrated in Figure 2, which presents AI-detection reports for the same manuscript generated by two different AI-detection tools. The reports show different proportions of AI-generated content, highlighting how detection results may vary across platforms and should be interpreted with caution.

The reports show different estimated proportions of AI-generated content, illustrating the variability that can occur across detection platforms. Note: Both AI-detection reports include a disclaimer stating that the results are not fully reliable and should be interpreted with caution, with human review recommended.
Concerns about accuracy have been widely discussed. Studies evaluating AI-detection systems have reported both false positives and false negatives.8,9 Human-written text may be incorrectly flagged as AI-generated, particularly when the writing is clear, structured, and grammatically consistent. At the same time, AI-generated content may sometimes go undetected. This creates uncertainty in how such results should be interpreted in editorial decisions.
An additional concern is the potential effect on writing behavior. Academic writing typically encourages clarity, conciseness, and structured presentation. However, some researchers may begin to adjust their writing style to avoid triggering AI-detection systems. This includes avoiding simplicity or refinement in language, even when such features improve readability. In such cases, the presence of detection tools may indirectly influence how researchers write, rather than simply evaluating their output.
Unlike similarity reports, which can be checked against visible sources, AI-detection outputs are harder to verify independently. This makes interpretation more dependent on trust in the system itself, increasing the importance of caution when using these tools in academic evaluation.
Impact on researchers
The growing use of similarity-detection and AI-detection tools affects researchers in different ways. While these systems are intended to support research integrity, their practical impact is often felt most directly at the level of individual writing and manuscript preparation. Early career researchers may be particularly affected. Many are still developing confidence in academic writing and may rely heavily on established phrases, templates, and commonly used scientific expressions. When such standard language is flagged by similar tools, it can create uncertainty about what is considered acceptable. Rewritings, this may lead to repeated rewriting of text that is already correct, simply to reduce similarity scores rather than to improve scientific content.
Researchers writing in a second language may face similar challenges. To ensure clarity and precision, they often depend on widely accepted academic structures and phrasing. However, because these expressions are frequently used across publications, they are more likely to appear in similar reports. This can create additional ambiguity, particularly when there is limited guidance on how such results should be interpreted in context.
AI-detection tools may add another layer of uncertainty. Since these systems do not provide transparent explanations for their decisions, researchers may find it difficult to understand why a particular text has been flagged. This can lead to hesitation in writing style, where authors may avoid certain forms of expression not because they are incorrect, but because they are unsure how the system might respond.
More broadly, the increasing visibility of these scores during the publication process can influence how researchers approach writing. Instead of focusing only on clarity, logic, and scientific contribution, some attention may shift toward how the text will be interpreted by automated systems. This does not necessarily change the scientific content, but it can affect the writing process itself and the confidence with which researchers present their work.
Towards a more balanced approach
The challenges discussed above do not suggest that similarity-detection and AI-detection tools are inherently problematic or unnecessary. On the contrary, these systems play an important role in supporting research integrity and assisting editors in managing the increasing volume of scholarly submissions. The key issue is not their presence in academic publishing, but the way their outputs are interpreted within editorial and writing practices. Although these tools typically include disclaimers noting that their results should not be used in isolation, their outputs are sometimes treated as decisive in editorial or assessment decisions.
A more balanced approach requires a shared understanding among journals, editors, and researchers that automated outputs function as screening signals rather than definitive measures of originality or authorship. Similarity scores, for example, reflect textual overlap and not academic misconduct in themselves. Their interpretation depends on context, including the section of the manuscript, disciplinary conventions, and the nature of the matched content. For instance, text overlaps occurring within Methods sections may warrant different consideration from overlaps identified in sections where originality of expression is more central. Recognizing this distinction can help ensure that standard scientific language and methodological descriptions are not misinterpreted as problematic text reuse.
In a similar way, AI-detection results need to be understood within the limits of the underlying technology. These systems provide probabilistic assessments rather than verifiable evidence of AI-generated writing. Their outputs can therefore vary depending on writing style and system design, and they may not consistently reflect actual authorship. A careful, context-based interpretation of such results is essential to avoid drawing conclusions that are not fully supported by the available evidence.
In practice, this means that editorial assessment benefits from combining automated reports with informed human judgment. When similarity reports indicate high overlap, the key question is not only the numerical value but also the nature and location of the overlap. Similarly, when AI-detection tools produce high scores, these should be viewed as prompts for closer examination rather than standalone indicators of concern. In both cases, the emphasis should remain on contextual evaluation rather than numerical thresholds.
The examples discussed above also point to potential directions for the continued evolution of plagiarism- and AI-detection technologies. Developers of plagiarism and AI-detection systems may consider improving transparency by providing section-specific analyses and confidence measures. For instance, text identified within Methods sections could be flagged differently from text appearing in the Introduction or Discussion, where originality of expression is generally expected. Similarly, AI-detection tools should report uncertainty estimates and explain which linguistic features contributed most strongly to classification, allowing users to make informed judgments rather than relying on a binary output.
Greater clarity in how these systems are understood across the academic community would also be beneficial. Many researchers, particularly those early in their careers, interpret similarity and AI-detection outputs as measures of writing quality or integrity, which can lead to unnecessary revisions or uncertainty. A more consistent understanding of these tools as preliminary screening mechanisms could help reduce such misinterpretations and support more confident academic writing.
Conclusion
Originality in academic work extends beyond textual similarity or algorithmic classification. It is reflected in the formulation of ideas, the design and execution of research, and the interpretation of findings. Automated tools can contribute valuable signals in assessing written work, but they cannot fully capture these deeper dimensions of scholarship. For this reason, human judgment remains central to ensuring that research is evaluated fairly and accurately.
Ethical approval:
Institutional Review Board approval is not required.
Declaration of patient consent:
Patient's consent is not required as there are no patients in this study.
Conflicts of interest:
There are no conflicts of interest.
Use of artificial intelligence (AI)-assisted technology for manuscript preparation:
The author confirms that there was no use of artificial intelligence (AI)-assisted technology for assisting in the writing or editing of the manuscript, and no images were manipulated using AI.
Financial support and sponsorship: Nil.
References
- Digital tools for academic writing: A systematic analysis of AI and technology-enhanced support. Türkiye Egitim Derg. 2025;10:275-97.
- [CrossRef] [Google Scholar]
- Handling an article produced by AI. 2025. COPE. Available from: https://publicationethics.org/guidance/case/handling-article-produced-ai [Last accessed 2026 Jun 23]
- [Google Scholar]
- Transparency mechanisms for generative AI use in higher education assessment: A systematic scoping review (2022-2026) Computers. 2026;15:111.
- [CrossRef] [Google Scholar]
- Text-based plagiarism in scientific publishing: Issues, developments and education. Sci Eng Ethics. 2013;19:1241-54.
- [CrossRef] [PubMed] [Google Scholar]
- Turnitin: Is it a text matching or plagiarism detection tool? Saudi J Anaesth. 2019;13:S48-51.
- [CrossRef] [PubMed] [Google Scholar]
- Comparing the similarity index across iThenticate, Ouriginal, and Turnitin plagiarism detection software. J Phys Educ Sport. 2024;24:419-4.
- [Google Scholar]
- AI-generated text detection: A comprehensive review of active and passive approaches. Comput Mater Contin. 2026;86:1-10.
- [CrossRef] [Google Scholar]
- Testing of detection tools for AI-generated text. Int J Educ Integr. 2023;19:26.
- [CrossRef] [Google Scholar]
- Evaluating the accuracy and reliability of AI content detectors in academic contexts. Int J Educ Integr. 2026;22:4.
- [CrossRef] [Google Scholar]
