Wordvice’s Large-Scale Study of False Positives in AI Detection Accepted to EMNLP 2026 Industry Track

ⓘ This article is third-party content and does not represent the views of this site. We make no guarantees regarding its accuracy or completeness.

Analysis of more than 135,000 pairs of original and edited academic manuscripts finds that professional human editing alone can change AI detection results

-- “AI detector scores alone cannot establish whether AI was used. Manuscript evaluation should consider the quality of the work and how it was written and edited.”


Wordvice, a global provider of proofreading and editing services and AI writing solutions, has announced a study examining the reliability of AI text detectors and their false-positive rates using more than 135,000 pairs of real-world academic manuscripts. The study was accepted to the EMNLP 2026 Industry Track, part of an international conference on natural language processing.


The study, “Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing,” analyzed 135,389 pairs of academic manuscripts, collected between 2018 and 2025, before and after professional editing. By comparing original and edited versions with the same authorship and research content, the researchers examined how changes in grammar, fluency, and academic style affect AI detection results.


The researchers evaluated 13 AI text detectors using various approaches, including methods based on token statistics, zero-shot detection, and classifiers. They found that false-positive rates for the same set of human-written academic manuscripts varied widely across detection methods, ranging from 0% to 100%.


In particular, AI detection results changed significantly after these manuscripts underwent professional human editing. For several major classifier-based detectors, including MAGE and RADAR, professional editing reduced the false-positive rate at which human-written text was classified as AI-generated. However, other detection methods exhibited the opposite pattern. In addition, more extensive editing was associated with greater changes in AI detection scores.


According to the researchers, these findings suggest that AI detectors can be influenced not only by who or what produced a text but also by various characteristics, such as writing style and linguistic quality. Because legitimate language improvements, such as professional English editing routinely used in academic publishing, can change detection results, the findings call for caution in treating AI detector scores as definitive evidence of AI use.


“An AI detector score reflects how a particular detection method evaluates a text. It cannot, on its own, determine how that text was actually written,” a Wordvice spokesperson stated. “Rather than revising their writing to lower AI detector scores, researchers should focus on improving the clarity, accuracy, academic quality, and publication readiness of their manuscripts.”


The study was conducted during Hyeonchu Park’s internship at Wordvice. Park is a researcher at Chung-Ang University’s Graduate School of Artificial Intelligence, supervised by Professor Bugeun Kim. The research provides empirical evidence to inform the interpretation of AI detector results, particularly in academic writing by nonnative English-speaking researchers. Previous studies have compared different groups of writers, whereas this study directly compared the same academic manuscripts.


About Wordvice

Wordvice is a global education technology company built on academic English editing services provided by native English-speaking editors with expertise across numerous research disciplines. In addition to professional editing, the company offers a range of AI writing solutions for researchers, including a grammar checker, translation, paraphrasing, summarization, and AI detection. Drawing on its expertise in professional human editing and AI technology, Wordvice helps researchers improve the quality of their academic writing and more efficiently prepare manuscripts for publication, while also studying the reliability and fairness of AI writing and detection technology.


The preprint is available on arXiv.


Paper title: “Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing”


Accepted to EMNLP 2026 Industry Track


Contact Info:
Name: Gahye Jeong
Email: Send Email
Organization: Wordvice
Website: https://wordvice.ai

Release ID: 89203597

If there are any errors, inconsistencies, or queries arising from the content contained within this press release that require attention or if you need assistance with a press release takedown, we kindly request that you inform us immediately by contacting error@releasecontact.com (it is important to note that this email is the authorized channel for such matters, sending multiple emails to multiple addresses does not necessarily help expedite your request). Our reliable team will be available to promptly respond within 8 hours, taking proactive measures to rectify any identified issues or providing guidance on the removal process. Ensuring accurate and dependable information is our top priority.

Report this content

If you believe this article contains misleading, harmful, or spam content, please let us know.

Report this article

More News

View More

Recent Quotes

View More
Symbol Price Change (%)
AMZN  251.19
+5.23 (2.13%)
AAPL  337.00
+4.59 (1.38%)
AMD  545.09
+32.59 (6.36%)
BAC  58.18
+0.28 (0.48%)
GOOG  343.68
+4.32 (1.27%)
META  682.31
+9.00 (1.34%)
MSFT  497.75
+7.45 (1.52%)
NVDA  219.34
+5.44 (2.54%)
ORCL  150.59
+7.43 (5.19%)
TSLA  366.20
+8.12 (2.27%)
Stock Quote API & Stock News API supplied by www.cloudquote.io
Quotes delayed at least 20 minutes.
By accessing this page, you agree to the Privacy Policy and Terms Of Service.