Acadwrite Lab · AI Detector Consistency Study

Why can the same academic passage differ by 100% across AI detectors?

Cross-platform test · 13 passages · 10 detection platforms

This experiment used 13 AI-generated academic passages across 10 detection platforms. When the same passage was submitted to multiple platforms, one could report nearly 100% while another reported 0%.

To evaluate a revision, test the before and after versions on the platform your institution will ultimately use.

Research summary: cross-platform results differ substantially

AI-content detectors use different models, decision thresholds, and text-segmentation methods. Although every platform reports a percentage, their calculation standards and the meaning of the numbers differ.

The study submitted each passage to multiple platforms and compared their results. Cross-platform agreement was low, with no stable relationship between third-party results and designated academic platforms such as CNKI and VIP.

100percentage-point maximum gap
Cross-platform percentages do not have a direct conversion

A drop from 80% to 20% on a third-party platform only shows a change on that platform. Any change on CNKI or VIP must be measured again on the corresponding platform.

Study design: 13 passages across 10 platforms

The study used 13 academic passages generated entirely by AI: 10 in Chinese and 3 in English. Each passage was submitted to multiple detection platforms, their reported AI-content rates were recorded, and the results for identical text were compared.

In the report appendix, some samples for CNKI, VIP, and Turnitin are marked as not tested. Therefore, 13 passages and 10 platforms describe the overall scope; the number of completed tests follows the actual appendix records.

How the experiment worked

01
Prepare 13 passages

10 Chinese and 3 English academic passages, all generated by AI

02
Submit them to multiple platforms

The experiment covered 10 detection platforms overall

03
Compare identical passages

Observe the direction and size of differences in reported percentages

Multiple samples spanned the full range from 0% to 100%

In the four extreme cases listed in the report, the lowest and highest readings for the same passage differed by 100 percentage points across platforms.

Test 1The same text was judged as either 0% or 100%
ZeroGPT · 0%
VIP · 100%
Test 8CNKI and VIP differed by 79 percentage points
ZeroGPT · 0%
VIP · 100%
Test 9CNKI and VIP differed by 95 percentage points
CNKI · 0%
QuillBot · 100%
Test 13The English passage showed the same extreme disagreement
Turnitin · 0%
ZeroGPT · 100%

Each line connects the lowest and highest readings for the same passage across platforms to show cross-platform disagreement.

Three main findings

For practical use, a reliable before-and-after comparison requires the same detection platform, version, and text range.

01

Changing platforms can completely change the number

The same version of the same passage can score nearly 100% on one platform and 0% on another. Cross-platform percentages use different standards and must be interpreted separately.

02

Third-party screening does not convert to a designated platform

If a passage is tested and revised using a third-party detector before being checked on CNKI or VIP, the change may reflect either the revision or simply different platform criteria.

03

Use the same platform to determine whether a revision worked

A more meaningful method is to test the same text range, on the same platform and version, before and after revision, then inspect the specific passages marked in the report.

Test 9: the same academic passage scored 100%, 95%, and 0% on three platforms

The original Test 9 passage was unchanged and submitted separately to QuillBot, VIP, and CNKI. Their results were 100%, 95%, and 0%. Merely changing the detector moved the same text across the entire range from almost entirely AI-generated to not detected.

Test 9 · Cross-platform results

Same text and version; data from the full research report

QuillBot
Third-party platform
100%
VIP
Chinese academic detection platform
95%
CNKI
Chinese academic detection platform
0%

All three numbers came from the same unedited passage; only the detection platform changed.

Conclusions and recommendations

Interpret results separately for each platform and use the specific passages marked in the report to identify problems and plan revisions.

Main conclusions

  • Platforms use different algorithms and thresholds, producing gaps of up to 100 percentage points.
  • CNKI and VIP differed by as much as 95 percentage points on the same text.
  • Self-check platforms do not correlate reliably with official platforms.
  • Cross-platform results are highly dispersed, and no industry technical benchmark has yet formed.

Recommendations

  • Use the reported percentage to locate expression issues, but judge paper quality from the content itself.
  • Use the designated platform ultimately adopted by the institution, such as CNKI or VIP, for retesting.
  • Return to the source text based on report markers and verify and revise it passage by passage.

How can before-and-after tests be comparable?

This comparison concerns the reported AI-content rate. To evaluate revisions, test the same text range before and after on the same platform and version. Third-party platforms can quickly reveal expression patterns, but the final result should follow the platform used by the institution.

  1. 01
    First confirm the final platform

    Find out whether the institution, journal, or organization requires CNKI, VIP, or another platform, and use it as the final retesting standard.

  2. 02
    Fix the platform and text range

    Retest the same text range before and after revision using the same platform and version, so numerical changes can be tied to edits.

  3. 03
    Read the report alongside marked passages

    The overall percentage provides an overview; specific revisions should return to marked passages to verify facts, terminology, citations, and sentence structure.

Already have an AI-content detection report?

Use the report to locate passages that need attention, revise them, and retest on the same detection platform.