Abstract
We find that large language models (LLMs) are more likely to modify human-written text than AI-generated text when tasked with rewriting. This tendency arises because LLMs often perceive AI-generated text as high-quality, leading to fewer modifications. We introduce a method to detect AI-generated content by prompting LLMs to rewrite text and calculating the editing distance of the output. We dubbed our geneRative AI Detection viA Rewriting method Raidar. Raidar significantly improves the F1 detection scores of existing AI content detection models – both academic and commercial – across various domains, including News, creative writing, student essays, code, Yelp reviews, and arXiv papers, with gains of up to 29 points. Operating solely on word symbols without high-dimensional features, our method is compatible with black box LLMs, and is inherently robust on new content. Our results illustrate the unique imprint of machine-generated text through the lens of the machines themselves.
Resources
Coverage
- Science News Explores A New Test Could Help Weed Out AI-Generated Text
- Columbia Engineering Who Wrote This? Columbia Engineers Discover Novel Method to Identify AI-Generated Text
- Columbia Magazine A New Way to Spot Text Written by AI, and Other Science News
- Columbia Engineering Turns Out, I’m Not Real: Detecting AI-Generated Videos
- 36Kr SORA、Gen-2、Pika也逃不过,文生视频检测新工具来了,准确率高达93.7%
- Science News Explores Google Now Adds Watermarks to All Its AI-Generated Content
- Tech Xplore Who Wrote This? Engineers Discover Novel Method to Identify AI-Generated Text
- Tech Xplore New Tool Detects AI-Generated Videos with 93.7% Accuracy