Case study 24 / 26
AI Slop Detector
An end-to-end AI-writing detector: scraped data, synthetic generation, a LoRA-tuned BERT classifier, and an honest benchmark.
Collaborative project — the dataset and model weights are published on Hugging Face by a co-author.
- Status
- Research
- Period
- 2026
- Domain
- ai
- Language
- Python
- Last push
- 01 APR 2026
- License
- MIT
- Source of claims
- README.md, training/src/train.py, raid/src/inference.py, run_detector.py
Human paragraphs scraped from Wikipedia are paired with GPT-5 Nano rewrites to build a balanced dataset; a LoRA adapter on bert-base-cased learns to separate them, and the model is evaluated on the external RAID benchmark to measure where it generalises — and where it does not.
01/The problem
Detecting machine-generated text needs paired data that differs only in authorship. Most public detectors are black boxes with unpublished limits.
02/The system
Human paragraphs scraped from Wikipedia are paired with GPT-5 Nano rewrites to build a balanced dataset; a LoRA adapter on bert-base-cased learns to separate them, and the model is evaluated on the external RAID benchmark to measure where it generalises — and where it does not.
- Input to Tokenizer
- Tokenizer to BERT + LoRA
- BERT + LoRA to Softmax
- Softmax to Result
03/Implementation
- 01Crawled Wikipedia and stored two random paragraphs per page — 10,001 human-written paragraphs.
- 02Generated AI counterparts in two passes with separate contexts (summarise, then rewrite) using GPT-5 Nano through the Batch API.
- 03Fine-tuned bert-base-cased with a LoRA adapter (r = 16, rank-stabilised, all linear layers) and a classification head; 95/5 train/validation split, 5 epochs, fp16, gradient accumulation of 8.
- 04Inference returns P(AI) from a two-class softmax; text above 0.5 is labelled AI-generated. The benchmark path scores each paragraph separately.
- 05Evaluated on the RAID benchmark training set by domain and generator family, reporting TPR at 1% FPR.
04/Engineering
Two-pass generation
Summarising and rewriting in separate contexts stops the model from paraphrasing the source sentence by sentence, which would make the task artificially easy.
Parameter-efficient tuning
LoRA trains a small adapter instead of the full network, so the published artefact is light and the base model stays untouched.
Measure transfer, not just accuracy
In-distribution validation accuracy is reported alongside an external benchmark, so the model's limits are explicit.
05/Interface
No product screenshots are published for this project. The visual above is a code-driven representation of how it behaves, built from the repository source — not a screenshot.
06/Tech stack
- Python
- PyTorch
- Hugging Face Transformers
- PEFT / LoRA
- BERT
- OpenAI Batch API
- uv
07/Result
Verified outcomes
- 96% validation accuracy on the held-out split (as reported in the repository).
- On RAID, better than random across all generator families and most domains.
- TPR at 1% FPR: about 90% on reviews; 30–40% on Wikipedia, arXiv abstracts and news; 1–20% on books, recipes, poetry and Reddit.
Known limitations
- Trained only on Wikipedia-style text generated by a single model.
- 512-token context, one paragraph at a time.
08/Links
Next case study
ML Trading Bot →