Humanit Benchmark - conducted and published by Humanit.
Version 1 · Humanit Benchmark Scoring Model v1
Detector outcomes are observational only and must not determine the overall winner.
At least 3 runs per selected item where affordable. Mean, median and variability disclosed. No cherry-picking.
Failed runs are recorded, not excluded. Failure rate shown where meaningful.
Blinded human review is performed. Humanit-staff review does not count as independent verification.
AI detectors can produce false positives and false negatives. Results should not be used alone to determine authorship or misconduct.
Equivalent settings are used where possible: plan, mode, tone, language, input length, temperature (where accessible), credits consumed, date, product version and processing time are all recorded.
All trademarks belong to their respective owners. Humanit is not affiliated with the compared providers unless expressly stated.
How can we help?