Skip to main content

Methodology

Humanit Benchmark - conducted and published by Humanit.

Scoring model

Version 1 · Humanit Benchmark Scoring Model v1

  • meaning preservation20%
  • writing quality20%
  • naturalness15%
  • factual consistency10%
  • tone and voice10%
  • readability10%
  • user experience5%
  • platform capability5%
  • privacy and transparency3%
  • detector analysis2%
  • Total100%

Detector outcomes are observational only and must not determine the overall winner.

Test categories

  • Writing quality
  • Meaning preservation
  • Naturalness
  • Tone and voice
  • Factual consistency
  • Readability
  • Detector analysis (observational)
  • Platform capability
  • User experience
  • Cost and value

Repetition and variance

At least 3 runs per selected item where affordable. Mean, median and variability disclosed. No cherry-picking.

Failure handling

Failed runs are recorded, not excluded. Failure rate shown where meaningful.

Human review

Blinded human review is performed. Humanit-staff review does not count as independent verification.

Detector safeguards

AI detectors can produce false positives and false negatives. Results should not be used alone to determine authorship or misconduct.

  • No guaranteed-undetectable claims
  • Detector outcomes are observational only
  • Detector weight is minor (2%)

Fair test settings

Equivalent settings are used where possible: plan, mode, tone, language, input length, temperature (where accessible), credits consumed, date, product version and processing time are all recorded.

FREE PLAN COMPARISONENTRY PAID PLAN COMPARISONBEST AVAILABLE MODEFEATURE SPECIFIC COMPARISON

All trademarks belong to their respective owners. Humanit is not affiliated with the compared providers unless expressly stated.