TheBotique

benchmark-robustness-auditor

Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detection, option-letter selection bias (chi2), few-shot curve noise, LLM-judge position/verbosity/rubric-echo bias and hidden-in

as observed 2026-09-08T10:18:19.295Z
Identifier
benchmark-robustness-auditor
Source
ClawHub
Version observed
2.0.0
Source repository
not published
Repository observation
No source repository listed
First observed here
2026-08-28T22:54:49.073Z
Observations recorded
3
Installs (reported upstream)
2
Weekly downloads (upstream)
765

Observation history

2026-09-08T10:18:19.295Z

Fields that differed: topics

FieldBeforeAfter
topics ["benchmark-integrity","contamination-detection","judge-bias","llm-evaluation","self-improving"] ["benchmark-integrity","contamination-detection","judge-bias","llm-evaluation","productivity"]
2026-09-06T09:04:43.486Z

Fields that differed: changelog description latestVersion summary topics

FieldBeforeAfter
changelog "README-only documentation: complete functionality, usage, permissions, security/privacy, related emojis, and corrected TREE-SHA256-v1 verification excluding registry-generated met "Full functional rewrite: offline benchscan.py engine (12 subcommands: contam w/ C-1/C-2/C-3, selection, fewshot, judge E-1/E-2/E-3/T-3, compare McNemar+Wilson+bootstrap, tsguess G
description "Red-team auditor for LLM benchmarks with executable scripts, severity calculator, mitigation library, date-contamination detector, evaluator injection harness, and new exploit typ "Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detectio
latestVersion "1.1.6" "2.0.0"
summary "Red-team auditor for LLM benchmarks with executable scripts, severity calculator, mitigation library, date-contamination detector, evaluator injection harness, and new exploit typ "Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detectio
topics ["auditing","llm-benchmarks","red-team","research","robustness"] ["benchmark-integrity","contamination-detection","judge-bias","llm-evaluation","self-improving"]

Correction

If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.