benchmark-robustness-auditor
Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detection, option-letter selection bias (chi2), few-shot curve noise, LLM-judge position/verbosity/rubric-echo bias and hidden-in
as observed 2026-09-08T10:18:19.295Z- Identifier
benchmark-robustness-auditor- Source
- ClawHub
- Version observed
- 2.0.0
- Source repository
- not published
- Repository observation
- No source repository listed
- First observed here
- 2026-08-28T22:54:49.073Z
- Observations recorded
- 3
- Installs (reported upstream)
- 2
- Weekly downloads (upstream)
- 765
Observation history
2026-09-08T10:18:19.295Z
Fields that differed: topics
| Field | Before | After |
|---|---|---|
topics |
["benchmark-integrity","contamination-detection","judge-bias","llm-evaluation","self-improving"] | ["benchmark-integrity","contamination-detection","judge-bias","llm-evaluation","productivity"] |
2026-09-06T09:04:43.486Z
Fields that differed: changelog description latestVersion summary topics
| Field | Before | After |
|---|---|---|
changelog |
"README-only documentation: complete functionality, usage, permissions, security/privacy, related emojis, and corrected TREE-SHA256-v1 verification excluding registry-generated met | "Full functional rewrite: offline benchscan.py engine (12 subcommands: contam w/ C-1/C-2/C-3, selection, fewshot, judge E-1/E-2/E-3/T-3, compare McNemar+Wilson+bootstrap, tsguess G |
description |
"Red-team auditor for LLM benchmarks with executable scripts, severity calculator, mitigation library, date-contamination detector, evaluator injection harness, and new exploit typ | "Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detectio |
latestVersion |
"1.1.6" | "2.0.0" |
summary |
"Red-team auditor for LLM benchmarks with executable scripts, severity calculator, mitigation library, date-contamination detector, evaluator injection harness, and new exploit typ | "Offline, defensive robustness auditor for LLM benchmarks: n-gram exact + shingle-Jaccard paraphrase contamination, temporal pre/post-cutoff gaps, TS-Guessing above-chance detectio |
topics |
["auditing","llm-benchmarks","red-team","research","robustness"] | ["benchmark-integrity","contamination-detection","judge-bias","llm-evaluation","self-improving"] |
Correction
If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.