TheBotique

eval-bench-builder

当用户说『怎么评测这个AI/技能好不好』『给我造个测试集』『评测用例怎么设计』『要可复现的评测』,或要给一个 agent/模型/技能建可复现评测基准时使用。从能力说明+边界用例生成结构化 eval 样本(输入/期望/判定标准),保证可复现、可回归。可运行脚本(bench_build 生成器)。理论根基:LGD 三律之有证(评测可复现、可核验)。触发词:评测基准、eval、测试集、benchmark、可复现评测、评测用例、怎么测AI。

as observed 2026-09-29T04:47:03.526Z
Identifier
eval-bench-builder
Source
ClawHub
Version observed
1.0.3
Source repository
not published
Repository observation
No source repository listed
First observed here
2026-09-12T14:17:54.398Z
Observations recorded
3
Installs (reported upstream)
1
Weekly downloads (upstream)
216
Declared license
MIT-0

Observation history

2026-09-29T04:47:03.526Z

Fields that differed: changelog latestVersion

FieldBeforeAfter
changelog "Eval Bench Builder v1.2.0 introduces benchmark automation for AI/agent evaluation.\n\n- Generates structured, reproducible evaluation samples from capability descriptions and edge "权利层第 3 轮:① §3.2 权属宣告统一块归一——清除块内他资产事实(此前误填 uibc-core 的 DOI / commit cb6f11b / Sigstore Rekor 锚),改按本资产真值填写;② 技能件许可口径过授权清除——理论文本由 CC BY 4.0 改为「保留所有权利」,.py 依 §3.2 标 MIT;③ 保留分层许可:代码 MI
latestVersion "1.0.0" "1.0.3"
2026-09-17T19:21:47.094Z

Fields that differed: license description

FieldBeforeAfter
license null "MIT-0"
description "当用户说『怎么评测这个AI/技能好不好』『给我造个测试集』『评测用例怎么设计』『要可复现的评测』,或要给一个 agent/模型/技能建可复现评测基准时使用。从能力说明+边界用例生成结构化 eval 样本(输入/期望/判定标准),保证可复现、可回归。可运行脚本(bench_build 生成器)。理论根基:LGD 三律之有证(评测可复现、可核验)。触发词:评测 null

Correction

If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.