eval-bench-builder
当用户说『怎么评测这个AI/技能好不好』『给我造个测试集』『评测用例怎么设计』『要可复现的评测』,或要给一个 agent/模型/技能建可复现评测基准时使用。从能力说明+边界用例生成结构化 eval 样本(输入/期望/判定标准),保证可复现、可回归。可运行脚本(bench_build 生成器)。理论根基:LGD 三律之有证(评测可复现、可核验)。触发词:评测基准、eval、测试集、benchmark、可复现评测、评测用例、怎么测AI。
as observed 2026-09-29T04:47:03.526Z- Identifier
eval-bench-builder- Source
- ClawHub
- Version observed
- 1.0.3
- Source repository
- not published
- Repository observation
- No source repository listed
- First observed here
- 2026-09-12T14:17:54.398Z
- Observations recorded
- 3
- Installs (reported upstream)
- 1
- Weekly downloads (upstream)
- 216
- Declared license
- MIT-0
Observation history
2026-09-29T04:47:03.526Z
Fields that differed: changelog latestVersion
| Field | Before | After |
|---|---|---|
changelog |
"Eval Bench Builder v1.2.0 introduces benchmark automation for AI/agent evaluation.\n\n- Generates structured, reproducible evaluation samples from capability descriptions and edge | "权利层第 3 轮:① §3.2 权属宣告统一块归一——清除块内他资产事实(此前误填 uibc-core 的 DOI / commit cb6f11b / Sigstore Rekor 锚),改按本资产真值填写;② 技能件许可口径过授权清除——理论文本由 CC BY 4.0 改为「保留所有权利」,.py 依 §3.2 标 MIT;③ 保留分层许可:代码 MI |
latestVersion |
"1.0.0" | "1.0.3" |
2026-09-17T19:21:47.094Z
Fields that differed: license description
| Field | Before | After |
|---|---|---|
license |
null | "MIT-0" |
description |
"当用户说『怎么评测这个AI/技能好不好』『给我造个测试集』『评测用例怎么设计』『要可复现的评测』,或要给一个 agent/模型/技能建可复现评测基准时使用。从能力说明+边界用例生成结构化 eval 样本(输入/期望/判定标准),保证可复现、可回归。可运行脚本(bench_build 生成器)。理论根基:LGD 三律之有证(评测可复现、可核验)。触发词:评测 | null |
Correction
If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.