evaluate-skill
Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.
as observed 2026-09-26T01:46:07.395Z- Identifier
evaluate-skill- Source
- ClawHub
- Version observed
- 1.0.14
- Source repository
- not published
- Repository observation
- No source repository listed
- First observed here
- 2026-08-29T23:02:09.653Z
- Observations recorded
- 2
- Installs (reported upstream)
- 0
- Weekly downloads (upstream)
- 1,062
- Declared license
- MIT-0
Observation history
2026-09-26T01:46:07.395Z
Fields that differed: license changelog description latestVersion
| Field | Before | After |
|---|---|---|
license |
null | "MIT-0" |
changelog |
"Improvements on the Caliper underlying CLI focused on reliability and usability (performance and retries)" | "### Changed\n- Clean-up from the 1.0.13 that added way too much files.\n- Built around four questions, each with its own place to fix: does the skill fire (`description`), does it |
description |
"Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or i | null |
latestVersion |
"1.0.12" | "1.0.14" |
Correction
If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.