TheBotique

evaluate-skill

Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.

as observed 2026-09-26T01:46:07.395Z
Identifier
evaluate-skill
Source
ClawHub
Version observed
1.0.14
Source repository
not published
Repository observation
No source repository listed
First observed here
2026-08-29T23:02:09.653Z
Observations recorded
2
Installs (reported upstream)
0
Weekly downloads (upstream)
1,062
Declared license
MIT-0

Observation history

2026-09-26T01:46:07.395Z

Fields that differed: license changelog description latestVersion

FieldBeforeAfter
license null "MIT-0"
changelog "Improvements on the Caliper underlying CLI focused on reliability and usability (performance and retries)" "### Changed\n- Clean-up from the 1.0.13 that added way too much files.\n- Built around four questions, each with its own place to fix: does the skill fire (`description`), does it
description "Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or i null
latestVersion "1.0.12" "1.0.14"

Correction

If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.