TheBotique

rag-eval-harness

RAG 问答系统的质量评估与防幻觉验证方法论,用于回答「知识库答得准不准」「它会不会胡编」「换了模型还稳不稳」这类问题

as observed 2026-09-01T06:16:18.534Z

Want to know when this changes? Sign in with an email address and we will tell you. Free, no card, watch up to 5.

Identifier
rag-eval-harness
Source
ClawHub
Version observed
1.1.2
Source repository
not published
Repository observation
No source repository listed
First observed here
2026-08-28T22:06:28.212Z
Observations recorded
2
Installs (reported upstream)
0
Weekly downloads (upstream)
133

Observation history

2026-09-01T06:16:18.534Z

Fields that differed: changelog description displayName latestVersion summary

FieldBeforeAfter
changelog "v1.1.0 首发:RAG 问答系统质量评估与防幻觉验证——正向命中+负面拒答双测试,拒答阈值+引用兜底堵幻觉;本地零积分检索评测与阈值扫描选优(实测 50 题合格率 92%→100%)" "- Removed the file skill-card.md.\n- No other functional or content changes in this version."
description "RAG 问答系统的质量评估与防幻觉验证方法论。当用户需要对基于本地知识库的 AI 问答系统(RAG)做上线前验收、语料变更后回归、模型更换后复测,或想量化'知识库答得好不好''会不会胡编'时使用。覆盖:正向评估集(区域×主题矩阵+关键词命中)、负面测试集(知识库外问题+拒答信号表)、防幻觉三机制(低相似度拒答/引用 "RAG 问答系统的质量评估与防幻觉验证方法论,用于回答「知识库答得准不准」「它会不会胡编」「换了模型还稳不稳」这类问题"
displayName "RAG 评估与防幻觉验证" "rag-eval-harness"
latestVersion "1.1.0" "1.1.2"
summary "RAG 问答系统的质量评估与防幻觉验证方法论。当用户需要对基于本地知识库的 AI 问答系统(RAG)做上线前验收、语料变更后回归、模型更换后复测,或想量化'知识库答得好不好''会不会胡编'时使用。覆盖:正向评估集(区域×主题矩阵+关键词命中)、负面测试集(知识库外问题+拒答信号表)、防幻觉三机制(低相似度拒答/引用 "RAG 问答系统的质量评估与防幻觉验证方法论,用于回答「知识库答得准不准」「它会不会胡编」「换了模型还稳不稳」这类问题"

Correction

If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.