TinkerClaw Memory Bench
Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against your live memory system. Runs entirely LOCALLY by default — no memory content leaves your machine, and excerpts are re
as observed 2026-09-09T11:17:27.097Z- Identifier
memory-bench-pioneer- Source
- ClawHub
- Version observed
- 2.1.3
- Source repository
- not published
- Repository observation
- No source repository listed
- First observed here
- 2026-09-08T10:18:19.295Z
- Observations recorded
- 2
- Installs (reported upstream)
- 24
- Weekly downloads (upstream)
- 902
Observation history
2026-09-09T11:17:27.097Z
Fields that differed: changelog description latestVersion summary
| Field | Before | After |
|---|---|---|
changelog |
"Fixes two real defects the audit found, not just wording. submit.sh interpolated the report path directly into python3 -c source in five places, so a crafted report filename could | "Exact request-body preview before OpenAI consent; schema-validated public report; no unattended bypass." |
description |
"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against y | "Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against y |
latestVersion |
"2.1.2" | "2.1.3" |
summary |
"Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against y | "Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Peer-review-grade evaluation (LLM-as-judge, nDCG/MAP/MRR with 95% CIs, ablations) against y |
Correction
If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.