百度文档解析pipeline-parser
调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。
as observed 2026-09-17T19:21:46.869Z- Identifier
baidu-doc-pipeline-parser- Source
- ClawHub
- Version observed
- 1.0.8
- Source repository
- not published
- Repository observation
- No source repository listed
- First observed here
- 2026-08-28T22:55:14.708Z
- Observations recorded
- 2
- Installs (reported upstream)
- 15
- Weekly downloads (upstream)
- 1,266
- Declared license
- MIT-0
Observation history
2026-09-17T19:21:46.869Z
Fields that differed: license changelog description latestVersion
| Field | Before | After |
|---|---|---|
license |
null | "MIT-0" |
changelog |
"baidu-doc-pipeline-parser 1.0.7\n\n- 切换为标准文档解析(pipeline-parser)Skill版本,替换原多模态VLM实现\n- 新增 scripts/baidu_doc_parser.py,移除 scripts/baidu_doc_vlm_parser.py\n- 支持18+文档格式及RAG场景分块,适用PDF、 | "- 移除 skill-card.md 文件。\n- SKILL.md 文档中,调整了免费额度表:企业实名认证用户额度由 1000 页改为 200 页。\n- 页面对象解析字段及部分类型补充、细化(如 page_num、text 字段描述、type/版面类型等)。\n- 其他内容未变。" |
description |
"调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。" | null |
latestVersion |
"1.0.7" | "1.0.8" |
Correction
If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.