TheBotique

百度文档解析pipeline-parser

调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。

as observed 2026-09-17T19:21:46.869Z
Identifier
baidu-doc-pipeline-parser
Source
ClawHub
Version observed
1.0.8
Source repository
not published
Repository observation
No source repository listed
First observed here
2026-08-28T22:55:14.708Z
Observations recorded
2
Installs (reported upstream)
15
Weekly downloads (upstream)
1,266
Declared license
MIT-0

Observation history

2026-09-17T19:21:46.869Z

Fields that differed: license changelog description latestVersion

FieldBeforeAfter
license null "MIT-0"
changelog "baidu-doc-pipeline-parser 1.0.7\n\n- 切换为标准文档解析(pipeline-parser)Skill版本,替换原多模态VLM实现\n- 新增 scripts/baidu_doc_parser.py,移除 scripts/baidu_doc_vlm_parser.py\n- 支持18+文档格式及RAG场景分块,适用PDF、 "- 移除 skill-card.md 文件。\n- SKILL.md 文档中,调整了免费额度表:企业实名认证用户额度由 1000 页改为 200 页。\n- 页面对象解析字段及部分类型补充、细化(如 page_num、text 字段描述、type/版面类型等)。\n- 其他内容未变。"
description "调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。" null
latestVersion "1.0.7" "1.0.8"

Correction

If you maintain this extension and believe anything above is inaccurate, request a correction. Corrections are published, and disputed entries are marked as disputed while under review.