# 图片处理完整流程分析 ## 流程概览 ``` ┌─────────────────────────────────────────────────────────────────────────────────┐ │ 图片处理完整流程 │ ├─────────────────────────────────────────────────────────────────────────────────┤ │ │ │ 1. 解析阶段 (MinerU) │ │ ┌─────────────┐ ┌──────────────────┐ ┌───────────────────────────┐ │ │ │ PDF/Word │────→│ MinerU 解析 │────→│ MinerUChunk 对象 │ │ │ │ 文件 │ │ parsers/mineru_ │ │ ├── content (文本/标题) │ │ │ └─────────────┘ │ parser.py │ │ ├── chunk_type │ │ │ └──────────────────┘ │ ├── table_html (表格HTML) │ │ │ │ ├── image_path (独立图片) │ │ │ │ └── images (关联图片列表) │ │ │ ↓ │ │ to_page_content() │ │ ↓ │ │ 返回 chunks 列表 │ │ │ ├─────────────────────────────────────────────────────────────────────────────────┤ │ │ │ 2. 存储阶段 (Knowledge Manager) │ │ ┌──────────────────┐ ┌────────────────────┐ ┌───────────────────┐ │ │ │ chunks 列表 │────→│ add_file_to_kb() │────→│ 向量库 metadata │ │ │ │ (MinerUChunk) │ │ knowledge/manager │ │ │ │ │ └──────────────────┘ │ .py │ │ ├── chunk_type │ │ │ │ │ │ ├── source │ │ │ │ ✅ 合并跨页表格 │ │ ├── page │ │ │ │ ✅ 序列化 images │ │ ├── images_json ✅│ │ │ │ ✅ 存储 image_path │ │ └── image_path ✅ │ │ │ └────────────────────┘ └───────────────────┘ │ │ │ ├─────────────────────────────────────────────────────────────────────────────────┤ │ │ │ 3. 召回阶段 (RAG 检索) │ │ ┌──────────────────┐ ┌────────────────────┐ ┌───────────────────┐ │ │ │ 用户查询 │────→│ 混合检索 │────→│ 检索结果 │ │ │ │ "表3.1 数据" │ │ search_hybrid() │ │ contexts = [{ │ │ │ └──────────────────┘ │ api/chat_routes.py │ │ "doc": "...", │ │ │ └────────────────────┘ │ "meta": {...} │ │ │ ↓ │ }] │ │ │ ↓ └───────────────────┘ │ │ ┌────────────────────┐ ↓ │ │ │ _extract_rich_media│ ┌───────────────────┐ │ │ │ api/chat_routes.py │ │ 返回给前端 │ │ │ │ │────→│ { │ │ │ │ 读取 images_json │ │ "images": [...],│ │ │ │ 读取 image_path │ │ "tables": [...],│ │ │ │ ✅ 正确处理 │ │ "answer": "..." │ │ │ └────────────────────┘ │ } │ │ │ └───────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────────────────┘ ``` ## 关键代码位置 ### 1. MinerU 解析 (parsers/mineru_parser.py) **MinerUChunk 数据结构 (第107-125行)**: ```python @dataclass class MinerUChunk: content: str # 文本内容 chunk_type: str # 类型: text, table, image, equation page_start: int = 1 page_end: int = 1 title: str = "" table_html: Optional[str] = None # 表格 HTML(如果是表格) image_path: Optional[str] = None # 图片路径(独立图片) images: Optional[List[Dict]] = None # 关联图片列表: [{"id": "abc.jpg", "order": 1}] ``` **表格切片创建 (第453-469行)**: ```python chunk = MinerUChunk( content=table_caption or "表格", # ⚠️ content 只有标题 chunk_type="table", table_html=table_body, # ✅ 完整表格在 table_html image_path=img_path, # 表格图片路径 images=table_images # 表格中嵌入的图片 ) ``` ### 2. 向量库存储 (knowledge/manager.py) **add_file_to_kb() 核心逻辑 (第228-270行)**: ```python for i, chunk in enumerate(chunks): # 1. 获取 chunk_type chunk_type = getattr(chunk, 'chunk_type', None) if not chunk_type: page_info = getattr(chunk, 'page_info', {}) or {} chunk_type = page_info.get('chunk_type', 'text') if chunk_type == 'table': # 2. 表格内容优先使用 table_html table_md = getattr(chunk, 'table_html', None) or chunk.content semantic_content = _build_semantic_content_for_table(...) # 3. 构建 metadata metadata = { 'chunk_type': chunk_type, 'source': filename, 'page': page_start, # ... } # 4. ✅ 序列化图片信息 if hasattr(chunk, 'images') and chunk.images: metadata['images_json'] = json.dumps(chunk.images, ensure_ascii=False) if hasattr(chunk, 'image_path') and chunk.image_path: metadata['image_path'] = chunk.image_path ``` **跨页表格合并 (第323-420行)**: ```python def _merge_cross_page_tables(self, chunks: list) -> list: """ 合并规则: 1. 相邻两个表格切片 2. 页码连续 (page_end + 1 == next.page_start) 3. 第二个表格标题包含"续表" """ # 合并 table_html current.table_html = curr_html + '\n' + next_html # 合并 image_path 到 images merged_images = [ {'id': curr_img, 'page': curr_page_end}, {'id': next_img, 'page': next_page_start} ] current.images = merged_images ``` ### 3. 富媒体召回 (api/chat_routes.py) **_extract_rich_media() 核心逻辑 (第724-843行)**: ```python def _extract_rich_media(contexts: List[Dict]) -> Dict[str, List]: images = [] tables = [] for ctx in contexts: meta = ctx.get("meta", {}) # 1. 独立图片切片 (image_path) - 图片/图表类型 if meta.get("chunk_type") in ("image", "chart") and meta.get("image_path"): img_id = os.path.basename(meta["image_path"]) images.append({"id": img_id, "url": f"/images/{img_id}", ...}) # 2. 关联图片 (images_json) - 表格/文本嵌入图片 if meta.get("images_json"): img_list = json.loads(meta["images_json"]) for img_info in img_list: images.append({"id": img_info["id"], ...}) # 3. 表格图片 (image_path) - 表格类型的图片形式 if meta.get("chunk_type") == "table" and meta.get("image_path"): img_id = os.path.basename(meta["image_path"]) images.append({"id": img_id, "type": "table_image", ...}) return {"images": images, "tables": tables} ``` ## 当前问题分析 ### 问题1: 表格显示"0行数据" **根因**: `_build_semantic_content_for_table()` 接收的 `table_md` 可能是空的 **验证点**: - MinerU 解析时 `table_html` 是否有值? - `manager.py` 第240行 `table_md = getattr(chunk, 'table_html', None)` 是否正确获取? ### 问题2: 跨页表格合并不生效 **根因**: 可能是页码不连续或标题匹配失败 **验证点**: - 检查 `_merge_cross_page_tables()` 的日志输出 - 验证两个表格切片的 `page_end` 和 `page_start` 是否连续 ### 问题3: 图片重复 **根因**: 可能是 `images_json` 和 `image_path` 同时存在导致重复 **验证点**: - 检查向量库中是否有同时存在 `images_json` 和 `image_path` 的切片 - `_extract_rich_media()` 中的去重逻辑是否有效 ## 数据存储位置 | 目录 | 用途 | |------|------| | `.data/images/` | 全局图片存储(哈希命名,去重) | | `.data/cache/vlm/` | VLM 图片描述缓存 | | `.data/docstore/` | 原始表格/图片 JSON 备份 | | `knowledge/vector_store/chroma/` | ChromaDB 向量数据库 | ## 测试验证步骤 ### 1. 检查向量库 metadata ```python from knowledge.manager import get_kb_manager kb = get_kb_manager() coll = kb.get_collection('my_ky') result = coll.get(limit=10, include=['metadatas']) for meta in result['metadatas']: print(f"chunk_type: {meta.get('chunk_type')}") print(f"images_json: {meta.get('images_json')}") print(f"image_path: {meta.get('image_path')}") print("---") ``` ### 2. 检查 MinerU 解析结果 ```python from parsers.mineru_parser import parse_with_mineru result = parse_with_mineru("tests/public/test_report.pdf") for chunk in result.get('chunks', []): if chunk.chunk_type == 'table': print(f"表格标题: {chunk.title}") print(f"table_html 长度: {len(chunk.table_html or '')}") print(f"image_path: {chunk.image_path}") print(f"images: {chunk.images}") print("---") ```