init: RAG 知识库服务初始提交
- 后端 API(Flask + Gunicorn) - RAG 引擎(混合检索 + 云端 Reranker + 引用溯源) - 文档解析(MinerU + 多格式支持) - Docker 生产部署配置 - 排除前端项目、敏感配置、模型文件
This commit is contained in:
233
docs/image_processing_flow.md
Normal file
233
docs/image_processing_flow.md
Normal file
@@ -0,0 +1,233 @@
|
||||
# 图片处理完整流程分析
|
||||
|
||||
## 流程概览
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────────────────────┐
|
||||
│ 图片处理完整流程 │
|
||||
├─────────────────────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ 1. 解析阶段 (MinerU) │
|
||||
│ ┌─────────────┐ ┌──────────────────┐ ┌───────────────────────────┐ │
|
||||
│ │ PDF/Word │────→│ MinerU 解析 │────→│ MinerUChunk 对象 │ │
|
||||
│ │ 文件 │ │ parsers/mineru_ │ │ ├── content (文本/标题) │ │
|
||||
│ └─────────────┘ │ parser.py │ │ ├── chunk_type │ │
|
||||
│ └──────────────────┘ │ ├── table_html (表格HTML) │ │
|
||||
│ │ ├── image_path (独立图片) │ │
|
||||
│ │ └── images (关联图片列表) │ │
|
||||
│ ↓ │
|
||||
│ to_page_content() │
|
||||
│ ↓ │
|
||||
│ 返回 chunks 列表 │
|
||||
│ │
|
||||
├─────────────────────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ 2. 存储阶段 (Knowledge Manager) │
|
||||
│ ┌──────────────────┐ ┌────────────────────┐ ┌───────────────────┐ │
|
||||
│ │ chunks 列表 │────→│ add_file_to_kb() │────→│ 向量库 metadata │ │
|
||||
│ │ (MinerUChunk) │ │ knowledge/manager │ │ │ │
|
||||
│ └──────────────────┘ │ .py │ │ ├── chunk_type │ │
|
||||
│ │ │ │ ├── source │ │
|
||||
│ │ ✅ 合并跨页表格 │ │ ├── page │ │
|
||||
│ │ ✅ 序列化 images │ │ ├── images_json ✅│ │
|
||||
│ │ ✅ 存储 image_path │ │ └── image_path ✅ │ │
|
||||
│ └────────────────────┘ └───────────────────┘ │
|
||||
│ │
|
||||
├─────────────────────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ 3. 召回阶段 (RAG 检索) │
|
||||
│ ┌──────────────────┐ ┌────────────────────┐ ┌───────────────────┐ │
|
||||
│ │ 用户查询 │────→│ 混合检索 │────→│ 检索结果 │ │
|
||||
│ │ "表3.1 数据" │ │ search_hybrid() │ │ contexts = [{ │ │
|
||||
│ └──────────────────┘ │ api/chat_routes.py │ │ "doc": "...", │ │
|
||||
│ └────────────────────┘ │ "meta": {...} │ │
|
||||
│ ↓ │ }] │ │
|
||||
│ ↓ └───────────────────┘ │
|
||||
│ ┌────────────────────┐ ↓ │
|
||||
│ │ _extract_rich_media│ ┌───────────────────┐ │
|
||||
│ │ api/chat_routes.py │ │ 返回给前端 │ │
|
||||
│ │ │────→│ { │ │
|
||||
│ │ 读取 images_json │ │ "images": [...],│ │
|
||||
│ │ 读取 image_path │ │ "tables": [...],│ │
|
||||
│ │ ✅ 正确处理 │ │ "answer": "..." │ │
|
||||
│ └────────────────────┘ │ } │ │
|
||||
│ └───────────────────┘ │
|
||||
│ │
|
||||
└─────────────────────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## 关键代码位置
|
||||
|
||||
### 1. MinerU 解析 (parsers/mineru_parser.py)
|
||||
|
||||
**MinerUChunk 数据结构 (第107-125行)**:
|
||||
```python
|
||||
@dataclass
|
||||
class MinerUChunk:
|
||||
content: str # 文本内容
|
||||
chunk_type: str # 类型: text, table, image, equation
|
||||
page_start: int = 1
|
||||
page_end: int = 1
|
||||
title: str = ""
|
||||
table_html: Optional[str] = None # 表格 HTML(如果是表格)
|
||||
image_path: Optional[str] = None # 图片路径(独立图片)
|
||||
images: Optional[List[Dict]] = None # 关联图片列表: [{"id": "abc.jpg", "order": 1}]
|
||||
```
|
||||
|
||||
**表格切片创建 (第453-469行)**:
|
||||
```python
|
||||
chunk = MinerUChunk(
|
||||
content=table_caption or "表格", # ⚠️ content 只有标题
|
||||
chunk_type="table",
|
||||
table_html=table_body, # ✅ 完整表格在 table_html
|
||||
image_path=img_path, # 表格图片路径
|
||||
images=table_images # 表格中嵌入的图片
|
||||
)
|
||||
```
|
||||
|
||||
### 2. 向量库存储 (knowledge/manager.py)
|
||||
|
||||
**add_file_to_kb() 核心逻辑 (第228-270行)**:
|
||||
```python
|
||||
for i, chunk in enumerate(chunks):
|
||||
# 1. 获取 chunk_type
|
||||
chunk_type = getattr(chunk, 'chunk_type', None)
|
||||
if not chunk_type:
|
||||
page_info = getattr(chunk, 'page_info', {}) or {}
|
||||
chunk_type = page_info.get('chunk_type', 'text')
|
||||
|
||||
if chunk_type == 'table':
|
||||
# 2. 表格内容优先使用 table_html
|
||||
table_md = getattr(chunk, 'table_html', None) or chunk.content
|
||||
semantic_content = _build_semantic_content_for_table(...)
|
||||
|
||||
# 3. 构建 metadata
|
||||
metadata = {
|
||||
'chunk_type': chunk_type,
|
||||
'source': filename,
|
||||
'page': page_start,
|
||||
# ...
|
||||
}
|
||||
|
||||
# 4. ✅ 序列化图片信息
|
||||
if hasattr(chunk, 'images') and chunk.images:
|
||||
metadata['images_json'] = json.dumps(chunk.images, ensure_ascii=False)
|
||||
|
||||
if hasattr(chunk, 'image_path') and chunk.image_path:
|
||||
metadata['image_path'] = chunk.image_path
|
||||
```
|
||||
|
||||
**跨页表格合并 (第323-420行)**:
|
||||
```python
|
||||
def _merge_cross_page_tables(self, chunks: list) -> list:
|
||||
"""
|
||||
合并规则:
|
||||
1. 相邻两个表格切片
|
||||
2. 页码连续 (page_end + 1 == next.page_start)
|
||||
3. 第二个表格标题包含"续表"
|
||||
"""
|
||||
# 合并 table_html
|
||||
current.table_html = curr_html + '\n' + next_html
|
||||
|
||||
# 合并 image_path 到 images
|
||||
merged_images = [
|
||||
{'id': curr_img, 'page': curr_page_end},
|
||||
{'id': next_img, 'page': next_page_start}
|
||||
]
|
||||
current.images = merged_images
|
||||
```
|
||||
|
||||
### 3. 富媒体召回 (api/chat_routes.py)
|
||||
|
||||
**_extract_rich_media() 核心逻辑 (第724-843行)**:
|
||||
```python
|
||||
def _extract_rich_media(contexts: List[Dict]) -> Dict[str, List]:
|
||||
images = []
|
||||
tables = []
|
||||
|
||||
for ctx in contexts:
|
||||
meta = ctx.get("meta", {})
|
||||
|
||||
# 1. 独立图片切片 (image_path) - 图片/图表类型
|
||||
if meta.get("chunk_type") in ("image", "chart") and meta.get("image_path"):
|
||||
img_id = os.path.basename(meta["image_path"])
|
||||
images.append({"id": img_id, "url": f"/images/{img_id}", ...})
|
||||
|
||||
# 2. 关联图片 (images_json) - 表格/文本嵌入图片
|
||||
if meta.get("images_json"):
|
||||
img_list = json.loads(meta["images_json"])
|
||||
for img_info in img_list:
|
||||
images.append({"id": img_info["id"], ...})
|
||||
|
||||
# 3. 表格图片 (image_path) - 表格类型的图片形式
|
||||
if meta.get("chunk_type") == "table" and meta.get("image_path"):
|
||||
img_id = os.path.basename(meta["image_path"])
|
||||
images.append({"id": img_id, "type": "table_image", ...})
|
||||
|
||||
return {"images": images, "tables": tables}
|
||||
```
|
||||
|
||||
## 当前问题分析
|
||||
|
||||
### 问题1: 表格显示"0行数据"
|
||||
|
||||
**根因**: `_build_semantic_content_for_table()` 接收的 `table_md` 可能是空的
|
||||
|
||||
**验证点**:
|
||||
- MinerU 解析时 `table_html` 是否有值?
|
||||
- `manager.py` 第240行 `table_md = getattr(chunk, 'table_html', None)` 是否正确获取?
|
||||
|
||||
### 问题2: 跨页表格合并不生效
|
||||
|
||||
**根因**: 可能是页码不连续或标题匹配失败
|
||||
|
||||
**验证点**:
|
||||
- 检查 `_merge_cross_page_tables()` 的日志输出
|
||||
- 验证两个表格切片的 `page_end` 和 `page_start` 是否连续
|
||||
|
||||
### 问题3: 图片重复
|
||||
|
||||
**根因**: 可能是 `images_json` 和 `image_path` 同时存在导致重复
|
||||
|
||||
**验证点**:
|
||||
- 检查向量库中是否有同时存在 `images_json` 和 `image_path` 的切片
|
||||
- `_extract_rich_media()` 中的去重逻辑是否有效
|
||||
|
||||
## 数据存储位置
|
||||
|
||||
| 目录 | 用途 |
|
||||
|------|------|
|
||||
| `.data/images/` | 全局图片存储(哈希命名,去重) |
|
||||
| `.data/cache/vlm/` | VLM 图片描述缓存 |
|
||||
| `.data/docstore/` | 原始表格/图片 JSON 备份 |
|
||||
| `knowledge/vector_store/chroma/` | ChromaDB 向量数据库 |
|
||||
|
||||
## 测试验证步骤
|
||||
|
||||
### 1. 检查向量库 metadata
|
||||
```python
|
||||
from knowledge.manager import get_kb_manager
|
||||
kb = get_kb_manager()
|
||||
coll = kb.get_collection('my_ky')
|
||||
result = coll.get(limit=10, include=['metadatas'])
|
||||
|
||||
for meta in result['metadatas']:
|
||||
print(f"chunk_type: {meta.get('chunk_type')}")
|
||||
print(f"images_json: {meta.get('images_json')}")
|
||||
print(f"image_path: {meta.get('image_path')}")
|
||||
print("---")
|
||||
```
|
||||
|
||||
### 2. 检查 MinerU 解析结果
|
||||
```python
|
||||
from parsers.mineru_parser import parse_with_mineru
|
||||
result = parse_with_mineru("tests/public/test_report.pdf")
|
||||
|
||||
for chunk in result.get('chunks', []):
|
||||
if chunk.chunk_type == 'table':
|
||||
print(f"表格标题: {chunk.title}")
|
||||
print(f"table_html 长度: {len(chunk.table_html or '')}")
|
||||
print(f"image_path: {chunk.image_path}")
|
||||
print(f"images: {chunk.images}")
|
||||
print("---")
|
||||
```
|
||||
Reference in New Issue
Block a user