Initial commit

This commit is contained in:
谭凯
2026-01-19 09:51:07 +08:00
commit 9ef82393be
174 changed files with 22285 additions and 0 deletions
+94
View File
@@ -0,0 +1,94 @@
# 更新日志 (CHANGELOG)
## [v0.08] - 2026-01-15 (The Optimizer)
### 🎯 占位符系统优化
- **前缀/后缀标签分离**: 文本首尾的纯格式标签(如 `<sup>`, `<sub>`)不再发送给 LLM,自动回填。
- **公式检测**: 数学变量和公式被识别为单一、不可翻译的占位符,显著降低 LLM 误翻译风险。
- **简化占位符格式**: 从全局唯一 `φcXXXXXφ` 简化为段落局部 `φ1φ`,每个段落独立编号。
### 🎨 纯中文排版模式
- **模式切换**: 支持 `--mode chinese` 生成保留原始排版的纯中文译本(默认 `--mode bilingual`)。
- **Format Extractor (替换法)**:
- 彻底重构格式提取逻辑,放弃 DOM 递归,采用基于正则的"标签序列替换法"。
- 能够完美处理任意深度的嵌套标签,将其合并为单一占位符。
- 100% 保留原始 HTML 属性(class, style, href 等),实现"像素级"格式还原。
- **Format Restorer (自愈系)**:
- 引入 `FormatRestorer` 模块,负责将占位符替换回原始 HTML 代码。
- **自动修复 Agent**: 当检测到 LLM 丢失占位符时,自动触发回退机制进行格式修复。
- **优雅降级**: 如果修复失败,系统会自动降级为纯文本,确保程序不崩溃。
### 🛡️ 深度优化
- **占位符升级**: 从易混淆的 `«c...»` 升级为 `φc...φ`,显著降低 LLM 误翻译概率。
- **容器样式继承**:
- 中文模式:直接替换 `inner_html`,完美保留外层容器属性。
- 双语模式:新建 `<p>` 标签时自动继承原文的 `class` 和其他属性。
- **智能测试**: `--test` 模式逻辑升级,智能识别章节边界,自动翻译完第一章。
### 🔧 修复
- 修复了 `LLMClient` 中正则表达式转义错误导致的 `FutureWarning`
- 修复了 `FormatExtractor` 循环引用问题。
- 解决了复杂科学书籍中上标/链接嵌套导致的校验失败问题。
---
## [v0.07] - 2026-01-13 (The Refinement)
### 🛡️ 安全与配置
- **环境隔离**: 引入 `.env` 支持,彻底移除了代码库中的硬编码 API Key。
- **配置升级**: `utils.py` 现自动加载 `.env` 并注入到配置中,支持任意 Provider 的环境变量覆盖 (如 `V3_API_KEY`, `OPENROUTER_API_KEY`)。
- **模板化**: 新增 `config.example.json``.env` 模板,提升部署安全性。
### 🚀 核心改进
- **V3 Provider 支持**: 验证并修复了对 V3 API (OpenAI 兼容格式) 的支持,全流程跑通。
- **EPUB 构建修复**: 解决了 `ebooklib` 在处理 TOC 时因缺少 UID 导致的 `Argument must be bytes or unicode` 崩溃问题。
- **缓存优化**:
- 缓存目录结构调整为 Hash 前缀 (`cache/translations/ab/...`),解决了按日期分目录导致的缓存频繁失效问题。
- 放宽了缓存验证逻辑,支持部分命中的缓存复用。
- **视觉优化**: 引入“盘古之白” (Pangu spacing),自动在中文与英文/数字之间添加空格,显著提升阅读体验。
### ⚡ 体验提升
- **断点续传提示**: 启动时自动检测并提示未完成的翻译进度。
- **详细统计**: 翻译完成后展示详细的成功/失败/跳过统计数据。
- **并发优化**: 移除了冗余的信号量控制,完全依赖 `RateLimiter`,逻辑更清晰高效。
---
## [v0.05] - 2026-01-12 (The Arena)
### 🌟 核心突破
- **书籍画像 (Book Profiler)**:
- 自动提取前言和正文采样。
- 生成 `Book Profile`,包含领域 (Genre)、文风 (Style)、目标受众 (Audience) 和翻译指令。
- 生成 `Glossary` (术语表),并支持自动注入 Prompt。
- **状态绑定**: Profile 和 Glossary 现在直接存储在每本书的 `manifest.json` 中,互不干扰。
- **模型竞技场 (Model Arena)**:
- 自动选取典型 Chunk,让多个候选模型 (Gemini, Llama, Qwen) 同台竞技。
- 引入 `Judge Agent` (基于 Smart 模型),从准确性、信达雅维度评选最佳模型。
- 自动锁定获胜模型用于全书翻译。
### 🏗️ 架构升级
- **LLM Client 重构**:
- **Syntax Fixes**: 彻底修复了正则构造中的语法错误。
- **Quote Safety**: 移除了所有 f-string 中的复杂正则,改用安全的字符串拼接。
- **Dual RateLimiters**: 引入主/副限流器,防止死锁。
- **配置增强**:
- `config.json` 支持 `arena_models``judge_model` 配置。
### 🔧 修复与优化
- 修复了 `unhashable type: 'dict'` 错误 (移除了错误的 `{{}}`)。
- 修复了多本书连续翻译时 Profile 串用的问题 (Profile 现已绑定至 Manifest)。
---
## [v0.03] - 2026-01-12
- **极简 ID 锚点系统**: 废弃复杂的 `[p_xxxxx]` 格式,使用纯净 ID,彻底解决残留问题。
- **智能术语一致性**: 引入 GlossaryManager。
- **结构完美保留**: 修复了 EPUB Spine 和 Metadata 丢失问题。
## [v0.02] - 2026-01-12
- **Manifest 驱动架构**: 引入 `ManifestManager` 作为单一真理源。
- **流程解耦**: 提取、翻译、构建三阶段分离。
## [v0.01] - 2026-01-10
- 初始版本,实现基本的并发翻译和 EPUB 解析。
+97
View File
@@ -0,0 +1,97 @@
# EPUB 双语翻译程序 v0.07
一个基于 OpenRouter/OpenAI API 的 EPUB 双语翻译工具,采用**全局编号系统**和**真并发翻译**。
## ✨ 核心特性
### 🎯 全局编号系统
- **每个段落分配全局唯一ID**(格式:`p_0001`, `p_0002`...
- **ID贯穿全流程**:提取 → 翻译 → 组装
- **精确对应保证**:绝不出现中英文错行问题
### ⚡ 真并发翻译
- **asyncio.gather 并发执行**:高效利用 API 速率限制
- **智能速率控制**:基于 Token 桶的 RateLimiter
- **实时进度显示**:Rich 进度条显示翻译状态
- **断点续传**:自动记录进度,随时中断随时继续
### 🛡️ 安全与稳定
- **环境隔离**:支持 `.env` 配置,API Key 不落地
- **鲁棒重试**:集成 `tenacity` 处理网络波动
- **缓存系统**:基于 Hash 的持久化缓存,跨天复用
### 🎨 极致排版
- **盘古之白**:自动在中文与西文数字间添加空格
- **样式注入**:注入专用 CSS 优化阅读体验
## 🚀 快速开始
### 1. 安装依赖
```bash
pip install -r requirements.txt
```
### 2. 配置环境
复制 `.env` 模板并填入你的 API Key
```bash
# .env 文件
V3_API_KEY=sk-xxxxxx
OPENROUTER_API_KEY=sk-or-xxxxxx
```
### 3. 开始翻译
```bash
# 默认使用 OpenRouter
python main.py input/book.epub
# 使用 V3 Provider
python main.py input/book.epub -p v3
# 测试模式(只翻译前3个块)
python main.py input/book.epub --test
```
## 📂 目录结构
```
.
├── config/ # 配置文件
│ ├── config.json # 主配置
│ └── prompts.json # 提示词模板
├── input/ # 输入 EPUB 目录
├── output/ # 输出 EPUB 目录
├── cache/ # 缓存目录 (Manifest, Translations)
├── logs/ # 运行日志
└── src/ # 源代码
```
## ⚙️ 核心配置 (config.json)
```json
{
"translation": {
"chunk_size": 5000,
"temperature": 0.3
},
"providers": {
"v3": {
"base_url": "https://api.gpt.ge/v1",
"models": { "fast": "gpt-4o-mini" },
"rate_limits": { "requests_per_minute": 500 }
}
}
}
```
## 📄 许可证
MIT License
---
**版本**: v0.07
**更新**: 2026-01-13
+54
View File
@@ -0,0 +1,54 @@
{
"translation": {
"chunk_size": 5000,
"temperature": 0.3,
"glossary": {
"enabled": true,
"auto_generate": true,
"sample_size": 3000
}
},
"output": {
"output_dir": "output",
"filename_suffix": "_bilingual"
},
"logging": {
"level": "INFO",
"file": "logs/translator.log",
"rotation": "10 MB",
"retention": "7 days"
},
"providers": {
"openrouter": {
"base_url": "https://openrouter.ai/api/v1",
"api_key": "YOUR_OPENROUTER_API_KEY",
"models": {
"fast": "google/gemini-2.0-flash-001",
"smart": "google/gemini-2.0-flash-thinking-exp:free"
},
"extra_headers": {
"HTTP-Referer": "https://github.com/epub-translator",
"X-Title": "EPUB Translator"
},
"rate_limits": {
"requests_per_minute": 60,
"concurrent_requests": 32
}
},
"v3": {
"base_url": "https://api.gpt.ge/v1",
"api_key": "YOUR_V3_API_KEY",
"models": {
"fast": "gpt-4o-mini",
"smart": "gpt-4o"
},
"extra_headers": {
"x-foo": "true"
},
"rate_limits": {
"requests_per_minute": 500,
"concurrent_requests": 50
}
}
}
}
+69
View File
@@ -0,0 +1,69 @@
{
"translation": {
"chunk_size": 5000,
"temperature": 0.3,
"glossary": {
"enabled": true,
"auto_generate": true,
"sample_size": 3000
}
},
"output": {
"output_dir": "output",
"filename_suffix": "_bilingual"
},
"logging": {
"level": "INFO",
"file": "logs/translator.log",
"rotation": "10 MB",
"retention": "7 days"
},
"providers": {
"openrouter": {
"base_url": "https://openrouter.ai/api/v1",
"api_key": "YOUR_OPENROUTER_API_KEY",
"models": {
"fast": "google/gemini-2.0-flash-001",
"smart": "google/gemini-2.0-flash-thinking-exp:free"
},
"extra_headers": {
"HTTP-Referer": "https://github.com/epub-translator",
"X-Title": "EPUB Translator"
},
"rate_limits": {
"requests_per_minute": 60,
"concurrent_requests": 32
}
},
"v3": {
"base_url": "https://api.gpt.ge/v1",
"api_key": "YOUR_V3_API_KEY",
"models": {
"fast": "gemini-3-flash-preview",
"smart": "gemini-3-pro-preview"
},
"extra_headers": {
"x-foo": "true"
},
"rate_limits": {
"requests_per_minute": 500,
"concurrent_requests": 50
}
},
"openai": {
"base_url": "http://127.0.0.1:8045/v1",
"api_key": "YOUR_ANTIGRAVITY_API_KEY",
"models": {
"fast": "gemini-3-flash",
"smart": "gemini-3-pro-high"
},
"extra_headers": {
"x-foo": "true"
},
"rate_limits": {
"requests_per_minute": 500,
"concurrent_requests": 50
}
}
}
}
+39
View File
@@ -0,0 +1,39 @@
{
"llm": {
"provider": "openrouter",
"base_url": "https://openrouter.ai/api/v1",
"api_key": "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",
"models": {
"fast": "google/gemini-3-flash-preview",
"smart": "google/gemini-3-pro-preview"
},
"arena_models": [
"google/gemini-3-flash-preview",
"openai/gpt-5-mini",
"anthropic/claude-haiku-4.5"
],
"judge_model": "openai/gpt-5.1",
"rate_limits": {
"requests_per_minute": 60,
"concurrent_requests": 32
}
},
"translation": {
"chunk_size": 5000,
"temperature": 0.3,
"strategy": "arena_winner"
},
"processing": {
"min_paragraph_length": 5
},
"output": {
"output_dir": "output",
"filename_suffix": "_bilingual"
},
"logging": {
"level": "INFO",
"file": "logs/translator.log",
"rotation": "10 MB",
"retention": "7 days"
}
}
+10
View File
@@ -0,0 +1,10 @@
{
"translation": {
"system": "你是一位精通中英文的专业翻译家。你的任务是翻译书籍内容。\n\n要求:\n1. 准确传达原文含义,语言流畅自然,符合中文阅读习惯。\n2. 严格保持【p_xxxxx】编号格式,不要遗漏,不要修改编号。\n3. 不要添加任何解释、注释或无关内容,只返回【编号】+【译文】。\n\n{{glossary_instruction}}",
"user_template": "请翻译以下段落:\n\n{{content}}"
},
"glossary_extraction": {
"system": "你是一位资深的文学编辑和领域专家。你的任务是分析书籍样本,提取关键术语并制定统一的译名表。",
"user_template": "请阅读以下书籍片段(包含前言和正文采样)。\n\n任务:\n1. 识别文中出现的人名(如 'Masa', 'Steve Jobs')、地名、机构名。\n2. 识别特定的行业术语或关键概念。\n3. 为上述词汇提供标准的中文译名。如果像 'Masa' 这样的昵称有对应的全名(如孙正义),请务必使用全名。\n\n请以 JSON 格式输出,格式如下:\n{\n \"Masa\": \"孙正义\",\n \"Apple\": \"苹果公司\",\n ...\n}\n\n书籍片段:\n\n{{content}}"
}
}
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env python3
import asyncio
import sys
import argparse
from pathlib import Path
from loguru import logger
from src.translator import EPUBTranslator
from src.utils import load_config, setup_logging
def parse_args():
parser = argparse.ArgumentParser(description="EPUB 双语翻译工具")
parser.add_argument("epub_path", help="输入 EPUB 文件路径")
parser.add_argument("--provider", "-p", default="openrouter", help="LLM 供应商 (config.json 中 providers 的 key)")
parser.add_argument("--mode", "-m", default="bilingual", choices=["bilingual", "chinese"],
help="输出模式: bilingual (双语对照) 或 chinese (纯中文,保留格式)")
parser.add_argument("--test", action="store_true", help="测试模式(仅翻译前几段)")
parser.add_argument("--output", "-o", help="输出目录")
parser.add_argument("--no-cache", action="store_true", help="禁用缓存(强制重新翻译)")
parser.add_argument("--clear-cache", action="store_true", help="清理所有缓存文件")
return parser.parse_args()
def flatten_provider_config(config: dict, provider_name: str) -> dict:
"""
将选定的 provider 配置扁平化到 config['llm'] 中,
以便下游模块统一调用。
"""
providers = config.get('providers', {})
if provider_name not in providers:
available = list(providers.keys())
logger.error(f"未找到供应商 '{provider_name}'。可用供应商: {available}")
sys.exit(1)
selected_config = providers[provider_name]
logger.info(f"使用 LLM 供应商: {provider_name} ({selected_config.get('base_url')})")
# 注入到 config['llm']
config['llm'] = selected_config
return config
async def run_translation(args):
try:
# 1. 加载配置
config = load_config()
# 2. 处理 Provider 选择
config = flatten_provider_config(config, args.provider)
# 3. 设置日志
setup_logging(config)
logger.info("程序启动")
# 4. 初始化翻译器
translator = EPUBTranslator(config, use_cache=not args.no_cache)
# 5. 执行翻译
await translator.translate_epub(
args.epub_path,
test_mode=args.test,
output_dir=args.output,
mode=args.mode
)
except Exception as e:
import traceback
traceback.print_exc()
logger.error(f"翻译失败: {e}")
sys.exit(1)
def main():
args = parse_args()
if args.clear_cache:
import shutil
cache_dir = Path("cache")
if cache_dir.exists():
shutil.rmtree(cache_dir)
print("缓存已清理")
sys.exit(0)
asyncio.run(run_translation(args))
if __name__ == "__main__":
main()
+34
View File
@@ -0,0 +1,34 @@
[project]
name = "epub-translator"
version = "0.07"
description = "EPUB双语翻译程序"
authors = [
{name = "Kaitan", email = "your-email@example.com"}
]
dependencies = [
"ebooklib>=0.19",
"beautifulsoup4>=4.12.0",
"lxml>=4.9.0",
"openai>=1.0.0",
"aiohttp>=3.9.0",
"pydantic>=2.0.0",
"loguru>=0.7.0",
"rich>=13.0.0",
"asyncio-throttle>=1.0.2",
"tenacity>=8.0.0",
"python-dotenv>=1.0.0",
]
requires-python = ">=3.9"
[project.scripts]
epub-translator = "main:main"
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.uv]
dev-dependencies = [
"pytest>=7.0.0",
"pytest-asyncio>=0.21.0",
]
+12
View File
@@ -0,0 +1,12 @@
ebooklib>=0.19
beautifulsoup4>=4.12.0
lxml>=4.9.0
openai>=1.0.0
aiohttp>=3.9.0
pydantic>=2.0.0
loguru>=0.7.0
rich>=13.0.0
asyncio-throttle>=1.0.2
tenacity>=8.0.0
python-dotenv>=1.0.0
socksio>=1.0.0
+32
View File
@@ -0,0 +1,32 @@
"""
EPUB 双语翻译程序
主要功能模块的初始化文件
"""
__version__ = "0.08"
__author__ = "Kaitan"
from .epub_parser import EPUBParser
from .translator import EPUBTranslator
from .llm_client import LLMClient as OpenRouterClient # Keep alias for compatibility
from .llm_client import LLMClient
from .text_processor import TextProcessor
from .bilingual_builder import BilingualEPUBBuilder
from .chinese_builder import ChineseEPUBBuilder
from .format_extractor import FormatExtractor
from .format_restorer import FormatRestorer
from .utils import load_config, setup_logging
__all__ = [
"EPUBParser",
"EPUBTranslator",
"LLMClient",
"OpenRouterClient",
"TextProcessor",
"BilingualEPUBBuilder",
"ChineseEPUBBuilder",
"FormatExtractor",
"FormatRestorer",
"load_config",
"setup_logging"
]
+180
View File
@@ -0,0 +1,180 @@
"""
双语 EPUB 构建器模块 - 安全的EPUB构建 (Manifest 兼容版)
"""
from ebooklib import epub
import ebooklib
from bs4 import BeautifulSoup
from typing import Dict, List
from pathlib import Path
from loguru import logger
import uuid
class BilingualEPUBBuilder:
"""双语 EPUB 构建器"""
def __init__(self, original_book, config: Dict):
self.original_book = original_book
self.config = config
self.output_config = config['output']
def create_bilingual_epub_with_mapping(self, translation_map: Dict[str, str],
paragraph_map: Dict[str, Dict],
output_path: str) -> str:
"""
创建双语 EPUB。使用 ordered_ids 确保与 Manifest 严格一致。
"""
try:
new_book = epub.EpubBook()
self._copy_metadata(new_book)
new_book.toc = self._sanitize_toc(self.original_book.toc)
# 准备每个文件的有序ID列表
file_ordered_ids = {}
sorted_pids = sorted(paragraph_map.keys(), key=lambda x: int(x.split('_')[1]))
for pid in sorted_pids:
info = paragraph_map[pid]
fname = info['file_name']
if fname not in file_ordered_ids:
file_ordered_ids[fname] = []
file_ordered_ids[fname].append(pid)
processed_item_ids = set()
item_map = {}
# 特殊处理:封面图片
cover_id_meta = self.original_book.get_metadata('OPF', 'cover')
if cover_id_meta:
cover_item = self.original_book.get_item_with_id(cover_id_meta[0][0])
if cover_item:
new_book.add_item(cover_item)
processed_item_ids.add(cover_item.id)
item_map[cover_item.id] = cover_item
# 复制资源
for item in self.original_book.get_items():
if item.get_type() != ebooklib.ITEM_DOCUMENT:
if item.id not in processed_item_ids:
new_book.add_item(item)
processed_item_ids.add(item.id)
item_map[item.id] = item
# 重建 Spine
new_spine = []
for spine_id, linear in self.original_book.spine:
item = self.original_book.get_item_with_id(spine_id)
if not item: continue
if item.get_type() == ebooklib.ITEM_DOCUMENT:
file_name = item.get_name()
if file_name in file_ordered_ids:
new_item = self._create_bilingual_document(
item, file_ordered_ids[file_name], translation_map
)
new_item.id = item.id
else:
new_item = item
if new_item.id not in processed_item_ids:
new_book.add_item(new_item)
processed_item_ids.add(new_item.id)
new_spine.append(new_item)
else:
if item.id in item_map:
new_spine.append(item_map[item.id])
new_book.spine = new_spine
new_book.add_item(epub.EpubNcx())
new_book.add_item(epub.EpubNav())
output_file = self._generate_output_filename(output_path)
epub.write_epub(output_file, new_book, {})
return output_file
except Exception as e:
logger.error(f"创建双语 EPUB 失败: {e}", exc_info=True)
raise
def _sanitize_toc(self, toc):
"""确保 TOC 中的所有节点都有 ID,修复 ebooklib 读取后写入的兼容性问题"""
for item in toc:
if isinstance(item, (epub.Link, epub.Section)):
if not getattr(item, 'uid', None):
item.uid = f"navPoint-{uuid.uuid4().hex[:8]}"
elif isinstance(item, tuple) and len(item) == 2:
# 处理 (Section, [children]) 结构
section, children = item
if isinstance(section, (epub.Link, epub.Section)):
if not getattr(section, 'uid', None):
section.uid = f"navPoint-{uuid.uuid4().hex[:8]}"
self._sanitize_toc(children)
return toc
def _copy_metadata(self, new_book):
try:
for namespace, meta_dict in self.original_book.metadata.items():
for name, values in meta_dict.items():
for value, other in values:
if name and hasattr(name, 'lower') and name.lower() == 'identifier': continue
new_book.add_metadata(namespace, name, value, other)
new_book.add_metadata('DC', 'language', 'zh-CN')
new_book.set_identifier(f"bilingual-{uuid.uuid4().hex[:12]}")
cover_id_meta = self.original_book.get_metadata('OPF', 'cover')
if cover_id_meta:
cover_item = self.original_book.get_item_with_id(cover_id_meta[0][0])
if cover_item:
new_book.add_item(cover_item)
new_book.set_cover(cover_item.get_name(), cover_item.get_content())
except Exception as e:
logger.error(f"元数据复制出错: {e}")
def _create_bilingual_document(self, original_item, ordered_ids: list, translation_map: dict):
try:
from .text_processor import TextProcessor
soup = BeautifulSoup(original_item.get_content().decode('utf-8'), 'html.parser')
self._add_style_link(soup)
# 使用与 TextProcessor 相同的过滤逻辑获取元素
text_elements = TextProcessor.get_valid_text_elements(soup)
current_para_index = 0
for element in text_elements:
if TextProcessor.is_navigation_element(element): continue
if not TextProcessor.clean_element_text(element): continue
if current_para_index < len(ordered_ids):
target_id = ordered_ids[current_para_index]
translation = translation_map.get(target_id)
if translation:
self._insert_translation(element, translation, soup)
current_para_index += 1
new_item = epub.EpubHtml(title=original_item.title, file_name=original_item.get_name(), lang='zh-CN')
new_item.set_content(str(soup).encode('utf-8'))
return new_item
except Exception as e:
logger.error(f"创建双语文档失败 {original_item.get_name()}: {e}")
return original_item
def _add_style_link(self, soup):
head = soup.find('head')
if head and not head.find('link', href='style/bilingual.css'):
head.append(soup.new_tag('link', rel='stylesheet', type='text/css', href='style/bilingual.css'))
def _insert_translation(self, element, translation: str, soup):
try:
translation_p = soup.new_tag('p')
translation_p.string = translation
translation_p['class'] = ['translation-text', 'chinese']
element.insert_after(translation_p)
except Exception as e:
logger.warning(f"插入翻译失败: {e}")
def _generate_output_filename(self, output_path: str) -> str:
from .utils import sanitize_filename
title = self.original_book.get_metadata('DC', 'title')
clean_title = sanitize_filename(title[0][0]) if title else "bilingual_book"
Path(output_path).mkdir(parents=True, exist_ok=True)
return str(Path(output_path) / f"{clean_title}_bilingual.epub")
+84
View File
@@ -0,0 +1,84 @@
'''Book Profiler Module
Features:
1. Automatically extract book samples to generate Book Profile (Genre, Style, Glossary).
'''
import json
import random
from pathlib import Path
from typing import Dict, List
from loguru import logger
from .manifest_manager import ManifestManager
from .llm_client import LLMClient
class BookProfiler:
def __init__(self, config: Dict, llm_client: LLMClient):
self.config = config
self.llm_client = llm_client
def extract_sample_text(self, manifest: ManifestManager, char_limit: int = 3000) -> str:
"""Extract sample text."""
items = manifest.get_items()
if not items: return ""
intro_text = []
for item in items[:50]:
if len(item.clean_text) > 50:
intro_text.append(item.clean_text)
body_text = []
body_items = [i for i in items[50:] if len(i.clean_text) > 80]
if body_items:
samples = random.sample(body_items, min(5, len(body_items)))
body_text = [i.clean_text for i in samples]
full_text = "\n\n".join(intro_text[:5] + body_text)
return full_text[:char_limit]
async def analyze_book(self, manifest: ManifestManager) -> Dict:
"""Generate Book Profile."""
existing_profile = manifest.data.get('metadata', {}).get('profile')
if existing_profile:
logger.info("Loaded existing Book Profile")
return existing_profile
sample = self.extract_sample_text(manifest)
if not sample: return {}
logger.info("Generating Book Profile...")
system_prompt = "You are a senior publishing editor. Analyze the text and output JSON."
user_prompt = f"""
Please analyze the following book excerpt.
Output JSON format:
{{
"genre": "Genre",
"style": "Style description",
"audience": "Target Audience",
"glossary": {{ "Term": "Chinese Translation" }},
"translation_instruction": "Specific instruction for translator"
}}
Excerpt:
{sample}
"""
try:
response = await self.llm_client.raw_chat_completion(system_prompt, user_prompt)
json_str = response.strip()
if "```json" in json_str:
json_str = json_str.split("```json")[1].split("```")[0].strip()
elif "```" in json_str:
json_str = json_str.split("```")[1].split("```")[0].strip()
profile = json.loads(json_str)
if 'metadata' not in manifest.data:
manifest.data['metadata'] = {}
manifest.data['metadata']['profile'] = profile
manifest.save()
return profile
except Exception as e:
logger.error(f"Profile generation failed: {e}")
return {}
+226
View File
@@ -0,0 +1,226 @@
"""
翻译缓存管理模块 - 简化版
基于全局ID和chunk的缓存系统
"""
import json
import hashlib
from pathlib import Path
from datetime import datetime, timedelta
from typing import Dict, Optional, List
from loguru import logger
class TranslationCache:
"""翻译缓存管理器 - 简化版"""
def __init__(self, config: Dict):
"""初始化缓存管理器"""
self.config = config
cache_config = config.get('cache', {})
self.enabled = cache_config.get('enabled', True)
self.cache_dir = Path(cache_config.get('directory', 'cache'))
self.max_age_days = cache_config.get('max_age_days', 30)
if self.enabled:
self.cache_dir.mkdir(parents=True, exist_ok=True)
self.translations_dir = self.cache_dir / 'translations'
self.translations_dir.mkdir(parents=True, exist_ok=True)
logger.info(f"翻译缓存已启用: {self.cache_dir}")
def get_chunk_translation(self, chunk: List[Dict], model: str) -> Optional[Dict[str, str]]:
"""
获取chunk的缓存翻译
Args:
chunk: 段落列表(带global_id
model: 模型名称
Returns:
{global_id: translation} 映射,如果不存在返回 None
"""
if not self.enabled:
return None
try:
cache_key = self._get_chunk_cache_key(chunk, model)
cache_file = self._get_cache_file_path(cache_key)
if not cache_file.exists():
return None
# 检查是否过期
file_age = datetime.now() - datetime.fromtimestamp(cache_file.stat().st_mtime)
if file_age > timedelta(days=self.max_age_days):
logger.debug(f"缓存已过期: {cache_key[:8]}...")
cache_file.unlink()
return None
# 读取缓存
with open(cache_file, 'r', encoding='utf-8') as f:
cache_data = json.load(f)
# 验证缓存
if (cache_data.get('success') and
cache_data.get('model') == model and
self._validate_cache_data(cache_data, chunk)):
logger.debug(f"缓存命中: {cache_key[:8]}... ({len(chunk)} 段落)")
return cache_data.get('translations', {})
return None
except Exception as e:
logger.warning(f"读取缓存失败: {e}")
return None
def save_chunk_translation(self, chunk: List[Dict], translations: Dict[str, str],
model: str, success: bool = True) -> None:
"""
保存chunk翻译到缓存
Args:
chunk: 段落列表(带global_id
translations: {global_id: translation} 映射
model: 模型名称
success: 是否翻译成功
"""
if not self.enabled:
return
try:
cache_key = self._get_chunk_cache_key(chunk, model)
cache_file = self._get_cache_file_path(cache_key)
# 构建缓存数据
cache_data = {
'global_ids': [p['global_id'] for p in chunk],
'translations': translations,
'model': model,
'timestamp': datetime.now().isoformat(),
'success': success,
'paragraph_count': len(chunk),
'cache_version': '3.0'
}
with open(cache_file, 'w', encoding='utf-8') as f:
json.dump(cache_data, f, ensure_ascii=False, indent=2)
logger.debug(f"缓存已保存: {cache_key[:8]}... ({len(chunk)} 段落)")
except Exception as e:
logger.warning(f"保存缓存失败: {e}")
def _get_chunk_cache_key(self, chunk: List[Dict], model: str) -> str:
"""
生成chunk缓存键(基于全局ID序列)
Args:
chunk: 段落列表
model: 模型名称
Returns:
缓存键
"""
# 使用全局ID序列作为缓存键的一部分
id_sequence = ",".join(p['global_id'] for p in chunk)
combined = f"{id_sequence}|{model}"
return hashlib.md5(combined.encode('utf-8')).hexdigest()
def _get_cache_file_path(self, cache_key: str) -> Path:
"""获取缓存文件路径"""
# 使用 hash 前缀分目录,避免单目录文件过多
subdir = cache_key[:2]
cache_subdir = self.translations_dir / subdir
cache_subdir.mkdir(parents=True, exist_ok=True)
return cache_subdir / f"{cache_key}.json"
def _validate_cache_data(self, cache_data: Dict, chunk: List[Dict]) -> bool:
"""验证缓存数据的有效性"""
# 检查ID序列是否匹配
cached_ids = cache_data.get('global_ids', [])
chunk_ids = [p['global_id'] for p in chunk]
if cached_ids != chunk_ids:
logger.debug("缓存ID序列不匹配")
return False
# 只要有翻译结果就认为有效,不要求数量完全匹配
translations = cache_data.get('translations', {})
if not translations:
logger.debug("缓存无翻译结果")
return False
return True
def clear_cache(self, older_than_days: Optional[int] = None) -> int:
"""清理缓存"""
if not self.enabled or not self.translations_dir.exists():
return 0
cleared_count = 0
cutoff_time = None
if older_than_days is not None:
cutoff_time = datetime.now() - timedelta(days=older_than_days)
try:
for cache_file in self.translations_dir.rglob('*.json'):
should_delete = False
if cutoff_time is None:
should_delete = True
else:
file_time = datetime.fromtimestamp(cache_file.stat().st_mtime)
should_delete = file_time < cutoff_time
if should_delete:
cache_file.unlink()
cleared_count += 1
# 清理空目录
for date_dir in self.translations_dir.iterdir():
if date_dir.is_dir() and not any(date_dir.iterdir()):
date_dir.rmdir()
logger.info(f"清理了 {cleared_count} 个缓存文件")
return cleared_count
except Exception as e:
logger.error(f"清理缓存失败: {e}")
return 0
def get_cache_stats(self) -> Dict:
"""获取缓存统计信息"""
if not self.enabled or not self.translations_dir.exists():
return {'enabled': False}
try:
cache_files = list(self.translations_dir.rglob('*.json'))
total_files = len(cache_files)
total_size = sum(f.stat().st_size for f in cache_files)
# 统计段落数
total_paragraphs = 0
for cache_file in cache_files:
try:
with open(cache_file, 'r', encoding='utf-8') as f:
data = json.load(f)
total_paragraphs += data.get('paragraph_count', 0)
except:
continue
return {
'enabled': True,
'total_files': total_files,
'total_paragraphs': total_paragraphs,
'total_size_mb': round(total_size / 1024 / 1024, 2),
'cache_directory': str(self.cache_dir),
'max_age_days': self.max_age_days
}
except Exception as e:
logger.error(f"获取缓存统计失败: {e}")
return {'enabled': True, 'error': str(e)}
+180
View File
@@ -0,0 +1,180 @@
"""
纯中文 EPUB 构建器模块
负责:
1. 复制原始 EPUB 结构
2. 使用译文替换原文
3. 调用 FormatRestorer 将译文占位符还原为 HTML 标签
"""
from ebooklib import epub
import ebooklib
from bs4 import BeautifulSoup
from typing import Dict, List
from pathlib import Path
from loguru import logger
import uuid
class ChineseEPUBBuilder:
"""纯中文 EPUB 构建器"""
def __init__(self, original_book, config: Dict):
self.original_book = original_book
self.config = config
self.output_config = config['output']
def create_chinese_epub_with_mapping(self,
items: List, # List[ManifestItem]
output_path: str) -> str:
"""
创建纯中文 EPUB。
"""
try:
new_book = epub.EpubBook()
self._copy_metadata(new_book)
new_book.toc = self._sanitize_toc(self.original_book.toc)
# 准备每个文件的有序项目列表
file_items = {}
for item in sorted(items, key=lambda x: x.global_id):
fname = item.source_file
if fname not in file_items:
file_items[fname] = []
file_items[fname].append(item)
processed_item_ids = set()
item_map = {}
# 特殊处理:封面图片
cover_id_meta = self.original_book.get_metadata('OPF', 'cover')
if cover_id_meta:
cover_item = self.original_book.get_item_with_id(cover_id_meta[0][0])
if cover_item:
new_book.add_item(cover_item)
processed_item_ids.add(cover_item.id)
item_map[cover_item.id] = cover_item
# 复制资源
for item in self.original_book.get_items():
if item.get_type() != ebooklib.ITEM_DOCUMENT:
if item.id not in processed_item_ids:
new_book.add_item(item)
processed_item_ids.add(item.id)
item_map[item.id] = item
# 重建 Spine
new_spine = []
for spine_id, linear in self.original_book.spine:
item = self.original_book.get_item_with_id(spine_id)
if not item: continue
if item.get_type() == ebooklib.ITEM_DOCUMENT:
file_name = item.get_name()
if file_name in file_items:
new_item = self._create_chinese_document(
item, file_items[file_name]
)
new_item.id = item.id
else:
new_item = item
if new_item.id not in processed_item_ids:
new_book.add_item(new_item)
processed_item_ids.add(new_item.id)
new_spine.append(new_item)
else:
if item.id in item_map:
new_spine.append(item_map[item.id])
new_book.spine = new_spine
new_book.add_item(epub.EpubNcx())
new_book.add_item(epub.EpubNav())
output_file = self._generate_output_filename(output_path)
epub.write_epub(output_file, new_book, {})
return output_file
except Exception as e:
logger.error(f"创建纯中文 EPUB 失败: {e}", exc_info=True)
raise
def _sanitize_toc(self, toc):
"""确保 TOC 中的所有节点都有 ID"""
for item in toc:
if isinstance(item, (epub.Link, epub.Section)):
if not getattr(item, 'uid', None):
item.uid = f"navPoint-{uuid.uuid4().hex[:8]}"
elif isinstance(item, tuple) and len(item) == 2:
section, children = item
if isinstance(section, (epub.Link, epub.Section)):
if not getattr(section, 'uid', None):
section.uid = f"navPoint-{uuid.uuid4().hex[:8]}"
self._sanitize_toc(children)
return toc
def _copy_metadata(self, new_book):
try:
for namespace, meta_dict in self.original_book.metadata.items():
for name, values in meta_dict.items():
for value, other in values:
if name and hasattr(name, 'lower') and name.lower() == 'identifier': continue
new_book.add_metadata(namespace, name, value, other)
new_book.add_metadata('DC', 'language', 'zh-CN')
new_book.set_identifier(f"chinese-{uuid.uuid4().hex[:12]}")
except Exception as e:
logger.error(f"元数据复制出错: {e}")
def _create_chinese_document(self, original_item, manifest_items: list):
try:
from .text_processor import TextProcessor
soup = BeautifulSoup(original_item.get_content().decode('utf-8'), 'html.parser')
# 获取文本元素
text_elements = TextProcessor.get_valid_text_elements(soup)
current_idx = 0
for element in text_elements:
if not TextProcessor.clean_element_text(element): continue
if current_idx < len(manifest_items):
m_item = manifest_items[current_idx]
# 只有当非导航元素时才尝试替换内容
if not TextProcessor.is_navigation_element(element):
# 优先使用带格式的翻译,降级到纯文本翻译
if m_item.translation_with_original_html:
self._replace_content(element, m_item.translation_with_original_html)
elif m_item.translation:
# 降级:使用纯文本翻译(无格式)
self._replace_content(element, m_item.translation)
current_idx += 1
new_item = epub.EpubHtml(title=original_item.title, file_name=original_item.get_name(), lang='zh-CN')
new_item.set_content(str(soup).encode('utf-8'))
return new_item
except Exception as e:
logger.error(f"创建中文文档失败 {original_item.get_name()}: {e}")
return original_item
def _replace_content(self, element, translated_html: str):
"""用译文替换元素的 inner_html"""
try:
# 将译文 HTML 字符串解析为 BeautifulSoup 对象
new_soup = BeautifulSoup(translated_html, 'html.parser')
# 清空原元素并填入新内容
element.clear()
# 重要:必须先转换为 list,否则 append 会修改 contents 导致跳过元素
for content in list(new_soup.contents):
element.append(content)
except Exception as e:
logger.error(f"替换内容失败: {e}")
def _generate_output_filename(self, output_path: str) -> str:
from .utils import sanitize_filename
title = self.original_book.get_metadata('DC', 'title')
clean_title = sanitize_filename(title[0][0]) if title else "chinese_book"
Path(output_path).mkdir(parents=True, exist_ok=True)
return str(Path(output_path) / f"{clean_title}_chinese.epub")
+164
View File
@@ -0,0 +1,164 @@
"""
EPUB 解析器模块 (EPUB Parser Module)
该模块负责读取 EPUB 文件,提取元数据和内容项目。
它使用 ebooklib 库来处理 EPUB 格式的底层细节。
Classes:
EPUBParser: 负责 EPUB 文件的加载、元数据提取和内容项遍历。
"""
import ebooklib
from ebooklib import epub
from bs4 import BeautifulSoup
from typing import List, Dict, Any
from pathlib import Path
from loguru import logger
class EPUBParser:
"""
EPUB 文件解析器。
负责加载 EPUB 文件,提取书籍元数据(如标题、作者),并提供方法来遍历和提取
书中的文档内容(HTML/XHTML)。
Attributes:
epub_path (Path): EPUB 文件的路径对象。
book (epub.EpubBook): ebooklib 加载的书籍对象。
metadata (Dict[str, str]): 提取的书籍元数据字典。
"""
def __init__(self, epub_path: str):
"""
初始化 EPUB 解析器。
Args:
epub_path (str): EPUB 文件的文件路径。
Raises:
FileNotFoundError: 如果指定的文件不存在。
Exception: 如果 EPUB 文件加载失败(格式错误等)。
"""
self.epub_path = Path(epub_path)
if not self.epub_path.exists():
raise FileNotFoundError(f"EPUB 文件不存在: {epub_path}")
try:
# ignore_ncx=True 是为了避免某些旧版 epub 的警告,但新版 ebooklib 可能行为不同
# 这里直接读取,让 ebooklib 处理
self.book = epub.read_epub(str(self.epub_path))
logger.info(f"成功加载 EPUB: {self.epub_path.name}")
except Exception as e:
logger.error(f"加载 EPUB 失败: {e}")
raise
self.metadata = self._extract_metadata()
def _extract_metadata(self) -> Dict[str, str]:
"""
从 EPUB 对象中提取标准元数据。
提取 Dublin Core (DC) 元数据,包括标题、作者和语言。
Returns:
Dict[str, str]: 包含 'title', 'author', 'language' 的字典。
如果提取失败,会使用默认值 ("Unknown", "en")。
"""
metadata = {}
try:
# get_metadata 返回的是 (value, dict) 的列表,我们取第一个结果
title_meta = self.book.get_metadata('DC', 'title')
metadata['title'] = title_meta[0][0] if title_meta else "Unknown"
author_meta = self.book.get_metadata('DC', 'creator')
metadata['author'] = author_meta[0][0] if author_meta else "Unknown"
lang_meta = self.book.get_metadata('DC', 'language')
metadata['language'] = lang_meta[0][0] if lang_meta else "en"
logger.info(f"书籍: {metadata['title']} - {metadata['author']}")
except Exception as e:
logger.warning(f"提取元数据时出错: {e}")
# 设置保底值
metadata.setdefault('title', 'Unknown')
metadata.setdefault('author', 'Unknown')
metadata.setdefault('language', 'en')
return metadata
def extract_all_content_items(self) -> List[Dict[str, Any]]:
"""
提取所有可翻译的内容项目(文档)。
遍历 EPUB 中的所有 Item,筛选出类型为 ITEM_DOCUMENT 的项目。
同时会进行简单的过滤,跳过内容过短(<100字符)或看起来像非正文的文件(如 nav, toc, cover)。
Returns:
List[Dict[str, Any]]: 内容项目列表。每个字典包含:
- item (epub.EpubItem): 原始 Item 对象。
- file_name (str): 文件名。
- content (str): 解码后的 HTML 内容。
- text_length (int): 纯文本长度(用于统计)。
"""
content_items = []
# 获取所有文档类型的项目
for item in self.book.get_items():
if item.get_type() == ebooklib.ITEM_DOCUMENT:
try:
# 获取内容 (bytes -> str)
content = item.get_content().decode('utf-8')
# 简单的内容验证:提取纯文本检查长度
soup = BeautifulSoup(content, 'html.parser')
text = soup.get_text().strip()
# 1. 跳过太短的内容(可能是只有图片的页面、空页面)
if len(text) < 100:
logger.debug(f"跳过短内容: {item.get_name()} ({len(text)} 字符)")
continue
# 2. 跳过明显的非正文内容 (根据文件名判断)
name_lower = item.get_name().lower()
skip_patterns = ['cover', 'copyright', 'titlepage', 'halftitle',
'nav.xhtml', 'toc.xhtml']
if any(pattern in name_lower for pattern in skip_patterns):
logger.debug(f"跳过非正文内容: {item.get_name()}")
continue
content_items.append({
'item': item,
'file_name': item.get_name(),
'content': content,
'text_length': len(text)
})
logger.debug(f"添加内容项: {item.get_name()} ({len(text)} 字符)")
except Exception as e:
logger.warning(f"处理项目失败 {item.get_name()}: {e}")
continue
logger.info(f"提取了 {len(content_items)} 个内容项目")
return content_items
def get_book_info(self) -> Dict[str, str]:
"""
获取书籍的摘要信息。
Returns:
Dict[str, str]: 包含文件名、标题、作者、语言和文档数量的字典。
"""
# 统计内容项
document_count = sum(1 for item in self.book.get_items()
if item.get_type() == ebooklib.ITEM_DOCUMENT)
return {
'filename': self.epub_path.name,
'title': self.metadata.get('title', 'Unknown'),
'author': self.metadata.get('author', 'Unknown'),
'language': self.metadata.get('language', 'en'),
'document_count': document_count
}
+305
View File
@@ -0,0 +1,305 @@
"""
格式提取模块 (优化版 v3)
核心优化:
1. 前缀/后缀标签分离:文本前后的标签不发送给 LLM,直接回填
2. 公式检测:将数学变量/公式作为整体占位符
3. 简化占位符:φ1φ 格式,每个段落独立编号
"""
import re
from bs4 import BeautifulSoup, Tag, NavigableString
from typing import Tuple, Dict, Any, List, Optional
class HeadingDetector:
"""标题与段落类型检测器"""
CHAPTER_PATTERNS = [
r'^(chapter|chap\.?|part)\s+([0-9]+|[ivxlc]+|[a-z])',
r'^(第\s*[0-9一二三四五六七八九十百]+\s*[章节部篇])',
r'^(\d+|[IVXLC]+|[A-Z])\.$'
]
EPIGRAPH_CLASSES = {
'epigraph', 'quote', 'blockquote', 'motto',
'dedication', 'verse', 'poetry', 'poem'
}
def detect(self, element: Tag, text: str) -> str:
if self._is_epigraph(element):
return "epigraph"
tag_name = element.name.lower()
if tag_name in ['h1', 'h2']:
return "chapter" if self._matches_chapter_pattern(text) else "section"
if tag_name == 'h3':
return "section"
if tag_name in ['h4', 'h5', 'h6']:
return "subsection"
if self._is_pseudo_heading(element, text):
return "subsection"
return "body"
def _is_epigraph(self, element: Tag) -> bool:
if element.name == 'blockquote':
return True
current = element
for _ in range(3):
if not current: break
classes = current.get('class', [])
if isinstance(classes, list):
classes = ' '.join(classes)
if any(k in classes.lower() for k in self.EPIGRAPH_CLASSES):
return True
current = current.parent
return False
def _matches_chapter_pattern(self, text: str) -> bool:
text = text.strip().lower()
for pattern in self.CHAPTER_PATTERNS:
if re.match(pattern, text, re.IGNORECASE):
return True
return False
def _is_pseudo_heading(self, element: Tag, text: str) -> bool:
if element.name != 'p':
return False
text = text.strip()
if not text or len(text) > 80:
return False
children = list(element.children)
if len(children) == 1 and isinstance(children[0], Tag):
if children[0].name in ['strong', 'b']:
return True
return False
class FormatExtractor:
"""
HTML 格式提取器 (优化版 v3)
核心改进:
1. 前缀/后缀标签分离 - 不发送给 LLM,自动回填
2. 公式元素整体替换
3. 简化占位符格式 φ1φ, φ2φ
"""
FORMULA_CHARS = re.compile(
r'^[\d\s\+\-\×\÷\=\(\)\[\]\{\}\<\>\^\*\/\.\,\;\:\'\"\`\~\@\#\$\%\&\|\\'
r'αβγδεζηθικλμνξοπρστυφχψωΑΒΓΔΕΖΗΘΙΚΛΜΝΞΟΠΡΣΤΥΦΧΨΩ'
r'a-zA-Z]+$'
)
def __init__(self):
self.detector = HeadingDetector()
def extract(self, element_html: str) -> Tuple[str, str, Dict[str, str], str]:
"""
提取格式信息
Returns:
clean_text: 纯文本
text_with_placeholders: 只包含内嵌占位符的文本(不含前缀/后缀标签)
placeholder_map: 占位符映射,包含特殊键 "_prefix""_suffix"
paragraph_type: 段落类型
"""
soup = BeautifulSoup(element_html, 'html.parser')
root = list(soup.children)[0] if list(soup.children) else soup
# 获取纯文本和段落类型
clean_text = root.get_text().strip()
clean_text = re.sub(r'\s+', ' ', clean_text)
p_type = self.detector.detect(root, clean_text) if isinstance(root, Tag) else "body"
# 获取内部 HTML
inner_html = root.decode_contents() if isinstance(root, Tag) else str(root)
# 智能提取(分离前缀/后缀)
text_with_ph, local_map = self._smart_extract_v3(inner_html)
return clean_text, text_with_ph, local_map, p_type
def _smart_extract_v3(self, inner_html: str) -> Tuple[str, Dict[str, str]]:
"""
智能提取 v3:分离前缀/后缀 + 合并内嵌公式块
核心逻辑:
1. 分离前缀(第一个可翻译文本之前)和后缀(最后一个可翻译文本之后)
2. 中间部分:检测"公式块"(连续标签+不可翻译文本),合并为单个占位符
3. 只有真正需要翻译的格式标签(如斜体包裹的长文本)才拆分
"""
# 使用正则分割标签和文本
parts = re.split(r'(<[^>]+>)', inner_html)
parts = [p for p in parts if p]
if not parts:
return "", {"_prefix": "", "_suffix": ""}
# 识别每个部分的类型
part_types = [] # 'tag', 'translatable', 'formula', 'whitespace'
for part in parts:
if part.startswith('<'):
part_types.append('tag')
elif not part.strip():
part_types.append('whitespace')
elif self._is_translatable_text(part):
part_types.append('translatable')
else:
part_types.append('formula')
# 找到第一个和最后一个可翻译文本的索引
first_trans_idx = None
last_trans_idx = None
for i, t in enumerate(part_types):
if t == 'translatable':
if first_trans_idx is None:
first_trans_idx = i
last_trans_idx = i
if first_trans_idx is None:
# 没有可翻译文本,全部作为前缀
return "", {"_prefix": inner_html, "_suffix": ""}
# 分割
prefix_parts = parts[:first_trans_idx]
middle_parts = parts[first_trans_idx:last_trans_idx + 1]
middle_types = part_types[first_trans_idx:last_trans_idx + 1]
suffix_parts = parts[last_trans_idx + 1:]
# 构建映射
local_map = {}
# 前缀
prefix_html = "".join(prefix_parts)
if prefix_html:
local_map["_prefix"] = prefix_html
# 后缀
suffix_html = "".join(suffix_parts)
if suffix_html:
local_map["_suffix"] = suffix_html
# 中间部分处理:使用配对占位符格式
#
# 策略:
# 1. 连续的 (tag|formula|whitespace) 不包含可翻译文本 → 合并为单个占位符 φ1φ
# 2. 开始标签后接可翻译文本 → 配对格式 φ2φ文本φ/2φ
#
placeholder_counter = 1
result_parts = []
tag_stack = [] # 追踪开放标签 [(id, opening_tag), ...]
i = 0
while i < len(middle_parts):
part = middle_parts[i]
ptype = middle_types[i]
if ptype == 'translatable':
# 可翻译文本,直接保留
result_parts.append(part)
i += 1
elif ptype == 'tag':
# 检查是开始标签还是结束标签
is_closing = part.startswith('</')
if is_closing:
# 结束标签
if tag_stack:
# 匹配最近的开放标签
open_id, open_tag = tag_stack.pop()
local_map[f"/{open_id}"] = part
result_parts.append(f"φ/{open_id}φ")
else:
# 没有匹配的开放标签,作为单独占位符
pid = str(placeholder_counter)
placeholder_counter += 1
local_map[pid] = part
result_parts.append(f"φ{pid}φ")
i += 1
else:
# 开始标签,检查后面是否有可翻译文本
has_translatable_after = False
for j in range(i + 1, len(middle_parts)):
if middle_types[j] == 'translatable':
has_translatable_after = True
break
elif middle_types[j] == 'tag' and middle_parts[j].startswith('</'):
# 遇到结束标签但还没遇到可翻译文本
break
if has_translatable_after:
# 开始配对模式
pid = str(placeholder_counter)
placeholder_counter += 1
local_map[pid] = part
result_parts.append(f"φ{pid}φ")
tag_stack.append((pid, part))
i += 1
else:
# 后面没有可翻译文本,合并为公式块
block_parts = []
while i < len(middle_parts) and middle_types[i] != 'translatable':
block_parts.append(middle_parts[i])
i += 1
if block_parts:
block_html = "".join(block_parts)
pid = str(placeholder_counter)
placeholder_counter += 1
local_map[pid] = block_html
result_parts.append(f"φ{pid}φ")
elif ptype in ('formula', 'whitespace'):
# 公式或空白,检查是否是连续块的开始
block_parts = []
while i < len(middle_parts) and middle_types[i] in ('formula', 'whitespace'):
block_parts.append(middle_parts[i])
i += 1
if block_parts:
block_html = "".join(block_parts)
pid = str(placeholder_counter)
placeholder_counter += 1
local_map[pid] = block_html
result_parts.append(f"φ{pid}φ")
else:
i += 1
text_with_ph = "".join(result_parts)
text_with_ph = re.sub(r'\s+', ' ', text_with_ph).strip()
return text_with_ph, local_map
def _is_translatable_text(self, text: str) -> bool:
"""判断文本是否需要翻译(包含可翻译的单词)"""
text = text.strip()
if not text:
return False
# 如果包含 4 个及以上连续字母,视为可翻译
if re.search(r'[a-zA-Z]{4,}', text):
return True
return False
def _is_formula_element(self, element: Tag, text_content: str) -> bool:
"""判断元素是否为公式元素(应整体保留不翻译)"""
if not text_content:
return True
if len(text_content) <= 3:
return True
if re.search(r'[a-zA-Z]{4,}', text_content):
return False
return bool(self.FORMULA_CHARS.match(text_content))
def _get_opening_tag(self, element: Tag) -> str:
"""获取元素的开始标签(含属性)"""
attrs_str = ""
for key, value in element.attrs.items():
if isinstance(value, list):
value = " ".join(value)
attrs_str += f' {key}="{value}"'
return f"<{element.name}{attrs_str}>"
def reset(self):
"""兼容旧接口"""
pass
+89
View File
@@ -0,0 +1,89 @@
"""
格式恢复模块 (优化版 v4)
负责:
1. 解析译文中的配对占位符 (φ1φ...φ/1φ)
2. 还原前缀和后缀标签(_prefix, _suffix
3. 从映射表中查找对应的 HTML 片段并替换
"""
import re
from typing import Dict, Tuple, List, Optional
from loguru import logger
class FormatRestorer:
"""
格式恢复器 (优化版 v4)
支持:
- 配对占位符 φ1φ...φ/1φ
- 前缀/后缀自动回填 (_prefix, _suffix)
"""
# 占位符正则: φ1φ, φ/1φ, φ12φ, φ/12φ (支持配对格式)
PLACEHOLDER_REGEX = re.compile(r'φ(/?\d+)φ')
def restore(self, text_with_placeholders: str, placeholder_map: Dict[str, str]) -> Tuple[str, bool]:
"""
将带占位符的文本还原为 HTML
自动处理 _prefix 和 _suffix 键,以及配对占位符 φ1φ...φ/1φ
Returns:
(html_string, success): 还原后的 HTML 和是否完全成功的标志
"""
if not placeholder_map:
return text_with_placeholders or "", True
if not text_with_placeholders:
# 没有文本,只有前缀/后缀(如空元素)
prefix = placeholder_map.get("_prefix", "")
suffix = placeholder_map.get("_suffix", "")
return prefix + suffix, True
# 提取前缀和后缀
prefix = placeholder_map.get("_prefix", "")
suffix = placeholder_map.get("_suffix", "")
# 创建只包含占位符键的映射(排除 _prefix, _suffix
inner_map = {k: v for k, v in placeholder_map.items() if not k.startswith("_")}
# 校验占位符
found_ids = set(self.PLACEHOLDER_REGEX.findall(text_with_placeholders))
expected_ids = set(inner_map.keys())
success = True
missing_ids = expected_ids - found_ids
if missing_ids:
logger.warning(f"格式还原警告: 丢失占位符 {missing_ids}")
success = False
unknown_ids = found_ids - expected_ids
if unknown_ids:
logger.warning(f"格式还原警告: 发现未知占位符 {unknown_ids}")
success = False # 未知占位符也视为问题
# 替换占位符
def replace_match(match):
pid = match.group(1) # 可能是 "1" 或 "/1"
if pid in inner_map:
return inner_map[pid]
else:
return "" # 删除未知占位符
try:
# 还原占位符
restored_inner = self.PLACEHOLDER_REGEX.sub(replace_match, text_with_placeholders)
# 添加前缀和后缀
restored_html = prefix + restored_inner + suffix
return restored_html, success
except Exception as e:
logger.error(f"格式还原失败: {e}")
return prefix + self._strip_placeholders(text_with_placeholders) + suffix, False
def _strip_placeholders(self, text: str) -> str:
"""移除所有 φ...φ 占位符"""
return self.PLACEHOLDER_REGEX.sub("", text)
+241
View File
@@ -0,0 +1,241 @@
"""
LLM Client Module - Generic OpenAI Compatible
Features:
1. Fully configurable via config.json (base_url, headers).
2. Mode-aware prompt building (bilingual vs chinese).
3. Format repair capability for chinese mode.
"""
import asyncio
import json
import re
from openai import AsyncOpenAI
from typing import List, Dict, Optional, Any
from loguru import logger
import time
from tenacity import retry, stop_after_attempt, wait_exponential
from .manifest_manager import ManifestItem
class RateLimiter:
"""Rate limiter for concurrency and RPM."""
def __init__(self, requests_per_minute: int, concurrent_requests: int):
self.semaphore = asyncio.Semaphore(concurrent_requests)
self.min_interval = 60.0 / requests_per_minute if requests_per_minute > 0 else 0
self.last_request_time = 0
self._lock = asyncio.Lock()
async def acquire(self):
await self.semaphore.acquire()
async with self._lock:
current_time = time.time()
wait_time = self.min_interval - (current_time - self.last_request_time)
if wait_time > 0:
await asyncio.sleep(wait_time)
self.last_request_time = time.time()
def release(self):
self.semaphore.release()
class LLMClient:
"""Generic OpenAI-compatible API Client."""
def __init__(self, config: Dict):
self.config = config
llm_config = config["llm"]
api_key = llm_config.get("api_key")
base_url = llm_config.get("base_url")
extra_headers = llm_config.get("extra_headers", {})
if not api_key:
raise ValueError("API Key is missing in config")
self.client = AsyncOpenAI(
base_url=base_url,
api_key=api_key,
default_headers=extra_headers
)
self.models = llm_config.get("models", {"fast": "gpt-3.5-turbo", "smart": "gpt-4"})
self.rate_limiter = RateLimiter(
llm_config["rate_limits"]["requests_per_minute"],
llm_config["rate_limits"]["concurrent_requests"]
)
self.prompts = self._load_prompts()
def _load_prompts(self) -> Dict:
try:
with open("config/prompts.json", "r", encoding="utf-8") as f:
return json.load(f)
except:
return {}
async def translate_chunk(self, items: List[ManifestItem], glossary: Dict = None,
instruction: str = None, model_type: str = "fast",
mode: str = "bilingual") -> Dict[str, str]:
"""
Translate a chunk of items.
Args:
items: List of ManifestItem to translate
glossary: Term dictionary
instruction: Style guide
model_type: "fast" or "smart"
mode: "bilingual" or "chinese"
"""
if not items: return {}
model = self.models.get(model_type, self.models.get("fast"))
prompt = self._build_prompt(items, mode)
try:
# Build System Prompt
base_sys_prompt = self.prompts.get("translation", {}).get("system", "You are a professional translator.")
# 中文模式:添加占位符保护指令
if mode == "chinese":
base_sys_prompt += """
Placeholder Instructions (CRITICAL):
1. Text contains PAIRED placeholders: φNφ (start) and φ/Nφ (end), like HTML tags.
2. Example: "φ1φTable Talkφ/1φ" means italic text, translate as "φ1φ桌谈φ/1φ"
3. Single placeholders φNφ without φ/Nφ are inline elements (footnotes, formulas) - keep them in place.
4. RULES:
- DO NOT create new placeholder numbers that don't exist in the original
- DO NOT remove or modify existing placeholders
- Keep placeholders in the SAME relative position in your translation
- If word order changes, keep placeholders with their associated text
5. Each line starts with paragraph ID (p_xxxxx). Preserve them.
"""
if instruction:
base_sys_prompt += f"\n\nBook Style Guide:\n{instruction}"
if glossary:
glossary_text = "\n".join([f"{k} -> {v}" for k, v in glossary.items()])
base_sys_prompt += f"\n\nTerminology:\n{glossary_text}"
# Strict formatting instructions
base_sys_prompt += "\n\nRequirements:\n1. Each line MUST start with ID (p_xxxxx).\n2. DO NOT modify IDs.\n3. Return only translations."
raw_response = await self._make_request(model, base_sys_prompt, prompt)
if not raw_response:
return {item.global_id: f"[Error - Empty Response]" for item in items}
return self._simple_parse(raw_response, items, mode)
except Exception as e:
logger.error(f"Translation failed ({model}): {e}")
return {item.global_id: f"[Error - {str(e)}]" for item in items}
async def repair_format(self, original_text: str, broken_translation: str) -> str:
"""
修复翻译格式:将占位符正确插入到译文中。
"""
model = self.models.get("fast")
system_prompt = "You are a format repair assistant. Your ONLY job is to insert placeholders into the translation."
user_prompt = f"""
Original Text (with placeholders):
{original_text}
Translation (placeholders missing/incorrect):
{broken_translation}
Task:
Please rewrite the Translation to include ALL placeholders (φcXXXXXφ) from the Original Text.
1. DO NOT translate again. Keep the meaning of the Translation.
2. Place φcXXXXXφ tags exactly where they correspond to the original format (bold, italic, links).
3. Output ONLY the fixed translation.
"""
try:
return await self._make_request(model, system_prompt, user_prompt)
except Exception as e:
logger.error(f"Format repair failed: {e}")
return broken_translation
async def raw_chat_completion(self, system_prompt: str, user_prompt: str, model_type: str = "smart") -> str:
"""Generic chat completion (for Profiler)."""
model = self.models.get(model_type, self.models.get("smart"))
return await self._make_request(model, system_prompt, user_prompt)
def _build_prompt(self, items: List[ManifestItem], mode: str = "bilingual") -> str:
"""构建翻译提示词"""
lines = []
for item in items:
if mode == "chinese":
# 中文模式:使用带占位符的文本和段落类型
text = item.text_with_placeholders if item.text_with_placeholders else item.clean_text
p_type = getattr(item, 'paragraph_type', 'body').upper()
lines.append(f"{item.global_id} [{p_type}] {text}")
else:
# 双语模式:使用纯文本
lines.append(f"{item.global_id} {item.clean_text}")
return "\n".join(lines)
def _simple_parse(self, response: str, items: List[ManifestItem], mode: str = "bilingual") -> Dict[str, str]:
"""解析 LLM 响应"""
results = {}
for i, item in enumerate(items):
current_id = item.global_id
start_idx = response.find(current_id)
if start_idx == -1: continue
end_idx = len(response)
if i + 1 < len(items):
next_id = items[i+1].global_id
next_found = response.find(next_id, start_idx + len(current_id))
if next_found != -1:
end_idx = next_found
content = response[start_idx:end_idx].strip()
clean_content = content[len(current_id):].strip()
clean_content = clean_content.lstrip(": \t")
# 移除类型标记 (如 [BODY])
if mode == "chinese":
clean_content = re.sub(r'^\[[A-Z]+\]\s*', '', clean_content)
if clean_content:
results[current_id] = clean_content
# Fallback: 逐行解析
if len(results) < len(items):
for line in response.split("\n"):
line = line.strip()
for item in items:
if item.global_id not in results and line.startswith(item.global_id):
res = line[len(item.global_id):].strip().lstrip(": ")
if mode == "chinese":
res = re.sub(r'^\[[A-Z]+\]\s*', '', res)
if res: results[item.global_id] = res
return results
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
async def _make_request(self, model: str, system_prompt: str, user_prompt: str) -> str:
await self.rate_limiter.acquire()
try:
resp = await self.client.chat.completions.create(
model=model,
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_prompt}
],
temperature=self.config['translation'].get('temperature', 0.2),
max_tokens=8000
)
return resp.choices[0].message.content.strip()
finally:
self.rate_limiter.release()
async def close(self):
await self.client.close()
# Alias for backward compatibility
OpenRouterClient = LLMClient
+166
View File
@@ -0,0 +1,166 @@
"""
Manifest 管理器模块 (Manifest Manager Module)
该模块是系统的单一真理源 (SSOT)。
它记录了每一段文本的原始状态、清洗后的文本、哈希值以及翻译状态。
所有对翻译流程的操作(提取、翻译、回填)都必须通过修改此 Manifest 进行。
"""
import json
import os
import hashlib
from typing import List, Dict, Optional, Any
from pathlib import Path
from loguru import logger
from dataclasses import dataclass, asdict, field
@dataclass
class ManifestItem:
"""代表一个翻译单元(通常是一个段落)"""
global_id: str
source_file: str
original_html: str
clean_text: str
text_hash: str
tag: str
tag_attrs: Dict[str, Any] = field(default_factory=dict) # 外层标签的属性 (class, style...)
translation: Optional[str] = None
status: str = "pending" # pending, translated, ignored, failed
error_msg: Optional[str] = None
model_used: Optional[str] = None # 记录使用的模型
quality_score: Optional[int] = None # 记录质量评分
# === 中文模式专用字段 ===
text_with_placeholders: str = "" # 带占位符的文本
placeholder_map: Dict[str, str] = field(default_factory=dict) # 占位符映射表 {id: html_string}
paragraph_type: str = "body" # 段落类型:chapter/section/subsection/epigraph/body
translation_with_placeholders: str = "" # 带占位符的译文
translation_with_original_html: str = "" # 还原后的最终 HTML (中文模式)
metadata: Dict[str, Any] = field(default_factory=dict)
def to_dict(self):
return asdict(self)
class ManifestManager:
"""
负责 Manifest 的生命周期管理。
"""
def __init__(self, manifest_path: str):
self.manifest_path = Path(manifest_path)
self.data: Dict[str, Any] = {
"book_id": "",
"metadata": {},
"items": []
}
self._items_by_id: Dict[str, ManifestItem] = {}
def load(self) -> bool:
"""从文件加载 Manifest。如果文件不存在则返回 False。"""
if self.manifest_path.exists():
try:
with open(self.manifest_path, 'r', encoding='utf-8') as f:
self.data = json.load(f)
# 重建对象映射
self._items_by_id = {
item['global_id']: ManifestItem(**item)
for item in self.data["items"]
}
logger.info(f"成功从 {self.manifest_path} 加载 Manifest, 包含 {len(self._items_by_id)} 个项目")
return True
except Exception as e:
logger.error(f"加载 Manifest 失败: {e}")
return False
return False
def save(self):
"""将当前状态保存到 Manifest 文件。"""
# 确保目录存在
self.manifest_path.parent.mkdir(parents=True, exist_ok=True)
# 同步 items 到 data 字典
self.data["items"] = [item.to_dict() for item in self._items_by_id.values()]
with open(self.manifest_path, 'w', encoding='utf-8') as f:
json.dump(self.data, f, ensure_ascii=False, indent=2)
# logger.debug(f"Manifest 已保存到 {self.manifest_path}")
def init_manifest(self, book_id: str, metadata: Dict):
"""初始化一个新的 Manifest。"""
self.data = {
"book_id": book_id,
"metadata": metadata,
"items": []
}
self._items_by_id = {}
self.save()
def add_item(self, source_file: str, original_html: str, clean_text: str, tag: str, metadata: Dict = None) -> ManifestItem:
"""添加一个新的翻译项并分配 ID。"""
# 生成全局 ID
new_index = len(self._items_by_id) + 1
global_id = f"p_{new_index:05d}"
# 生成内容哈希 (用于排重和缓存)
text_hash = hashlib.sha256(clean_text.encode('utf-8')).hexdigest()
item = ManifestItem(
global_id=global_id,
source_file=source_file,
original_html=original_html,
clean_text=clean_text,
text_hash=text_hash,
tag=tag,
metadata=metadata or {}
)
self._items_by_id[global_id] = item
return item
def get_items(self, status: str = None, file_name: str = None) -> List[ManifestItem]:
"""按状态或文件名查询项目。"""
items = list(self._items_by_id.values())
if status:
items = [i for i in items if i.status == status]
if file_name:
items = [i for i in items if i.source_file == file_name]
# 必须按 ID 顺序返回以保证分块正确
return sorted(items, key=lambda x: x.global_id)
def update_item(self, global_id: str, translation: str, status: str = "translated", error: str = None, model: str = None, score: int = None):
"""更新翻译结果。"""
if global_id in self._items_by_id:
item = self._items_by_id[global_id]
if translation is not None:
item.translation = translation
item.status = status
if error:
item.error_msg = error
if model:
item.model_used = model
if score is not None:
item.quality_score = score
else:
logger.warning(f"尝试更新不存在的 ID: {global_id}")
@property
def stats(self) -> Dict:
"""获取翻译进度统计。"""
total = len(self._items_by_id)
if total == 0: return {"progress": "0%"}
translated = sum(1 for i in self._items_by_id.values() if i.status == "translated")
ignored = sum(1 for i in self._items_by_id.values() if i.status == "ignored")
failed = sum(1 for i in self._items_by_id.values() if i.status == "failed")
return {
"total": total,
"translated": translated,
"ignored": ignored,
"failed": failed,
"pending": total - translated - ignored - failed,
"progress_percent": round((translated + ignored) / total * 100, 1)
}
+87
View File
@@ -0,0 +1,87 @@
"""
Quality Manager Module
Responsible for evaluating translation quality and deciding on re-translation.
"""
import json
import random
from typing import List, Dict, Any, Tuple
from loguru import logger
from .manifest_manager import ManifestItem
from .llm_client import LLMClient
class QualityManager:
def __init__(self, config: Dict, llm_client: LLMClient):
self.config = config
self.llm_client = llm_client
self.qc_config = config['translation'].get('quality_control', {})
self.pass_score = self.qc_config.get('pass_score', 7)
self.sample_size = self.qc_config.get('sample_size', 2)
async def evaluate_chunk(self, chunk: List[ManifestItem]) -> Tuple[bool, int, str]:
"""
Evaluate a chunk of translations.
Returns:
(passed: bool, average_score: int, reason: str)
"""
if not self.qc_config.get('enabled', False):
return True, 10, "QC Disabled"
# 1. Sample items
# Filter for items that actually have content and translations
valid_items = [item for item in chunk if item.translation and len(item.clean_text) > 20]
if not valid_items:
return True, 10, "No valid items to sample"
sample_items = random.sample(valid_items, min(len(valid_items), self.sample_size))
# 2. Build Prompt
prompt = self._build_evaluation_prompt(sample_items)
# 3. Call LLM (Smart)
try:
response = await self.llm_client.raw_chat_completion(
system_prompt="You are a professional translation editor.",
user_prompt=prompt,
model_type="smart"
)
# 4. Parse JSON
# Clean potential markdown
json_str = response.strip()
if "```json" in json_str:
json_str = json_str.split("```json")[1].split("```")[0].strip()
elif "```" in json_str:
json_str = json_str.split("```")[1].split("```")[0].strip()
result = json.loads(json_str)
score = result.get('score', 0)
reason = result.get('reason', 'No reason provided')
passed = score >= self.pass_score
return passed, score, reason
except Exception as e:
logger.error(f"QC evaluation failed: {e}")
# If QC fails, we default to PASS to avoid blocking progress, but log it
return True, 0, f"QC Error: {e}"
def _build_evaluation_prompt(self, items: List[ManifestItem]) -> str:
content = ""
for i, item in enumerate(items, 1):
content += f"Item {i}:\nOriginal: {item.clean_text}\nTranslation: {item.translation}\n\n"
return f"""Please evaluate the following translations (English to Chinese).
Focus on accuracy, fluency, and terminology consistency.
Items to evaluate:
{content}
Return a JSON object with:
- \"score\": An integer from 1 to 10 (10 being perfect).
- \"reason\": A brief explanation of the score.
JSON Output:"""
+188
View File
@@ -0,0 +1,188 @@
"""
文本处理器模块 (Text Processor Module) - Manifest 驱动版
该模块专注于 HTML 文档的遍历和段落提取。
它不再维护全局状态,而是将提取的内容注册到 ManifestManager 中。
"""
import re
from bs4 import BeautifulSoup
from typing import List, Dict, Any
from loguru import logger
from .manifest_manager import ManifestManager
from .format_extractor import FormatExtractor
class TextProcessor:
"""
负责从 HTML 中识别有效段落并进行清洗。
"""
def __init__(self, config: Dict):
"""
Args:
config (Dict): 全局配置。
"""
self.config = config
self.chunk_size = config['translation'].get('chunk_size', 5000)
self.format_extractor = FormatExtractor()
def extract_to_manifest(self, html_content: str, source_file: str, manifest: ManifestManager, mode: str = "bilingual"):
"""
解析 HTML 内容,并将识别出的段落注册到 Manifest 中。
Args:
html_content (str): HTML 源码。
source_file (str): 来源文件名。
manifest (ManifestManager): 清单管理器实例。
mode (str): 翻译模式 - "bilingual""chinese"
"""
try:
soup = BeautifulSoup(html_content, 'html.parser')
# 1. 移除不需要的元素
for element in soup(['script', 'style', 'meta', 'link']):
element.decompose()
# 2. 获取有效的文本元素 (使用静态过滤逻辑)
text_elements = self.get_valid_text_elements(soup)
# 3. 注册到 Manifest
for element in text_elements:
clean_text = self.clean_element_text(element)
# 过滤逻辑
if not clean_text:
continue
status = "pending"
# 如果是导航元素,标记为 ignored
if self.is_navigation_element(element):
status = "ignored"
# 提取格式信息(中文模式)
text_with_ph = ""
placeholder_map = None
p_type = "body"
if mode == "chinese":
# FormatExtractor 接受 HTML 字符串
clean_text, text_with_ph, placeholder_map, p_type = self.format_extractor.extract(str(element))
# 注册
item = manifest.add_item(
source_file=source_file,
original_html=str(element),
clean_text=clean_text,
tag=element.name,
metadata={"status": status}
)
item.tag_attrs = element.attrs # 存储外层标签属性
# 更新中文模式字段
if mode == "chinese":
item.text_with_placeholders = text_with_ph
item.placeholder_map = placeholder_map
item.paragraph_type = p_type
# 同步更新 manifest 状态 (如果需要过滤)
if status == "ignored":
manifest.update_item(item.global_id, translation=None, status="ignored")
except Exception as e:
logger.error(f"{source_file} 提取段落失败: {e}")
@staticmethod
def get_valid_text_elements(soup) -> List:
"""获取不含嵌套子块的叶子级文本容器元素。"""
tags = ['p', 'div', 'h1', 'h2', 'h3', 'h4', 'h5', 'h6', 'blockquote', 'li', 'td']
all_candidates = soup.find_all(tags)
candidate_set = set(all_candidates)
final_elements = []
for element in all_candidates:
# 如果包含其他候选标签,说明是容器,跳过
if any(d in candidate_set for d in element.find_all(tags)):
continue
final_elements.append(element)
return final_elements
@staticmethod
def clean_element_text(element) -> str:
"""清理 HTML 元素,提取纯净的待翻译文本。"""
element_copy = element.__copy__()
# 移除脚注引用等
for tag in element_copy.find_all(['sup', 'sub']):
tag.decompose()
footnote_patterns = re.compile(r'footnote|endnote|reference|note|super|sub', re.I)
for tag in element_copy.find_all(['a', 'span', 'div'], class_=footnote_patterns):
tag.decompose()
# 移除仅包含数字的 span
for tag in element_copy.find_all('span'):
if re.match(r'^(\[\d+\]|\(\d+\)|\d+)$', tag.get_text().strip()):
tag.decompose()
text = element_copy.get_text().strip()
# 正则清理残留引用标识 (如 sentence.2)
text = re.sub(r'(\.|。||,)\s*(\[\d+\]|\d+)(?=\s|$)', r'\1', text)
text = re.sub(r'\s+', ' ', text)
return text
@staticmethod
def is_navigation_element(element) -> bool:
"""判断是否是无翻译价值的导航、页码元素。"""
classes = element.get('class', [])
nav_classes = ['nav', 'navigation', 'toc', 'menu', 'header', 'footer', 'page-number']
class_str = ' '.join(classes).lower() if isinstance(classes, list) else str(classes).lower()
if any(nc in class_str for nc in nav_classes):
return True
# 检查父级
parent = element.parent
if parent:
p_classes = parent.get('class', [])
p_class_str = ' '.join(p_classes).lower() if isinstance(p_classes, list) else str(p_classes).lower()
if any(nc in p_class_str for nc in nav_classes):
return True
return False
def create_chunks_from_manifest(self, manifest: ManifestManager, mode: str = "bilingual") -> List[List[Any]]:
"""
从 Manifest 中筛选待翻译项目并分块。
Args:
manifest: ManifestManager 实例
mode: 翻译模式 (保留参数以供将来使用)
"""
pending_items = manifest.get_items(status="pending")
if not pending_items:
return []
chunks = []
current_chunk = []
current_size = 0
for item in pending_items:
# 中文模式使用带占位符的文本长度
if mode == "chinese" and item.text_with_placeholders:
text_len = len(item.text_with_placeholders)
else:
text_len = len(item.clean_text)
if current_size + text_len > self.chunk_size and current_chunk:
chunks.append(current_chunk)
current_chunk = []
current_size = 0
current_chunk.append(item)
current_size += text_len
if current_chunk:
chunks.append(current_chunk)
logger.info(f"分块完成: 共有 {len(pending_items)} 个待翻译项,分为 {len(chunks)} 个块")
return chunks
+248
View File
@@ -0,0 +1,248 @@
"""
EPUB Translator Core Module - v0.08 (Mode Support)
"""
import asyncio
import traceback
from typing import List, Dict, Any
from pathlib import Path
from loguru import logger
from rich.console import Console
from rich.progress import Progress, SpinnerColumn, TextColumn, BarColumn, TimeElapsedColumn
from .epub_parser import EPUBParser
from .llm_client import LLMClient
from .text_processor import TextProcessor
from .bilingual_builder import BilingualEPUBBuilder
from .chinese_builder import ChineseEPUBBuilder
from .manifest_manager import ManifestManager
from .book_profiler import BookProfiler
from .cache import TranslationCache
from .format_restorer import FormatRestorer
from .utils import add_spacing_between_cn_and_en_num
class EPUBTranslator:
def __init__(self, config: Dict, use_cache: bool = True):
self.config = config
self.console = Console()
self.use_cache = use_cache
self.parser = None
self.llm_client = LLMClient(config)
self.text_processor = TextProcessor(config)
self.profiler = BookProfiler(config, self.llm_client)
self.cache = TranslationCache(config) if use_cache else None
self.restorer = FormatRestorer()
self.manifest_dir = Path("cache/manifests")
self.manifest_dir.mkdir(parents=True, exist_ok=True)
async def translate_epub(self, epub_path: str, test_mode: bool = False,
output_dir: str = None, mode: str = "bilingual") -> str:
"""
翻译 EPUB 文件
Args:
epub_path: EPUB 文件路径
test_mode: 测试模式(只翻译前几块)
output_dir: 输出目录
mode: 翻译模式 - "bilingual" (双语) 或 "chinese" (纯中文)
"""
try:
epub_path = Path(epub_path)
self.parser = EPUBParser(str(epub_path))
# 1. Manifest - 根据 mode 使用不同的 manifest 文件
manifest_suffix = "_chinese" if mode == "chinese" else ""
manifest_path = self.manifest_dir / f"{epub_path.stem}{manifest_suffix}_manifest.json"
manifest = ManifestManager(str(manifest_path))
if not manifest.load() or not self.use_cache:
self.console.print(f"[yellow]Initializing Manifest (Mode: {mode})...[/yellow]")
manifest.init_manifest(book_id=epub_path.name, metadata=self.parser.get_book_info())
content_items = self.parser.extract_all_content_items()
for item in content_items:
# 传递 mode 参数
self.text_processor.extract_to_manifest(
item['content'], item['file_name'], manifest, mode=mode
)
manifest.save()
stats = manifest.stats
self.console.print(f"[green]Manifest loaded: {stats['total']} paragraphs[/green]")
if stats['pending'] < stats['total'] and stats['translated'] > 0:
self.console.print(
f"[yellow]Detected incomplete translation task, completed {stats['translated']}/{stats['total']}, "
f"continuing translation of remaining {stats['pending']} paragraphs...[/yellow]"
)
# 2. Profile (Glossary)
profile = {}
if not test_mode:
self.console.print("[yellow]Generating Book Profile...[/yellow]")
profile = await self.profiler.analyze_book(manifest)
self.console.print(f"Genre: {profile.get('genre')} | Style: {profile.get('style')}")
# 3. Translate - 传递 mode 参数
chunks = self.text_processor.create_chunks_from_manifest(manifest, mode=mode)
if test_mode:
self.console.print("[yellow]Test mode enabled: Translating only first 10 chunks...[/yellow]")
chunks = chunks[:10]
if chunks:
await self._translate_concurrently(chunks, manifest, profile, mode=mode)
# 4. Build - 根据 mode 选择正确的 Builder
self.console.print(f"\n[yellow]Building {mode} EPUB...[/yellow]")
output_path = output_dir or self.config['output']['output_dir']
if mode == "chinese":
builder = ChineseEPUBBuilder(self.parser.book, self.config)
result_file = builder.create_chinese_epub_with_mapping(
manifest.get_items(), output_path
)
else:
builder = BilingualEPUBBuilder(self.parser.book, self.config)
translation_map = {item.global_id: item.translation for item in manifest.get_items() if item.translation}
paragraph_map = {item.global_id: {
"file_name": item.source_file,
"text": item.clean_text,
"html_element": item.original_html,
"tag_attrs": item.tag_attrs
} for item in manifest.get_items()}
result_file = builder.create_bilingual_epub_with_mapping(
translation_map, paragraph_map, output_path
)
final_stats = manifest.stats
self.console.print(f"""
[green]✅ Translation complete![/green]
- Mode: {mode}
- Total Paragraphs: {final_stats['total']}
- Successfully Translated: {final_stats['translated']}
- Skipped: {final_stats['ignored']}
- Failed: {final_stats['failed']}
- Output File: {result_file}
""")
return result_file
except Exception as e:
traceback.print_exc()
logger.error(f"Translation flow failed: {e}")
raise
async def _translate_concurrently(self, chunks: List[List[Any]], manifest: ManifestManager,
profile: Dict, mode: str = "bilingual"):
total_chunks = len(chunks)
glossary = profile.get('glossary', {})
instruction = profile.get('translation_instruction', "")
with Progress(
SpinnerColumn(),
TextColumn("[progress.description]{task.description}"),
BarColumn(),
TextColumn("[progress.percentage]{task.percentage:>3.0f}%"),
TimeElapsedColumn(),
console=self.console
) as progress:
task_id = progress.add_task(f"[cyan]Translating ({mode})...", total=total_chunks)
async def worker(chunk, idx):
try:
model_name = self.llm_client.models.get('fast')
chunk_dicts = [item.to_dict() for item in chunk]
results = None
if self.cache:
results = self.cache.get_chunk_translation(chunk_dicts, model=model_name)
if not results:
# 传递 mode 参数给 LLM
results = await self.llm_client.translate_chunk(
chunk,
glossary=glossary,
instruction=instruction,
mode=mode
)
if self.cache and results:
self.cache.save_chunk_translation(chunk_dicts, results, model=model_name)
for item in chunk:
if item.global_id in results:
raw_translation = results[item.global_id]
# 错误检测:如果翻译结果包含错误标记,视为失败
if raw_translation.startswith("[Error") or "Error -" in raw_translation:
logger.warning(f"Translation error for {item.global_id}: {raw_translation}")
manifest.update_item(item.global_id, None, status="failed", error=raw_translation)
continue
# 格式化翻译文本(盘古之白)
processed_translation = add_spacing_between_cn_and_en_num(raw_translation)
if mode == "chinese":
# 检查是否有内嵌占位符(排除 _prefix, _suffix
inner_placeholders = {k: v for k, v in item.placeholder_map.items()
if not k.startswith("_")} if item.placeholder_map else {}
if not inner_placeholders:
# 没有内嵌占位符,直接使用译文(清除 LLM 可能虚构的占位符)
clean_translation = self.restorer._strip_placeholders(processed_translation)
item.translation_with_placeholders = clean_translation
# 还原时只添加前缀后缀
restored_html, success = self.restorer.restore(
clean_translation,
item.placeholder_map
)
else:
# 有内嵌占位符,正常还原
item.translation_with_placeholders = processed_translation
restored_html, success = self.restorer.restore(
processed_translation,
item.placeholder_map
)
# 格式修复逻辑
if not success:
logger.warning(f"格式丢失 (ID: {item.global_id}),尝试自动修复...")
try:
fixed_translation = await self.llm_client.repair_format(
item.text_with_placeholders,
processed_translation
)
restored_html_2, success_2 = self.restorer.restore(
fixed_translation,
item.placeholder_map
)
if success_2:
logger.info(f"格式修复成功! (ID: {item.global_id})")
restored_html = restored_html_2
item.translation_with_placeholders = fixed_translation
else:
logger.error(f"格式修复失败 (ID: {item.global_id}),保留原始译文")
except Exception as e:
logger.error(f"修复过程出错: {e}")
item.translation_with_original_html = restored_html
# translation 字段存储纯文本
item.translation = self.restorer._strip_placeholders(item.translation_with_placeholders)
manifest.update_item(item.global_id, item.translation)
else:
# 双语模式
item.translation = processed_translation
manifest.update_item(item.global_id, item.translation)
else:
manifest.update_item(item.global_id, None, status="failed", error="Missing")
manifest.save()
except Exception as e:
logger.error(f"Chunk {idx} failed: {e}")
finally:
progress.update(task_id, advance=1)
tasks = [worker(chunk, i) for i, chunk in enumerate(chunks)]
await asyncio.gather(*tasks)
+208
View File
@@ -0,0 +1,208 @@
"""
工具函数模块
提供配置加载、日志设置等通用功能
"""
import json
import os
from pathlib import Path
from typing import Dict, Any
from loguru import logger
import sys
from dotenv import load_dotenv
def load_config(config_path: str = "config/config.json") -> Dict[str, Any]:
"""
加载配置文件
Args:
config_path: 配置文件路径
Returns:
配置字典
"""
# 加载 .env 文件
load_dotenv()
try:
with open(config_path, 'r', encoding='utf-8') as f:
config = json.load(f)
# 从环境变量获取 API Key
for provider_name, provider_config in config.get('providers', {}).items():
env_key = f"{provider_name.upper()}_API_KEY"
if env_key in os.environ:
provider_config['api_key'] = os.environ[env_key]
return config
except FileNotFoundError:
raise FileNotFoundError(f"配置文件未找到: {config_path}")
except json.JSONDecodeError as e:
raise ValueError(f"配置文件格式错误: {e}")
def load_prompts(prompts_path: str = "config/prompts.json") -> Dict[str, str]:
"""
加载提示词模板
Args:
prompts_path: 提示词文件路径
Returns:
提示词字典
"""
try:
with open(prompts_path, 'r', encoding='utf-8') as f:
return json.load(f)
except FileNotFoundError:
raise FileNotFoundError(f"提示词文件未找到: {prompts_path}")
def setup_logging(config: Dict[str, Any]) -> None:
"""
设置日志配置
Args:
config: 配置字典
"""
log_config = config.get('logging', {})
# 移除默认处理器
logger.remove()
# 添加控制台输出
logger.add(
sys.stdout,
level=log_config.get('level', 'INFO'),
format="<green>{time:YYYY-MM-DD HH:mm:ss}</green> | <level>{level: <8}</level> | <cyan>{name}</cyan>:<cyan>{function}</cyan>:<cyan>{line}</cyan> - <level>{message}</level>"
)
# 添加文件输出
if 'file' in log_config:
log_file = log_config['file']
# 确保日志目录存在
Path(log_file).parent.mkdir(parents=True, exist_ok=True)
logger.add(
log_file,
level=log_config.get('level', 'INFO'),
rotation=log_config.get('rotation', '10 MB'),
retention=log_config.get('retention', '7 days'),
encoding='utf-8',
format="{time:YYYY-MM-DD HH:mm:ss} | {level: <8} | {name}:{function}:{line} - {message}"
)
def ensure_output_dir(output_dir: str) -> Path:
"""
确保输出目录存在
Args:
output_dir: 输出目录路径
Returns:
输出目录的 Path 对象
"""
output_path = Path(output_dir)
output_path.mkdir(parents=True, exist_ok=True)
return output_path
def sanitize_filename(filename: str) -> str:
"""
清理文件名,移除非法字符
Args:
filename: 原始文件名
Returns:
清理后的文件名
"""
import re
# 移除或替换非法字符
filename = re.sub(r'[<>:"/\\|?*]', '_', filename)
# 移除多余的空格和点
filename = re.sub(r'\s+', ' ', filename).strip('. ')
return filename
def format_file_size(size_bytes: int) -> str:
"""
格式化文件大小显示
Args:
size_bytes: 字节数
Returns:
格式化的大小字符串
"""
if size_bytes == 0:
return "0B"
size_names = ["B", "KB", "MB", "GB"]
import math
i = int(math.floor(math.log(size_bytes, 1024)))
p = math.pow(1024, i)
s = round(size_bytes / p, 2)
return f"{s} {size_names[i]}"
def estimate_tokens(text: str) -> int:
"""
估算文本的 token 数量
Args:
text: 输入文本
Returns:
估算的 token 数量
"""
# 简单估算:英文约 4 字符/token,中文约 1.5 字符/token
import re
# 分离中英文
chinese_chars = len(re.findall(r'[\u4e00-\u9fff]', text))
other_chars = len(text) - chinese_chars
# 估算 tokens
estimated_tokens = chinese_chars / 1.5 + other_chars / 4
return int(estimated_tokens)
def truncate_text(text: str, max_length: int = 100) -> str:
"""
截断文本用于显示
Args:
text: 原始文本
max_length: 最大长度
Returns:
截断后的文本
"""
if len(text) <= max_length:
return text
return text[:max_length-3] + "..."
def add_spacing_between_cn_and_en_num(text: str) -> str:
"""
在中文和英文/数字之间添加空格(盘古之白)
Args:
text: 原始文本
Returns:
处理后的文本
"""
import re
if not text:
return text
# 中文-英文/数字
text = re.sub(r'([\u4e00-\u9fff])([a-zA-Z0-9])', r'\1 \2', text)
# 英文/数字-中文
text = re.sub(r'([a-zA-Z0-9])([\u4e00-\u9fff])', r'\1 \2', text)
return text