Files
deep_research/.opencode/commands/dr-finalize.md

138 lines
4.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
description: Phase 4 - 成稿(v0.6)。dr-editor-in-chief 写 ES/Abstract/Glossary,然后调 Python 脚本链路:translate → build_glossary → apply_glossary → polish → build_report。用法:/dr-finalize [slug]
agent: dr-editor-in-chief
---
你是 dr-editor-in-chief。用户执行了 `/dr-finalize $ARGUMENTS`,进入 Phase 4 成稿链路(v0.6 架构)。
## 架构变更说明(v0.6
**Phase 4 的翻译/润色/出稿已从 LLM agent 改为 Python 脚本**。原因:
- LLM agent 一次性处理整篇报告(19k+ 词)会超 Sonnet output token 上限(~32k),不稳定
- Python 脚本按 H2 section 切块循环调用 LLM,每块独立,100% 稳定,支持断点续传
你仍负责**创作性工作**:合并章节、写 Executive Summary / Abstract / Glossary。其余机械工作全部交给脚本。
## Step 1: 定位项目与健康检查
- `$ARGUMENTS` 非空:用该 slug
- 空:取最近项目
读取 `projects/<slug>/manifest.json`
- `phase2.status == "completed"`
- `phase3.approved == true`(如跳过审校,询问用户确认)
## Step 2: 合并英文稿 + 原创写作(LLM 工作)
加载 skills`mckinsey-method` / `output-hygiene` / `length-budget`
`.opencode/agents/dr-editor-in-chief.md` §Step 3-7 的方式:
1. 合并 `phase2/drafts/ch01.md...chN.md``phase4/final_en.md`
2. 写 Executive Summary800-1000 英文词,融合式 SCQA
3. 写 Abstract500-600 英文词)
4. 写 Glossary(双语对照表,按字母序)
5. 插入占位符:
- `## Table of Contents\n\n[TOC will be generated at final rendering.]`
- `## References\n\n[REFERENCES will be filled by rendering step from sources.jsonl.]`
**禁止**
- 改写 dr-analyst 写好的章节正文
- 给每章强加 SCQA 或小节标题
- 保留调度元数据(字数配额/研究员/quota 等)
## Step 3: 翻译(Python 脚本)
```bash
uv run python scripts/translate.py <slug>
```
完成条件:`phase4/final_zh.md` 生成且字数 ≥ 目标字数的 90%。如未达标,`--force` 强制重跑。
## Step 4: 术语表核查(强烈推荐)
```bash
uv run python scripts/build_glossary.py <slug> --workers 4
```
完成后查看 `phase4/glossary.json`
- `confidence == "high"``issue` 非空的条目:说明发现了错误,需要回塑到正文
- 关注公司名 / 机构名 / 产品名类,它们最容易有拼写错误
## Step 5: 应用术语修正(Python 脚本)
```bash
# 先预览
uv run python scripts/apply_glossary.py <slug> --input phase4/final_zh.md --dry-run
# 确认无误后应用
uv run python scripts/apply_glossary.py <slug> --input phase4/final_zh.md
```
这会把 glossary 中发现的拼写错误 / 错译直接替换进 `final_zh.md`
如润色后仍需二次复核,可手动对 `final_zh_polished.md` 再运行一次:
`uv run python scripts/apply_glossary.py <slug> --input phase4/final_zh_polished.md --dry-run`
## Step 6: 润色(Python 脚本)
```bash
uv run python scripts/polish.py <slug>
```
输出:`phase4/final_zh_polished.md`。查看 `phase4/polish_notes.jsonl` 了解模型标记的异常点。
## Step 7: 出稿(Python 脚本)
```bash
uv run python scripts/build_report.py <slug>
```
自动:
-`manifest.report_title` 命名输出(`<Title>.pdf` + `<Title>.docx`
- PDF 自动插 TOC + 从 `phase2/sources.jsonl` 生成参考文献
## Step 8: 更新 manifest
```json
{
"phase4": {
"status": "completed",
"started_at": "...",
"completed_at": "...",
"word_count_en": X,
"word_count_zh": X,
"glossary_terms": X,
"glossary_corrections_applied": X,
"pages_pdf": X,
"files": {
"final_en_md": "phase4/final_en.md",
"final_zh_md": "phase4/final_zh.md",
"final_zh_polished_md": "phase4/final_zh_polished.md",
"glossary_json": "phase4/glossary.json",
"pdf": "phase4/<Title>.pdf",
"docx": "phase4/<Title>.docx"
}
}
}
```
## Step 9: 汇报
向用户展示:
- 各阶段耗时和成本
- glossary 核查发现的问题数 + 自动修复数
- PDF 页数 / 文件大小
- 如有 low-confidence 术语,提示人工复核
## 失败处理
- translate.py 中断:直接重跑(断点续传)
- build_glossary 大量失败:通常是代理/网络问题,降 workers 到 3 重跑
- polish.py 某块失败:用 `--only N,M` 单独重跑
- build_report 参考文献缺失:查看 warning 列表,补 sources.jsonl
## 关键提示(不变)
- **不要给每章强加 SCQA**v0.4 老问题)
- **元数据清理是合并阶段的事**,不要把章节 frontmatter 或 quota 带进 final_en.md
- **Exa 在 macOS + Clash socks 代理下需要 `trust_env=False`**(已在 SearchClient 处理)