Files

3.9 KiB

Translation Layer Technical Documentation

This document details the architecture of the Translation Layer, which operates primarily on manifest.json.

Core Principle: The Translation Layer is decoupled from the EPUB file format. It reads translatable units from manifest.json, processes them using an LLM, and writes translations back to manifest.json. It does not read or parse the EPUB file directly.


1. Architecture Overview

Data Flow

graph LR
    A[manifest.json] -->|Load| B(ManifestManager)
    B -->|Entries| C{Translator Orchestrator}
    C -->|Sample Text| D[BookProfiler]
    D -->|Style Guide| C
    C -->|Chunks| E[LLMClient]
    E -->|Translation| C
    C -->|Update| B
    B -->|Save| A[manifest.json]
  1. Input: manifest.json (Generated by Preprocessing Layer).
  2. Process:
    • Profiling: Analyze text samples to generate a BookProfile (style, tone, terminology).
    • Translation: Batch entries into chunks, send to LLM, receive translations.
  3. Output: manifest.json (Updated with translated_text fields).

2. Core Modules (src/translation/)

2.1 ManifestManager (manifest_manager.py)

  • Role: The interface for the "Source of Truth".
  • Responsibility:
    • Load manifest.json.
    • Provide list of ManifestEntry objects.
    • Save updates back to disk.
  • Key Method: update_translation(entry_id, translation)

2.2 Translator Engine (translator_engine.py)

  • Role: Orchestrates the translation process.
  • Responsibility:
    • Filtering: Identify untranslated entries.
    • Grouping: Group entries by chapter (file path) to maintain context.
    • Chunking: Create character-based chunks (default ~5000 chars) that do not cross chapter boundaries.
    • Concurrency: Manage async workers (default 5 concurrent tasks).
  • Input: List[ManifestEntry], BookProfile.
  • Output: Updates ManifestEntry objects in-place.

2.3 LLM Client (llm_client.py)

  • Role: Handles raw communication with the LLM Provider (OpenAI compatible).
  • Responsibility:
    • Prompt Engineering: Construct Short-ID based prompts.
    • Rate Limiting: Control RPM (Requests Per Minute).
    • Retry Logic: Exponential backoff for API failures.
    • Logging: Save raw request/response pairs to work/{book}/chunks/ for debugging.

2.4 Book Profiler (../preprocessing/profiler.py)

  • Note: While located in preprocessing, it is often invoked at the start of the translation phase.
  • Role: Generates a style guide.
  • Mechanism: Extracts a sample (Intro + Random segments) from manifest.json entries and asks the LLM to analyze author style.

3. Short ID Strategy

To optimize token usage and ensuring mapping accuracy, we use a Short ID system for LLM interaction.

Prompt Format:

#1: First paragraph text...
#2: Second paragraph text...

Response Format:

#1: 第一段翻译...
#2: 第二段翻译...

The LLMClient maintains a mapping of Short ID (#N) <-> Entry ID (File#UUID) for each chunk lifecycle.


4. Pipeline Usage

The translation is executed via the standalone pipeline script:

python pipeline/02_translate.py --input-epub inputs/my_book.epub
  • --input-epub: Uses the filename to locate the work/ directory.
  • --book-name: Alternatively, specify the book folder name directly.

Dependencies

  • Environment variables must be set in .env:
    • OPENAI_API_KEY
    • OPENAI_BASE_URL (Optional)

5. Development & Debugging

  • Chunk Logs: Check .work/{book}/chunks/ to see exactly what was sent to and received from the LLM.
  • Idempotency: The translation script skips entries that already have translated_text. To re-translate, you must manually clear translated_text in manifest.json or delete the manifest (to restart from preprocessing).