Files
translation-toolkit/skills/mpi-pptx-translate/SKILL.md
T
iacore 894d769051 refactor(skills): align MPI skills with Agent Skills best practices
The skill metadata had drifted: every SKILL.md name field lacked the
mpi- prefix, contradicting the directory names and the Agent Skills
specification. Descriptions were also missing negative triggers, making
it easy for the agent to load the wrong skill.

Rewrote the pdf-to-docx skill to follow progressive disclosure: the main
SKILL.md dropped from 318 lines to 80, with detailed code examples moved
to on-demand references. Added uv run instructions and /// script PEP 723
metadata so dependencies are declared inline and installed automatically.
Fixed the pptx skill's script paths and added CLI usage messages to both
pptx scripts and the normalize script.

Removed the empty self-review directory that was superseded by the unified
translation-review skill.
2026-07-14 23:18:55 +08:00

4.0 KiB
Raw Blame History

name, description, category, compatibility
name description category compatibility
mpi-pptx-translate Translate PowerPoint files between Chinese and English — extract strings to YAML, translate, quality review, and write back with font-shrink + auto-fit for layout. Use only for .pptx files. Do not use for .ppt, Google Slides exports, or PDFs. productivity Requires Python 3.9+ and uv. Dependencies (python-pptx, pyyaml) are declared in the scripts' /// script metadata.

PPTX Translation

Translate .pptx files between Chinese and English. Covers the full pipeline: extraction → translation → review → write-back.

Workflow

1. Extract strings to YAML

Run with uv:

uv run skills/mpi-pptx-translate/scripts/extract.py original.pptx strings.yaml

uv reads the /// script metadata block and installs python-pptx and pyyaml automatically. Produces YAML with entries:

- slide: 1
  shape: 0
  run: 0
  kind: title
  zh: 开启生命的富足
  en: ""
  • slide — 1-based slide number
  • shape — 0-based shape index within slide
  • run — 0-based paragraph index within text frame (or computed index for tables)
  • kindtitle | subtitle | center_title | body | table | notes
  • zh — source text
  • en — translation target (initially empty)

Table run index formula: num_rows * col + row. Reverse with row = run % num_rows, col = run // num_rows.

Speaker notes use shape: -1.

2. Translate

Do NOT call external translation APIs. Translate directly — the agent IS the model. The user corrects this: "Why do you call external models to do it? You can do it yourself!"

Fill in the en field for every entry. Batch if needed, but translate in your response, not via API calls.

Terminology guidance for Buddhist/gratitude content:

  • 感恩=gratitude, 缘起=dependent origination, 众生=sentient beings
  • 因缘=causes and conditions, 三宝=Three Jewels, 福报=merit/blessings
  • 座上=formal practice, 座下=daily life practice, 共修=group practice
  • 上报四重恩=repaying the four great kindnesses

3. Quality review

Scan for:

  • Terminology consistency (same zh term → same en term throughout)
  • Ellipsis convention — English uses 3 dots ..., zh may use 6
  • Buddhist term accuracy
  • Missing translations
  • Overly literal renderings

4. Write back with layout fixes

uv run skills/mpi-pptx-translate/scripts/build.py strings.yaml original.pptx translated.pptx

The script:

  • Replaces text in matching paragraphs (clears all runs, sets first run)
  • Replaces table cell text (using row/col from computed index)
  • Reduces font size by 18% (FONT_SCALE = 0.82) on all translated shapes and tables
  • Sets auto_size = TEXT_TO_FIT_SHAPE on text frames to handle overflow
  • English text is ~1.31.5× longer than Chinese — font shrink + auto-fit handles most cases

Alternate scripts

The absorbed pptx-translation skill had alternate script names: extract_pptx.py and build_pptx.py. These are functionally equivalent to extract.py and build.py with minor formatting differences (docstrings, variable naming). If the primary scripts fail, the alternates are available in the archive at ~/.hermes/skills/.archive/pptx-translation/scripts/.

Pitfalls

  • "run" in the YAML is actually the paragraph index within a text frame, not the OOXML text-run index. python-pptx iterates paragraphs, not runs.
  • Font shrink only applies to runs that have an explicit font.size — inherited sizes from paragraph/layout defaults are skipped.
  • After write-back, verify with python -m markitdown translated.pptx to check text landed correctly.
  • Tables: font shrink is applied per-cell text frame. Each cell is its own text frame.
  • markitdown may fail with ModuleNotFoundError: dotenv — run pip install python-dotenv first.

Scripts

  • uv run skills/mpi-pptx-translate/scripts/extract.py — extract strings from PPTX to YAML
  • uv run skills/mpi-pptx-translate/scripts/build.py — write translations back with font shrink + auto-fit