Merge other-review and self-review

This commit is contained in:
iacore
2026-07-06 00:08:06 +08:00
parent 55c8835885
commit b62a488e11
21 changed files with 661 additions and 987 deletions
-496
View File
@@ -1,496 +0,0 @@
---
name: other-review
description: Editorial polish for CN→EN Buddhist translations you did NOT translate — voice, flow, readability. Load self-review first for terminology/accuracy.
inputs: bilingual.dj (or source.dj + target.dj), or Google Doc
outputs: review-comments.dj
---
{% Serves AGENTS.md Workflow B2 (Other-Review) %}
# Other-Review — Editorial Polish(审他稿)
Review someone else's CN→EN translation. Focus on voice, flow, readability.
You did NOT translate this — suggest, don't command. Load `self-review` first
for terminology/accuracy checks.
Do NOT edit `bilingual.dj` — output to `review-comments.dj` instead.
Use `patch` (mode='replace') for any file edits — not regex-replace.
## Core Principles
Three principles drive all the specific rules below:
1. **Reader engagement** — the English should invite the reader in. Use "we," active
voice, conversational tone. The original is an interview/dialogue; the translation
should sound like one.
2. **Simplicity over sophistication** — prefer everyday words over academic ones.
If a construction feels heavy, stiff, or formal, simplify it.
3. **Source faithfulness** — never add content not in the Chinese, and never drop
content that is. But restructure freely for natural English flow — don't mirror
Chinese word order when it produces awkward English.
These patterns were distilled from 473 editor comments across three manuscripts:
| Document | Comments | Period |
|----------|----------|--------|
| 61 人生佛教在当代的弘扬 | 121 | 2026-06-1920 |
| 附录 我的判教观 | 151 | 2026-06-1316 |
| 16 觉醒的艺术 | 201 | 2026-04-1705-24 |
## Rule Checklist
Apply these in order. Each rule includes the signal (what to look for) and the fix
pattern. Examples are drawn from actual editor comments.
---
### R1: Active Voice + "We" Subject
**Signal**: Passive voice, impersonal "one," "it," abstract subjects.
**Fix**: Use "we" or a concrete human subject.
```
- "the propagation of Buddhism must move toward modernisation"
+ "we must bring Buddhism into the modern age"
- "it becomes even harder to have an environment for further study"
+ "students find it even harder to have a place for further study"
- "One must further develop the mind of renunciation"
+ "We must further develop the mind of renunciation"
```
**Check**: Scan for "one" (as pronoun), passive "be + past-participle," and sentences
where the subject is an abstract noun phrase. Replace with "we" + active verb.
---
### R2: Noun → Verb Conversion
**Signal**: Heavy noun phrases where a verb would be lighter and clearer.
Especially common with: propagation, promotion, engagement, commercialisation,
modernisation, adjustment, continuation, cultivation, development.
**Fix**: Replace the noun with its verb form. Restructure the sentence around a
human subject (usually "we").
```
- "the adjustment of focus in" → delete entirely or rephrase
- "for the propagation of the Dharma" → "to spread the Dharma"
- "the promotion of Humanistic Buddhism" → "as we actively promote..."
- "engagement with the worldly realm" → "it engages with worldly affairs"
- "commercialisation" → "society has become increasingly commercialized"
- "the continuation of" → "flow"
```
**Check**: Search for `tion of`, `ment of`, `isation of`. Try deleting the phrase
first — often the sentence flows fine without it.
---
### R3: Simplify Vocabulary
**Signal**: Formal, academic, or literary words that an everyday reader would trip on.
**Fix**: Replace with the simplest word that preserves the meaning.
```
Academic → Everyday:
"provisions" (资粮) → "spiritual resources" or rephrase
"constitutes" → "is" / "forms" / "makes up"
"cognizes" → "understands"
"transient" → "changing" / "impermanent"
"ancestral masters" → "great masters"
"eminent monks and great masters" → pick one (redundant)
"mode of existence" → "way of living"
"the degree of this entrapment" → "how tightly we cling"
"corresponding" → "proper" / "solid" or delete
"uncritically" → delete (accepting covers it)
"sublimated" → simpler word
"encompassing" → simpler word
"inherently unsustainable" → "they will not last"
```
**Check**: For each sentence, ask: would I say this word in conversation? If not,
find the everyday equivalent.
---
### R4: Break Long Sentences
**Signal**: Single sentence with 3+ clauses, especially with lists, parallel
structures, or multiple "and"/"while"/"with" connectors.
**Fix**: Split into 23 shorter sentences. Each sentence should convey one idea.
```
Before (one sentence):
"Influenced by society, many within the Buddhist community have also become
enthusiastic about power and economic gain, treating tourism and ritual services
as the main focus programme of temples, while neglecting their proper
responsibility of carrying on the Buddha's mission."
After (two sentences):
"Influenced by society, some in the Buddhist community have also become
enthusiastic about power and money. They make tourism and ritual services the
main focus of temple activities, while neglecting their true duty to carry on
the Buddha's mission."
```
**Check**: Count clauses. If >2, consider splitting. Especially split:
- When a list starts mid-sentence
- When "while" / "with" introduces a new subject
- When a long parenthetical could stand alone
---
### R5: Remove Unnecessary Words
**Signal**: Modifiers that don't add meaning, filler words, redundant pairs.
**Fix**: Delete them. The sentence should stand without them.
```
Common deletions:
"so-called" — almost always deletable
"generally" — deletable unless making a specific contrast
"phenomenon" — "this phenomenon" → "this"
"fundamentally" — deletable
"necessarily" — deletable
"solely" — deletable
"briefly summarise" — "briefly" is redundant with "summarise"
"and influenced" — deletable
"the practice of" — often deletable
"The development of"— often deletable
"and investigating" — if already "studying," drop
"practical" (in "practical methods") — "methods" alone covers it
```
**Check**: For each word: does removing it change the meaning? If not, remove it.
---
### R6: Conversational / Interview Tone
**Signal**: Formal register, stiff constructions, academic phrasing in a dialogue
context. The Chinese originals are interviews; the English should read as spoken word.
**Fix**: Imagine you're saying it aloud. Adjust accordingly.
```
Formal → Conversational:
"What do you think is the → "Why do you think that is?"
reason for this?"
"outdated" → "old-fashioned" (in spoken context)
"It is summarised in the → Use direct speech
famous expression: ..."
"the common point of the → "what all these issues have
above-mentioned issues" in common"
"From the perspective of → "Looking at Buddhist history,"
Buddhist history,"
"besides remaining grounded → "In the modernisation of Buddhism,
in the principle of adapting while we must stay true to..."
both to Buddhist teachings
and to the needs of the people"
```
**Check**: Read the passage aloud (mentally). Does it flow like speech? If it feels
like a lecture transcript, lighten it.
---
### R7: No -ly Adverbs
**Signal**: Adverbs ending in -ly.
**Fix**: Delete them, or restructure to avoid them. Modern English style avoids
adverb clutter.
```
- "briefly summarise" → "summarise"
- "fundamentally new" → "new"
- "truly able to" → "able to"
- "extremely limited" → "very small number" (restructure)
- "gradually established" → "established" (ok in context)
```
**Check**: Search for `\w+ly\b`. Evaluate each — most can be dropped.
---
### R8: Concrete over Abstract
**Signal**: Abstract noun phrases that could be expressed more directly.
**Fix**: Make the idea concrete. Use examples, metaphors, or simpler terms.
```
Abstract → Concrete:
"the mistaken identification → "the mistaken belief in a fixed
and attachment to the self" self and our clinging to it"
"corresponding mental functions" → "the mindset that aligns with them"
"discern the universal sequence → "understand how the teachings
from shallow to deep" progress from basic to advanced"
"mode of existence centered → "a life centered on the Three Jewels"
on the Three Jewels"
"a field of specialisation" → delete (unnecessary abstraction)
"ideological content" → "ideas"
```
**Check**: If a noun phrase has 3+ words and feels like jargon, find a plainer way.
---
### R9: Terminology Alignment
**Signal**: Terms that have established translations in the MPI terms database or
prior publications (坐看云起, etc.).
**Fix**: Use the established translation. Load `terms-search` skill and check:
`python3 $MPI_PROJECT_ROOT/terms-search/search.py <term>`
```
Key distinctions from the 3 manuscripts:
人生佛教 → "Buddhism of Human Life" (NOT "Humanistic Buddhism")
人间佛教 → "Humanistic Buddhism"
禅宗 → "Chan" (not "Zen")
根机 → "faculties" (not "dispositions", not "root")
高僧大德 → pick one: "eminent monks" or "great masters" (not both)
假 (三性) → "false" (from emptiness perspective)
monks → "monastics" (includes both monks and nuns — 含比丘比丘尼)
惑业苦 → "delusion, karma, and suffering" (logical order preserved)
出世乐 → keep consistent with "supreme bliss of Nirvana" used earlier
戒体 → don't use "precept essence" (hard to understand); expand to
"the mind of upholding the precepts"
```
**Check**: For any Buddhist technical term, verify against the terms DB before
using a new translation. Consistency across manuscripts matters.
---
### R10: Missing Content Detection
**Signal**: Chinese paragraph has multiple clauses/ideas, but English stops after
12 sentences. Quoted speech, poems, rhetorical climaxes in CN are absent from EN.
**Fix**: Flag as "Missing Content." Translate the missing portion.
```
Common patterns:
- 中文有4-5个短句,英文只有1-2句 → 漏翻了
- CN has a paragraph with quoted speech → EN paragraph ends before the quote
- 一些具有相当成就的前辈高僧... → entire passage skipped
- 内修与外弘也是同样... → entire paragraph missing
- 民众在接受教育期间,无法从教科书上对佛教获得正面了解 → line missing
```
**Check**: Compare CN and EN paragraph lengths. If CN is substantially longer but
EN seems complete, re-read the CN carefully — something was likely dropped.
---
### R11: Source Faithfulness
**Signal**: English adds concepts, metaphors, or explanations not present in the
Chinese. Or English drops specific details that the Chinese includes.
**Fix**: Remove additions. Translate dropped content. But restructure freely for
natural English — don't mirror Chinese word order if it produces awkward results.
```
Additions to remove:
- "this addresses one of the most fundamental philosophical questions"
(not in CN → delete)
- "preserving the potential for future expression" (not in CN → delete)
- "making this school widely known in China" (not in CN → delete)
Things to preserve:
- Quoted speech, idioms, and rhetorical questions — translate fully
- Logical connectors between sentences — don't skip them
- The speaker's distinctive voice and argument structure
- 讲法风格 — the master's teaching style should come through
```
**Check**: For each paragraph, ask: is there content in EN that has no basis in CN?
Is there content in CN that has no corresponding EN? Fix both.
---
### R12: Sentence Structure Clarity
**Signal**: English reads as "translationese" — word order follows Chinese,
parallel structures are misaligned, logical flow is broken.
**Fix**: Restructure for natural English. Reorder clauses. Align parallel items.
```
- "Despite different methods, your focus has always been Humanistic Buddhism"
→ "Although the methods vary, they all center on Humanistic Buddhism"
(methods vs. focus — misaligned comparison)
- "the root cannot be established"
→ "we cannot establish the root" (add human subject)
- Lists: keep verb forms parallel
"First, focus on building... Second, keep liberation at the center...
Third, highlight the unique spirit..."
(imperative verbs aligned)
```
**Check**: Read each sentence and ask: would a native English speaker structure it
this way? If not, reorder.
---
### R13: Specific Word Fixes
**Signal**: Certain English words have wrong connotations for the context.
**Fix**: Use the recommended replacement.
```
"comeback" → "return" (comeback = return to popularity, not return to roots)
"evokes" → "brings" / "gives rise to" (evokes is too formal)
"entertain" → "have" (entertain thoughts → have thoughts)
"falls" → "fades" (faith falls → faith fades)
"reverse" → "transform" (gentler)
"escape" → softer word in context
"dismantle" → "challenges" / "undoes" (dismantle is too aggressive)
"head" → "beginning" (head of the Eightfold Path → beginning)
"none" → "not many" (when the claim is "未提供多少" not "none at all")
"function" → "purpose" (in non-technical contexts)
"object" → "focus" (object of practice → focus of practice)
"uncommon" → "unique" / "distinctive" (不共 = not shared with other paths)
"conduct and propagate" → "uphold and spread"
"attainment of Buddhahood" → "enlightenment" (in some contexts, to avoid
Buddha/Buddhahood repetition)
"liberation" → "liberation" is fine, but check if context means "解脱" or "出世"
```
---
### R14: Positive Feedback — Praise the Translator (随喜)
**Signal**: A passage reads well — accurate, natural, fluent. Or the translation
as a whole shows consistent quality.
**Action**: Praise the first-pass translator by name. The editor consistently opens
or closes review sections with affirmations directed at the translator:
```
"随喜师兄整体翻译还是很好的,前后连贯,清晰"
"随喜师兄翻译 准确流畅!"
"随喜师兄上段和本段翻译 准确流畅"
"随喜师兄查证出处"
"整段翻译的很好"
"整段翻译还是很棒的"
"翻译一气呵成 很精彩"
"很赞叹 简洁有力!"
"这个功不唐捐翻译的真好"
"随喜这个翻译!衔接的很好"
"这个翻译很赞!"
"这段翻译也很好"
"整段翻译的还是很好的"
"翻译的意思都表达出来了"
"本段翻译的都很好,随喜赞叹!"
```
This matters for three reasons:
1. It tells the translator what to preserve — don't accidentally "fix" what works
2. It follows the deliberation protocol — rejoicing (随喜) comes first, before
any criticism. The translator should feel their effort is seen and valued
3. It names the translator explicitly — "随喜师兄" — making the praise personal,
not generic. When reviewing, use the translator's name from the document title
(e.g. from the document title)
## Workflow
Two output modes depending on target format:
### Mode A: .dj file review (use `patch`)
```
1. Read source.dj + target.dj (full files)
2. Apply R1R13 in order, scanning the English line by line
3. For each issue found, apply with patch or write to translation-findings.dj
4. Note well-translated passages (R14)
5. Final pass: read the full English aloud — does it flow?
```
### Mode B: Google Docs review (output to review-comments.dj)
Use when the translation manuscript is a Google Doc and you must add comments
without editing the original text. The agent reads the document, produces a
comment file, and the user manually inserts comments into the Google Doc.
**Why manual insertion is required**: The Google Drive API `comments.create`
with an `anchor` field is silently ignored by Google Docs editor apps. The
kix anchor format used by the Google Docs UI is an undocumented internal
format that has never been reverse-engineered. Line-based anchors
(`{"region": {"kind": "drive#commentRegion", "line": N, "rev": "head"}}`)
also do not display as anchored in the Google Docs UI. This is a known
limitation since 2016 with no resolution. The only reliable way to add
anchored comments is through the browser UI (select text → Insert → Comment).
**Workflow**:
```
1. Load google-workspace skill (ensure OAuth is set up)
2. Read the Google Doc body via Docs API: documents.get
3. Read existing comments via Drive API: comments.list (paginate fully)
4. Identify the translator from the document title (e.g., "maple 初翻")
5. Review per R1R14, scanning the English paragraphs
6. For each issue, extract the EXACT quoted text from the document:
- Use the full document body to find the precise, case-exact string
- Each quoted snippet must be UNIQUE — long enough to find with Ctrl+F
(at minimum 10+ words or a complete short sentence)
- Test: can you find exactly one match when searching the document?
7. Write review-comments.dj to the article directory:
$MPI_PROJECT_ROOT/translate-files/<article>/review-comments.dj
8. Format each entry:
Quoted text: `<exact text from document>`
<comment content>
9. Tell the user: "Comments written to <path>. For each, select the quoted
text in the Google Doc → Insert → Comment → paste the content."
```
**Quoted text rules** (critical — the user must be able to find the text):
- Must match the document text EXACTLY: same case, same punctuation,
same whitespace (including non-breaking spaces like `\xa0`)
- Must be UNIQUE within the document — verify by searching the full body
- Must be long enough to disambiguate: prefer a full short sentence or a
10+ word phrase over single words like "Consequently"
- Capitalize as it appears in the source. If the source has
"Consequently, while countless...", quote that, not "consequently"
**Example output format**:
```
## 10. Conversational Tone (R6)
Quoted text: `even discarded. Consequently, while countless companies emerged, they collapsed just as quickly.`
整体语气适合对话场景。个别词可以更口语化:'Consequently' → 'So'...
```
## Pitfalls
- Don't apply this skill before `self-review` — terminology must be correct first.
- Don't over-correct: some formal constructions are appropriate for certain passages.
Judge per context.
- When converting nouns to verbs, keep the meaning intact — don't simplify away
doctrinal nuance.
- The "we" subject is generally preferred but not absolute. If a passage describes
a general principle without a specific agent, passive/impersonal may be correct.
- Positive feedback (R14) is not optional fluff — it guides what to preserve.
- **Google Docs anchor limitation**: The Drive API cannot create anchored comments
on Google Docs. Use Mode B (output to review-comments.dj) for Google Doc targets.
The user manually inserts comments in the browser UI. See the journey log at
`references/google-docs-comment-journey.dj` for full details.
- **Quoted text must be exact and unique**: when writing review-comments.dj, each
quoted snippet must match the document text case-exactly and be long enough to
find unambiguously with Ctrl+F. Short words like "consequently" are unacceptable
— quote the full sentence or a 10+ word phrase.
- **Mine review comments for reusable patterns**: when a review-comments.dj is produced, scan it for recurring calques and Buddhist term register issues and add them to `../self-review/references/translation-pitfalls.md` and `../self-review/references/buddhist-terminology.md`. The *人生百问* review produced a reusable pattern library — follow that example.
-252
View File
@@ -1,252 +0,0 @@
---
name: self-review
description: Review your own CN→EN translations — three-pass review (terminology → mechanical → flow). Edit commented.dj with patch. For reviewing someone else's work, load other-review.
inputs: bilingual.dj, or CSV/XLSX (Chinese + English columns)
outputs: commented.dj, translation-findings.dj (optional)
---
{% Serves AGENTS.md Workflow B1 (Self-Review) %}
# Self-Review(自审)
Review YOUR OWN translations. You created the English — you can edit freely.
For reviewing someone else's work, load `other-review` instead.
Two input formats: CSV/XLSX batch, or `.dj` comparison file.
`bilingual.dj` is script-generated ground truth — **never edit it.**
Copy to `commented.dj` for all review work. Always use `patch` (mode='replace')
for edits — not regex-based string replacement in `execute_code`.
## Workflow A: CSV/XLSX batch review
Use when input is a CSV/XLSX with `Chinese`/`English` columns. Produces an `edit-suggestions.dj` file.
### 1. Get the data into CSV
If XLSX, export to CSV (or use `openpyxl`). CSV is faster.
### 2. Read the full file
Use `read_file` with offsets for complete coverage. Don't sample.
### 3. Write a systematic analysis script
Write to `/tmp/script.py`, run with `python3 /tmp/script.py`. No heredocs or `-c`.
The script should:
- Parse CSV with `csv.DictReader`
- Apply detection rules per category
- Collect issues: row number, CN text, EN text, problem, suggested fix
- Group/deduplicate identical issues
Common detection categories:
- **Buddhist terminology**: 正念→mindfulness (not "righteous thoughts"), 布施→generosity (not "alms")
- **Identity terms**: 学士/修士/胜士/智士 are practice stages, not titles
- **Literal machine translations**: "hard drive" for 硬盘 (endurance)
- **四摄法 terms**: 同事→"acting in harmony", 爱语→"kind speech"
- **Grammar**: subject-verb agreement, unbalanced quotes
- **Typos/formatting**: Chinese punctuation in English, "IOS"→"iOS"
- **Inconsistency**: same CN term translated differently across rows
### 4. Write edit-suggestions.dj
Format:
```
# 1
original: <Chinese text or key term>
translated: <current English>
<Explanation and suggested fix.>
# 2
...
```
One entry per problem category, not per row. Mention affected row numbers.
## Workflow B: .dj file review (self-review)
Use when you translated the text and want to review your own work.
### File rules
1. `bilingual.dj` — extracted from script. **Never edit.**
2. `cp bilingual.dj commented.dj` — all edits go here
3. `{% ... %}` comments document non-obvious translation choices
### 1. Read the full file
Use `read_file` with `offset` and `limit` for full coverage of large files
(this session had 322 lines, 93KB). For files >200 lines, paginate
explicitly rather than reading the whole thing at once.
If you have already read part of the file with `read_file` earlier in the
session, use `terminal: cat` (or `sed -n 'A,Bp'`) to get an un-deduped view
of the rest. `read_file` deduplicates within a session.
### 2. Scan for problems — three passes, in order
The review is best done in three distinct passes, each catching a different
category of error. Don't try to catch everything in one scan.
**Pass 1 — terminology + consistency + line count** (fast, mechanical):
- **Line count**: source and target must match exactly. Mismatch means paragraphs were dropped, merged, or split.
- Same CN term translated differently across the file (e.g. 人生佛教/人间佛教 conflation, 恨→resentment, 修行/修学)
- Buddhist terminology against the MPI terms DB (see terms-db-alignment below)
- Mistranslation of key terms, wrong proper names, garbled text
- Mid-paragraph truncation: CN covers 35 clauses but EN stops after 12 sentences. Signal: CN has quoted speech, poems, or a rhetorical climax absent from EN. Flag as "Missing Content" not "Incomplete."
- Use `references/common-issues-taxonomy.md` as a structured checklist for accuracy issues
**Pass 2 — mechanical/formatting** (also mechanical, but easy to skip):
- TOC format: AGENTS.md says TOC must be plain bullet list, no link targets. Strip `[I. Heading](#...)` markdown links if present.
- Double words, double punctuation, capitalisation typos, processing artifacts, stray spacing in Chinese text
- Numbering mismatches between CN and EN headings
- Redundant English calques: when the target mirrors a Chinese grammar pattern literally, it can read as a typo (e.g. "mind of death-mindfulness" for 念死之心 — should be "mindfulness of death"). Clunky idioms: 一念之差 → "a single thought of difference" is unidiomatic. Standard renderings exist (e.g. "a single errant thought," "a moment's carelessness," or rephrase as "a single thought can make all the difference").
**Pass 3 — flow/tonal/calques** (read the whole English as prose):
- Re-read the full English target. Does it hang together as prose, or does it read as "translationese"?
- Dramatic verbs that are calques of Chinese: "draw forth," "into full play," "shoulder," "look to with hope." See `references/translation-pitfalls.md` for the full calque checklist.
- New: review the *人生百问* pattern tables in `translation-pitfalls.md` for stiff calques ("keen on," "wisdom culture," "more ultimate," "choice difficulty," "serve as reference") and idiom renderings ("tree wishing stillness," "straddling two boats").
- Subject-shift calques: English substitutes a concrete agent (practitioners, people) for an abstract system noun (Buddhism, religion) — see pitfalls.
- Factual inconsistencies across paired descriptions of the same person/place/thing.
- Tonal coherence inside parallel lists: verb choice should be identical across First/Second/Third items.
- Intensifier drift: same intensifier ("profoundly important") used 3+ times in one section reads as over-translation.
- The user may explicitly request this pass ("are the words together nicely?", "does it read well?"). Treat such prompts as a signal to do the full re-read, not just spot-check.
**Terms database drift** (cross-cutting — apply during Pass 1):
- Cross-reference glossary terms against the MPI terms database
- CLI preferred: `python3 $MPI_PROJECT_ROOT/terms-search/search.py <query>`. For a review, batch many queries in one `execute_code` script (subprocess loop) — one terminal call per term is slow and noisy.
- Source priority: DoT定稿 > 内部特色词 > 佛教术语 > 经论名
- Fix both glossary comments AND body text
- See `references/terms-db-alignment.md` for batch-lookup patterns
### 3. Dump findings
Write to `translation-findings.dj` for issues that don't fit as inline fixes:
```
Finding N — Title (line numbers)
Chinese: ...
English: ...
Issue: description
```
### 4. Apply fixes with `patch`
Surgical string replacement in `commented.dj` with `patch` (mode='replace').
Never use regex-based string replacement in `execute_code` for .dj edits.
`patch` is safer, surfaces conflicts, and produces a reviewable diff.
Verify with `cat` — never rely on `read_file` (session dedup).
### 5. Add inline edit suggestions
Copy `bilingual.dj``commented.dj`, then apply inline corrections + `{% %}` comments.
### 6. Final sweep
Run `python3 scripts/sweep.py <source.dj> <target.dj> [--stale term1,term2] [--new term1,term2]`. This runs all mechanical checks in one call: line parity, heading parity, Unicode em/en-dashes, Markdown bold, Chinese punctuation, TOC link artifacts, unbalanced quotes, and stale/new term assertions. Run even when no content patches were needed — it serves as final validation.
## Buddhist terminology reference
See `references/buddhist-terminology.md` for Chinese-English term mappings and common pitfalls.
## Workflow C: Typeset proofread (DOCX manuscript vs PDF layout)
Use when the user gives a manuscript DOCX and a typeset PDF and asks to proofread.
Goal: catch typesetting errors (missing text, typos, wrong special characters, bad line
breaks), not translation quality.
### 0. Clarify scope FIRST
Before any extraction: ask what they want checked. "Proofread" can mean:
- Text accuracy (missing/doubled words, typos introduced by typesetter)
- Special characters (quotes, dashes, ellipses)
- Formatting (page numbers, headers, TOC layout)
- All of the above
Do not run extraction pipelines until scope is clear.
### 1. Extract text
- DOCX → plain: `pandoc file.docx -f docx -t plain --wrap=none`
- PDF → plain: `pdftotext -layout file.pdf` (preserves positional info)
### 2. Clean PDF artifacts
- Strip InDesign slug lines, page headers, page numbers
- Join hyphenated line breaks (line ending `-` + next line starting lowercase)
- Fix drop-cap artifacts (e.g. `L iving``Living`)
### 3. Compare
- Extract English paragraphs from DOCX (skip Chinese lines, match blank-line pattern)
- Check each DOCX paragraph exists as substring in PDF body text
- Flag paragraphs not found; investigate each (may be heading renumbering, not missing)
### Pitfalls specific to this workflow
- **PDF paragraph joining is lossy** — page breaks split paragraphs. Don't expect
perfect paragraph matching; check content coverage, not paragraph identity.
- **Heading numbering differs** — DOCX has `1.`, `(1)`; PDF has `I`, `1)`. Ignore
heading-only differences.
- **InDesign PDFs insert extra spaces** around drop caps and special characters.
Normalize multi-space to single space before comparison.
### Pitfalls
- **Never edit `bilingual.dj`** — it's script-generated ground truth. Copy to `commented.dj` first.
- **Use `patch`, not regex** — for all `.dj` edits. `patch` surfaces conflicts and produces diffs.
- **`commented.dj` comments must not split paragraphs** — always place `{% %}` after the FULL EN paragraph, not mid-sentence. Scan for merged comments after insertions and split them. Collapse triple+ blank lines created by comment insertions.
- **Clarify scope before diving into extraction pipelines** — if the user says
"proofread this" or "校对这篇文章", ask what specifically they want checked
before running pandoc/pdftotext. Getting interrupted mid-pipeline wastes
context.
- **Don't use heredocs or `-c`** — write to `/tmp/script.py` first
- **Deduplicate aggressively** — group by problem type, not per-row
- **Buddhist terminology is technical** — don't guess. When uncertain, flag for review
- **Never delete .dj comparison files** — intentional work artifacts
- **Verify patches with `cat`** — `read_file` dedup makes it unreliable
- **Re-read before fixing** — user may have made interim edits
- **Batch terminology lookups** — when checking many terms against the terms DB, run them in one `execute_code` script that loops over a query list and calls `search.py` via `subprocess.run`. One terminal call per term floods the context with repetitive output.
- **Mine review comments for patterns** — when a `review-comments.dj` is produced, scan it for recurring calques and term choices and add them to `references/translation-pitfalls.md` and `references/buddhist-terminology.md` so the skill improves with each review.
## Human Review Protocol (审议)
When giving feedback to human translators — whether in a review team or as an AI
assistant flagging issues — follow the Oriental Translation Workshop protocol.
See `references/deliberation-protocol.md` for full guidance.
Key points:
- **Rejoice first** (随喜): affirm what works before flagging issues
- **Three-tier issues**: Level 1 (spelling/grammar/format) — fix directly. Level 2
(omission/mistranslation/wordiness) — suggest or fix with tracked changes. Level 3
(citation versions / marginal wording) — discuss with translator
- **Tone**: questions, not commands; collaborative inquiry, not correction
- **Address translator as 菩萨** (Bodhisattva) — respectful peer
## Common Issues Taxonomy
Use `references/common-issues-taxonomy.md` as a structured checklist when reviewing.
Categories:
- **Accuracy**: omission, mistranslation (over-free, over-literal, misunderstanding,
wrong word choice), overtranslation, terminology errors
- **Readability**: redundancy (long sentences, passive voice, nominalization),
poor structure (top-heavy sentences), wrong register, weak transitions
## References
- `references/buddhist-terminology.md` — Chinese-English Buddhist term mappings and pitfalls
- `references/terms-db-alignment.md` — Batch-aligning glossary terms against the MPI terms database
- `references/translation-pitfalls.md` — Recurring CN→EN mistranslation patterns (关爱→compassion, 生生增上, etc.)
- `references/proofreading-patterns.md` — DOCX/PDF extraction techniques, block-based pairing, common manuscript issues
- `references/docx-md-extraction.md` — Extracting `.docx.md` (pandoc markdown) to bilingual, TOC guards, CN/EN boundary regex
- `references/edit-suggestions-in-bilingual.md` — Inline edit suggestions + `{% %}` comments, two-file comparison (commented.dj / bilingual.dj)
- `references/deliberation-protocol.md` — Oriental Translation Workshop review protocol: tiers, rejoicing, tone
- `references/common-issues-taxonomy.md` — Structured taxonomy of accuracy and readability issues with examples
## Scripts
- `scripts/sweep.py` — Mechanical validation sweep for completed reviews
- `scripts/review_csv.py` — Batch CSV/XLSX translation review
-78
View File
@@ -1,78 +0,0 @@
#!/usr/bin/env python3
"""Reference: reusable detection rules for Chinese→English translation review.
Adapt the rules list for each project's domain vocabulary."""
import csv, re, sys
def review_csv(csv_path):
rows = []
with open(csv_path) as f:
reader = csv.DictReader(f)
for r in reader:
rows.append((r.get('page', '') or '', r['Chinese'] or '', r['English'] or ''))
issues = []
def add(row_idx, cn, en, problem, suggestion):
issues.append({
'row': row_idx + 2,
'page': rows[row_idx][0],
'cn': cn, 'en': en,
'problem': problem,
'suggestion': suggestion
})
for i, (page, cn, en) in enumerate(rows):
if not cn or not en:
continue
# ── Add project-specific detection rules below ──
# Example: "Is we" → machine translation artifact
if re.search(r'\bIs we\b', en):
add(i, cn, en,
"'Is we' — literal MT of 是否/如果. Should be 'if we' or 'whether we'",
re.sub(r'\bIs we\b', 'if we', en))
# Example: Chinese punctuation in English text
if re.search(r'[,。;:!?、]', en):
add(i, cn, en,
"Chinese punctuation in English text",
"[Replace with English punctuation]")
# Example: unbalanced HTML tags
if en.count('<b>') != en.count('</b>'):
add(i, cn, en,
f"Unbalanced <b> tags (open={en.count('<b>')}, close={en.count('</b>')})",
"[Balance tags]")
# Example: unbalanced double quotes
if en.count('"') % 2 != 0:
add(i, cn, en,
f"Unbalanced quotes ({en.count(chr(34))} total)",
"[Balance quotation marks]")
# Example: term inconsistency check
# if re.search(r'TermA', en) and re.search(r'TermB', en) and ...
# ── Sanity checks ──
for i, (page, cn, en) in enumerate(rows):
if cn and not en:
print(f"WARNING row {i+2}: CN present but EN empty: {cn[:80]}")
cn_chars = re.findall(r'[\u4e00-\u9fff]', en)
if cn_chars:
print(f"WARNING row {i+2}: Chinese chars in EN: {cn_chars}")
if cn and en and cn.strip() == en.strip():
print(f"WARNING row {i+2}: CN==EN (untranslated): {cn[:60]}")
return issues
if __name__ == '__main__':
issues = review_csv(sys.argv[1])
print(f"Issues found: {len(issues)}")
for iss in issues:
print(f"\nCSV_ROW_{iss['row']} [{iss['page']}]")
print(f" CN: {iss['cn'][:120]}")
print(f" EN: {iss['en'][:120]}")
print(f" PROBLEM: {iss['problem']}")
print(f" SUGGEST: {iss['suggestion'][:150]}")
-139
View File
@@ -1,139 +0,0 @@
#!/usr/bin/env python3
"""Mechanical sweep for .dj translation review — run after patches or as final verification.
Usage: python3 sweep.py <source.dj> <target.dj> [--stale term1,term2] [--new term1,term2]
Checks:
1. Non-empty line count parity (source == target)
2. Heading count parity
3. Zero Markdown bold (**) in target (djot uses single *)
5. Zero common Chinese punctuation in target
6. Zero [text](#anchor) link artifacts in target TOC area (first 15 lines)
7. Zero unbalanced double-quotes in target
8. --stale: each listed string must appear ZERO times in target
9. --new: each listed string must appear at least once in target
"""
import re
import sys
CN_PUNCT = re.compile(r'[\u3000-\u303f\uff00-\uffef\u201c\u201d\u2018\u2019]')
def read_nonempty(path):
with open(path) as f:
return [l for l in f.read().rstrip('\n').split('\n') if l.strip()]
def main():
if len(sys.argv) < 3:
print("Usage: sweep.py <source.dj> <target.dj> [--stale a,b,c] [--new x,y,z]")
sys.exit(2)
src_path = sys.argv[1]
tgt_path = sys.argv[2]
stale_terms = []
new_terms = []
i = 3
while i < len(sys.argv):
if sys.argv[i] == '--stale' and i + 1 < len(sys.argv):
stale_terms = [t.strip() for t in sys.argv[i+1].split(',') if t.strip()]
i += 2
elif sys.argv[i] == '--new' and i + 1 < len(sys.argv):
new_terms = [t.strip() for t in sys.argv[i+1].split(',') if t.strip()]
i += 2
else:
i += 1
src_lines = read_nonempty(src_path)
tgt_lines = read_nonempty(tgt_path)
tgt_raw = open(tgt_path).read()
errors = 0
# 1. Line count
if len(src_lines) != len(tgt_lines):
print(f"[FAIL] Line count: src={len(src_lines)} tgt={len(tgt_lines)}")
errors += 1
else:
print(f"[OK] Line count: {len(src_lines)}")
# 2. Heading count
src_h = sum(1 for l in src_lines if l.startswith('## '))
tgt_h = sum(1 for l in tgt_lines if l.startswith('## '))
if src_h != tgt_h:
print(f"[FAIL] Headings: src={src_h} tgt={tgt_h}")
errors += 1
else:
print(f"[OK] Headings: {src_h}")
# 3. Unicode em/en-dash
em = tgt_raw.count('\u2014')
en = tgt_raw.count('\u2013')
if em or en:
print(f"[FAIL] Unicode dashes: em-dash={em} en-dash={en}")
errors += 1
else:
print("[OK] No Unicode em/en-dashes")
# 4. Markdown bold
bold = sum(1 for l in tgt_lines if '**' in l)
if bold:
print(f"[FAIL] Markdown bold (**): {bold} lines")
errors += 1
else:
print("[OK] No Markdown bold")
# 5. Chinese punctuation
cn = [(i+1, l[:60]) for i, l in enumerate(tgt_lines) if CN_PUNCT.search(l)]
if cn:
print(f"[FAIL] Chinese/smart punct: {len(cn)} lines")
for ln, snippet in cn[:5]:
print(f" L{ln}: {snippet}")
errors += 1
else:
print("[OK] No Chinese punctuation")
# 6. TOC link artifacts (first 15 lines)
toc_links = sum(1 for l in tgt_lines[:15] if re.search(r'\[.*?\]\(#', l))
if toc_links:
print(f"[FAIL] TOC has [text](#anchor) links: {toc_links}")
errors += 1
else:
print("[OK] TOC clean (no link artifacts)")
# 7. Unbalanced quotes
for i, l in enumerate(tgt_lines):
if l.count('"') % 2 != 0:
print(f"[FAIL] L{i+1}: Unbalanced quotes: {l[:80]}")
errors += 1
if errors == sum(1 for l in tgt_lines if l.count('"') % 2 != 0):
pass # errors already counted above
elif not any(l.count('"') % 2 != 0 for l in tgt_lines):
print("[OK] No unbalanced quotes")
# 8. Stale terms (must be absent)
for term in stale_terms:
count = tgt_raw.count(term)
if count > 0:
print(f"[FAIL] Stale term '{term}' still present: {count}")
errors += 1
else:
print(f"[OK] Stale term '{term}' absent")
# 9. New terms (must be present)
for term in new_terms:
count = tgt_raw.count(term)
if count == 0:
print(f"[FAIL] New term '{term}' not found")
errors += 1
else:
print(f"[OK] New term '{term}' found: {count}")
print(f"\n{'ALL CLEAN' if errors == 0 else f'{errors} ISSUE(S) FOUND'}")
sys.exit(0 if errors == 0 else 1)
if __name__ == '__main__':
main()
+4
View File
@@ -10,6 +10,10 @@ outputs: target.dj (English djot), bilingual.dj, edit-suggestions.dj
Core translation technique for Buddhist/Dharma content. Conventions (djot format,
terms DB query, workflows, output format) are in AGENTS.md.
## Source context
Before translating, understand the source's format and delivery context. Is it a transcript of an oral talk, a book excerpt, a guided meditation script, a Q&A, a written article, or another genre? The register shapes the translation. If the context is not clear from the file path or source content, ask the user before proceeding.
## Mindfulness Bell Corpus
English Buddhist prose — register/style reference for translation.