translate: 安宁疗护与临终关怀的时代意义
fix translation: compassion→care, making a living→survival-level, 生生增上 life-after-life, 因病返贫 remove 'back'
This commit is contained in:
@@ -15,13 +15,12 @@ When translating Chinese↔English Buddhist/Dharma content, consult TWO resource
|
||||
|
||||
### Terms DB
|
||||
- DB: `/home/user/documents/mpi/terms-search/termlib.duckdb`
|
||||
- CLI: `terms-search <query>`
|
||||
- HTTP API: port 8910 (`/search?q=...&src=...&limit=...`)
|
||||
- CLI: `/home/user/documents/mpi/terms-search/search.py <query>`
|
||||
- Source tables priority: DoT定稿 > 内部特色词 > 佛教术语 > 经论名
|
||||
- See `terms-search` skill for full API details
|
||||
- See `/home/user/documents/mpi/terms-search/SKILL.md` for full API details
|
||||
|
||||
### Mindfulness Bell Corpus
|
||||
- Location: `~/meta/www.files/public/The Mindfulness Bell/`
|
||||
- Location: `/home/user/meta/www.files/public/The Mindfulness Bell/`
|
||||
- 6 issues: MB92–MB97 (2023–2026), each ~84–116 pages
|
||||
- Index: `index.yaml` (article titles, authors, pages)
|
||||
- **Articles**: `~/documents/jingxin-lessons/Mindfulness Bell/articles/MB{92..97}/*.md` — 93 individual markdown files with YAML frontmatter. Read directly with `read_file`.
|
||||
|
||||
@@ -1,38 +0,0 @@
|
||||
#!/usr/bin/env fish
|
||||
# install.fish — Create symlinks from ~/.hermes/skills/ to ./skills/
|
||||
# Run from the skills/ directory or set SKILLS_DIR.
|
||||
|
||||
set -q SKILLS_DIR; or set SKILLS_DIR (realpath (dirname (status filename)))
|
||||
set HERMES_SKILLS "$HOME/.hermes/skills"
|
||||
|
||||
# skill name → category subdirectory under ~/.hermes/skills/
|
||||
set -l names \
|
||||
pptx-translate \
|
||||
chinese-text-normalize \
|
||||
translation-review \
|
||||
dharma-translation
|
||||
set -l cats \
|
||||
productivity \
|
||||
productivity \
|
||||
productivity \
|
||||
research
|
||||
|
||||
for i in (seq (count $names))
|
||||
set name $names[$i]
|
||||
set cat $cats[$i]
|
||||
set src "$SKILLS_DIR/$name"
|
||||
set dst "$HERMES_SKILLS/$cat/$name"
|
||||
|
||||
if not test -d "$src"
|
||||
echo "SKIP: source not found: $src"
|
||||
continue
|
||||
end
|
||||
|
||||
if test -e "$dst" -o -L "$dst"
|
||||
rm -rf "$dst"
|
||||
end
|
||||
|
||||
mkdir -p (dirname "$dst")
|
||||
ln -s "$src" "$dst"
|
||||
echo "LINKED: $dst -> $src"
|
||||
end
|
||||
+11
-6
@@ -1,10 +1,15 @@
|
||||
Translation-related skills.
|
||||
# MPI Skills
|
||||
|
||||
## For Hermes
|
||||
These skills are loaded via `skills.external_dirs` in `~/.hermes/config.yaml`.
|
||||
|
||||
For every skill here, place a symlink to it under ~/.hermes/skills/ —
|
||||
run `fish install.fish` to set up all symlinks.
|
||||
## Setup
|
||||
|
||||
If you create a skill related to translation, place it here in this directory
|
||||
and add its entry to install.fish.
|
||||
Add this to `~/.hermes/config.yaml`:
|
||||
|
||||
```yaml
|
||||
skills:
|
||||
external_dirs:
|
||||
- /home/user/documents/mpi/skills
|
||||
```
|
||||
|
||||
{% `hermes config set` stores list values as strings. Edit the YAML file directly to ensure proper list syntax. %}
|
||||
|
||||
@@ -0,0 +1,76 @@
|
||||
---
|
||||
name: terms-search
|
||||
description: Full-text search across the MPI term database. Use when translating or looking up Chinese-English Buddhist/MPI terminology.
|
||||
category: research
|
||||
---
|
||||
|
||||
# Terms Search
|
||||
|
||||
Database: `/home/user/documents/mpi/terms-search/termlib.duckdb`
|
||||
CLI: `/home/user/documents/mpi/terms-search/search.py`
|
||||
Server: `/home/user/documents/mpi/terms-search/server.py`
|
||||
|
||||
## CLI (preferred)
|
||||
|
||||
```
|
||||
/home/user/documents/mpi/terms-search/search.py <query> [limit]
|
||||
```
|
||||
|
||||
Multi-word queries are ANDed. Searches both `zh` and `en` columns.
|
||||
|
||||
## HTTP API (use only when CLI is insufficient)
|
||||
|
||||
Start: `python3 /home/user/documents/mpi/terms-search/server.py` (port 8910)
|
||||
|
||||
- `GET /` — plain HTML UI (form + results table, no CSS)
|
||||
- `GET /` — plain HTML UI (form + results table, no CSS)
|
||||
- `GET /search?q=...&loc=...&src=...&limit=...` — JSON `{count, results: [{zh, en, loc, source}]}`
|
||||
- `GET /sources` — JSON array of `{source, count}` for all source tables
|
||||
|
||||
All params optional. Omit `limit` for all results. Query terms are ANDed across zh+en.
|
||||
Errors return `{"error": "..."}` with HTTP 500 (API) or shown inline (UI).
|
||||
|
||||
## Source tables
|
||||
|
||||
| src | rows | description |
|
||||
|---|---|---|
|
||||
| BAICKZ | 7,679 | Main term bank with example sentences |
|
||||
| 佛教术语 | 1,795 | Buddhist terminology from 定稿书目术语库 |
|
||||
| DoT定稿 | 896 | DoT final translation decisions |
|
||||
| DoT初步 | 412 | DoT preliminary queries |
|
||||
| 偈颂经文名言 | 263 | Verses and sutra quotes |
|
||||
| 成语俗语 | 184 | Idioms and common expressions |
|
||||
| 经论名 | 89 | Sutra/shastra titles |
|
||||
| 内部特色词 | 87 | MPI internal terminology |
|
||||
| 海内外建筑名称 | 28+9 | MPI building/place names |
|
||||
| MPI组织架构 | 4+28 | MPI org structure |
|
||||
| 导师金句 | 24 | Teacher quotes |
|
||||
| 静心学堂课程 | 17+20 | Course names |
|
||||
| 禅意项目 | 14+11 | Zen program terms |
|
||||
| 公案 | 8 | Chan koans |
|
||||
|
||||
## Direct DuckDB
|
||||
|
||||
```
|
||||
duckdb /home/user/documents/mpi/terms-search/termlib.duckdb
|
||||
```
|
||||
|
||||
Key tables: `unified_terms_flat` (zh, en, loc, source), individual source tables, `unified_terms` view.
|
||||
|
||||
## Rebuilding
|
||||
|
||||
Terms data comes from `/home/user/documents/mpi/guide/03 术语库/`. To rebuild:
|
||||
1. Convert source xlsx/ods → CSV+YAML in `_output/`
|
||||
2. Rebuild DuckDB from CSVs
|
||||
3. Materialize `unified_terms_flat` view → table for performance
|
||||
|
||||
**Full rebuild pipeline:** See `references/termbase-rebuild.md` (absorbed from the `termbase-management` skill).
|
||||
|
||||
## Translation Alignment
|
||||
|
||||
When aligning translated djot files against the term database, load `references/translation-alignment.md` for the full workflow. Summary:
|
||||
1. Extract Chinese terms from `{% "..." %}` glossary blocks in the translated file
|
||||
2. Batch-search via HTTP API (`/search?q=...`)
|
||||
3. Prioritize DoT定稿 > 内部特色词 > 佛教术语
|
||||
4. Fix both glossary comments AND body-text occurrences
|
||||
5. Verify with grep
|
||||
@@ -0,0 +1,127 @@
|
||||
# Termbase Rebuild (from absorbed termbase-management)
|
||||
|
||||
How to rebuild the terms DuckDB from source spreadsheets. This is the full pipeline from the now-archived `termbase-management` skill.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
```bash
|
||||
pip install openpyxl odfpy pyyaml duckdb
|
||||
```
|
||||
|
||||
## Step 1: Inspect spreadsheet structure
|
||||
|
||||
```python
|
||||
from openpyxl import load_workbook
|
||||
wb = load_workbook(path, read_only=True, data_only=True)
|
||||
for sn in wb.sheetnames:
|
||||
ws = wb[sn]
|
||||
rows = [list(r) for r in ws.iter_rows(min_row=1, max_row=6, values_only=True)]
|
||||
print(f"[{sn}] {sum(1 for _ in ws.iter_rows())} rows, cols: {len(rows[0]) if rows else 0}")
|
||||
for r in rows[:5]: print(f" {r}")
|
||||
```
|
||||
|
||||
For ODS files, use odfpy:
|
||||
```python
|
||||
from odf.opendocument import load as odf_load
|
||||
from odf.table import Table, TableRow, TableCell
|
||||
from odf.text import P
|
||||
|
||||
doc = odf_load(path)
|
||||
for table in doc.getElementsByType(Table):
|
||||
for row in table.getElementsByType(TableRow):
|
||||
cells = row.getElementsByType(TableCell)
|
||||
vals = []
|
||||
for cell in cells:
|
||||
text = ''
|
||||
for p in cell.getElementsByType(P):
|
||||
for node in p.childNodes:
|
||||
if node.nodeType == node.TEXT_NODE:
|
||||
text += node.data
|
||||
vals.append(text.strip() if text else None)
|
||||
```
|
||||
|
||||
## Step 2: Convert to CSV + YAML
|
||||
|
||||
### Cleaning
|
||||
- Strip trailing None/empty values from each row: `while row and not row[-1]: row.pop()`
|
||||
- Skip entirely empty rows
|
||||
- Pad all rows to the max column count
|
||||
|
||||
### Duplicate header handling
|
||||
Some sheets have duplicate column names (e.g., two `英文` columns in paired layout). Deduplicate with suffixes:
|
||||
```python
|
||||
from collections import Counter
|
||||
def dedup_headers(headers):
|
||||
seen = Counter()
|
||||
result = []
|
||||
for h in headers:
|
||||
s = str(h) if h else ''
|
||||
if s in seen:
|
||||
seen[s] += 1
|
||||
result.append(f"{s}_{seen[s]}")
|
||||
else:
|
||||
seen[s] = 1
|
||||
result.append(s)
|
||||
return result
|
||||
```
|
||||
|
||||
### Output formats
|
||||
- **CSV**: `csv.writer` — column-major, preserves all raw data
|
||||
- **YAML**: `yaml.dump(data, allow_unicode=True, default_flow_style=False, sort_keys=False, width=200)` — list of dicts
|
||||
|
||||
## Step 3: Load into DuckDB
|
||||
|
||||
```python
|
||||
import duckdb
|
||||
con = duckdb.connect('termlib.duckdb')
|
||||
|
||||
# Simple CSVs work with auto-detect:
|
||||
con.execute("""
|
||||
CREATE TABLE table_name AS
|
||||
SELECT * FROM read_csv_auto('file.csv', header=true, all_varchar=true)
|
||||
""")
|
||||
```
|
||||
|
||||
### Pitfall: Multiline CSV fields
|
||||
CSV files with embedded newlines (common in glossary example-sentence columns) break DuckDB's auto-sniffer. Fall back to Python csv.reader:
|
||||
|
||||
```python
|
||||
import csv
|
||||
with open(path, 'r', encoding='utf-8') as f:
|
||||
rows = list(csv.reader(f))
|
||||
|
||||
headers = rows[0]
|
||||
data = rows[1:]
|
||||
|
||||
col_defs = ', '.join(f'"{h}" VARCHAR' for h in cleaned_headers)
|
||||
con.execute(f'CREATE TABLE "{table}" ({col_defs})')
|
||||
|
||||
batch_size = 500
|
||||
for i in range(0, len(data), batch_size):
|
||||
batch = data[i:i+batch_size]
|
||||
placeholders = ', '.join(['(' + ', '.join(['?' for _ in headers]) + ')' for _ in batch])
|
||||
flat = [v for row in batch for v in row]
|
||||
con.execute(f'INSERT INTO "{table}" VALUES {placeholders}', flat)
|
||||
```
|
||||
|
||||
### Pitfall: DuckDB CLI opens in-memory by default
|
||||
Running plain `duckdb` gives an empty database. Always pass the file path:
|
||||
```
|
||||
duckdb path/to/termlib.duckdb
|
||||
```
|
||||
|
||||
## Step 4: Create unified views
|
||||
|
||||
See `references/unified-view.sql` for the pattern. Key patterns:
|
||||
- `UNION ALL` across all source tables
|
||||
- Normalize column names to `zh`, `en`, `loc` (出处), `source`
|
||||
- For paired-column sheets (e.g., `中文/英文` + `补充内容/英文_1`), emit two UNION branches
|
||||
- Filter out rows where zh or en is NULL/empty
|
||||
|
||||
## Pitfalls
|
||||
|
||||
- **ODS reading**: Must traverse `odf.text.P` child elements, not direct text nodes
|
||||
- **Duplicate headers**: JSON/YAML dict silently overwrites duplicate keys — always deduplicate
|
||||
- **Multiline CSV + DuckDB**: `read_csv_auto` fails on CSVs with quoted newlines — use Python csv.reader
|
||||
- **DuckDB path**: Always explicit file path; `duckdb` alone is in-memory
|
||||
- **`execute_code` sandbox**: Does NOT share pip-installed packages — use `terminal` for Python scripts
|
||||
@@ -0,0 +1,58 @@
|
||||
# Translation Alignment Workflow
|
||||
|
||||
How to align translated djot files against the MPI terms database.
|
||||
|
||||
## When
|
||||
|
||||
After producing a first-pass translation, or when the user asks to check terminology. Any time a `.dj` file contains glossary-style `{% "..." %}` blocks with Chinese→English term pairs.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Read the full translated file. Extract all Chinese terms from `{% "TERM" (pinyin) = ENGLISH ... %}` blocks.
|
||||
|
||||
2. Start the search server: `python3 /home/user/documents/mpi/terms-search/server.py &` (port 8910). It may already be running — check with `curl -s http://localhost:8910/`.
|
||||
|
||||
3. Batch-search each term via the HTTP API:
|
||||
```
|
||||
curl -s "http://localhost:8910/search?q=TERM&limit=5"
|
||||
```
|
||||
Prefer `src=DoT定稿` filter for authoritative hits, but also check without filter to catch 内部特色词 and 佛教术语 entries.
|
||||
|
||||
4. For each term, compare the database `en` against the file's translation. A mismatch exists when the core term translation differs (ignore explanatory commentary in glossary blocks).
|
||||
|
||||
5. Priority order for which source to trust:
|
||||
- DoT定稿 (highest authority — final translation decisions)
|
||||
- 内部特色词 (MPI internal terminology)
|
||||
- 佛教术语 (general Buddhist terms)
|
||||
- 经论名 (sutra/shastra titles)
|
||||
|
||||
6. Apply fixes with `patch` tool. Fix BOTH the glossary comment AND all body-text occurrences. Use `replace_all=true` for terms that appear identically in multiple places.
|
||||
|
||||
7. After fixing, grep for remaining old forms to verify nothing was missed.
|
||||
|
||||
## Pitfalls
|
||||
|
||||
- Glossary blocks often include commentary after the term (pinyin, explanations). Compare only the core term translation, not the full comment.
|
||||
- Some terms appear in body text without glossary blocks — scan the body for domain terms too.
|
||||
- The search.py CLI does NOT support `src:` or `loc:` filter syntax directly; use the HTTP API or direct DuckDB queries instead.
|
||||
- Replace-all can create double articles ("the The Eight Steps...") when body text already has the article before the term. Check each replacement site.
|
||||
- Escaped quotes in patch old_string/new_string cause false "Escape-drift" errors. Use unescaped `"` characters from the actual file content.
|
||||
- Terms may have different translations in different contexts (e.g., standalone 心灯 = "lamp of awakening" vs compound 点亮心灯 = "illuminate one's heart"). Use the standalone form for glossary entries.
|
||||
- Sutra quote conventions (e.g., Diamond Sutra "lives" not "bodies" for 身体布施) are not always in the database. Apply standard English Buddhist idiom.
|
||||
|
||||
## Example
|
||||
|
||||
Searching 三无漏学:
|
||||
```
|
||||
curl -s "http://localhost:8910/search?q=三无漏学&limit=5"
|
||||
→ DoT定稿: "three uncontaminated forms of training"
|
||||
→ Current file: "the three undefiled studies"
|
||||
→ MISMATCH → fix
|
||||
```
|
||||
|
||||
## Report
|
||||
|
||||
After alignment, the user may ask for a report. Save it as `alignment-report.dj` in the project directory with:
|
||||
- Summary line (source breakdown)
|
||||
- Per-term before/after table with source annotation
|
||||
- "Not Changed" section listing terms checked and found acceptable
|
||||
Reference in New Issue
Block a user