An OmegaT machine-translation plugin that forwards segments to a local AI service for translation, with full glossary enforcement, TM fuzzy-match context, and document-level summarization. Works with local models via Ollama or cloud models via Anthropic/Google APIs.
- AI translation with glossary enforcement, fuzzy-match context, and surrounding-segment context
- Server-side translation memory cache — repeat segments return instantly with no extra LLM call
- Batch pre-translation — on file open, a popup offers to pre-translate all untranslated segments in the background; each segment returns instantly from cache when you reach it
- Optional QA self-critique pass — verifies each translation against the glossary and style rules, auto-correcting violations
- Automatic document summarization, injected into each translation request for better context
- Glossary extraction: LLM identifies candidate terms from each file; authoritative EN↔FR translations looked up from a local terminology index (import Termium/OQLF open-data once via CLI)
- Works with any model via Ollama (local, free) or Anthropic/Google APIs (cloud)
- OmegaT 6+
- Java 11+ and Maven (to build the plugin)
- Python 3.13+ and uv
- Ollama (optional, for local models)
- An Anthropic or Google API key (optional, for cloud models)
Build it:
cd plugin
mvn packageThis produces target/ai-translate-plugin-0.1.0.jar. Copy it to your OmegaT plugins folder:
| OS | Path |
|---|---|
| macOS | ~/Library/Preferences/OmegaT/plugins/ |
| Windows | %APPDATA%\OmegaT\plugins\ |
| Linux | ~/.omegat/plugins/ |
cd service
uv sync
cp .env.example .env # edit .env with your model/API key choices
./start.shThe service listens on http://localhost:8000 by default.
All service settings live in service/.env (see service/.env.example for the full list with descriptions):
| Setting | Default | Description |
|---|---|---|
ANTHROPIC_API_KEY |
(empty) | Required only if using an anthropic:... model |
AI_MODEL |
ollama:mistral-nemo |
Model used for translation |
GLOSSARY_MODEL |
(falls back to AI_MODEL) |
Model used for glossary web research — benefits from a stronger model |
OLLAMA_BASE_URL |
http://localhost:11434 |
Local Ollama instance |
STYLE_RULES_PATH |
(unset) | Path to a global style-rules file injected into the translation prompt — copy service/ai_style_rules.example.txt to get started |
STATE_DB_PATH |
platform user data dir | SQLite DB for glossary/summary/translation-memory state |
TM_CACHE_ENABLED |
true |
Server-side translation memory cache; set false to force a fresh LLM call per segment (gates both cache read and write) |
QA_ENABLED |
false |
Opt-in QA self-critique pass that checks each new translation against glossary + style rules and auto-corrects violations |
QA_MODEL |
(falls back to AI_MODEL) |
Model used for the QA pass — benefits from a stronger model |
GLOSSARY_MAX_TERMS |
20 |
Max candidate terms sent for terminology lookup |
GLOSSARY_MAX_PAGE_CHARS |
3000 |
Max characters of source text passed to the LLM per lookup |
The OmegaT plugin's service URL is configurable via the OmegaT preferences key
ai_translation_service_url (default http://localhost:8000) — set it in
omegat.prefs if you run the service on a different host or port.
Glossary extraction looks up candidate terms in a local SQLite index — fast, offline, no per-request network call. The index is empty until you import data. Import Termium and/or OQLF once using the CLI:
cd service
# OQLF Grand dictionnaire terminologique (single CSV, all domains)
curl -L "https://www.donneesquebec.ca/recherche/dataset/1c6567bf-8995-40b9-84a4-50faabae12f4/resource/c3ce0af4-7c0f-4dd2-b53a-6dc7fb3ea5ef/download/fiches_recentes_signees_oqlf_2026-01-19.csv" \
-o oqlf.csv
uv run python -m glossary.cli import-terminology oqlf.csv --preset oqlf
# Termium (one ZIP per subject — repeat for each subject you need)
curl -L "https://donnees-data.tpsgc-pwgsc.gc.ca/bt1/tp-tp/domaine-subject-construction.zip" \
-o construction.zip && unzip construction.zip
uv run python -m glossary.cli import-terminology "Construction_*.csv" --preset termiumBoth sources are open data (OGL-Canada / Données Québec). The index persists in
state.db — re-run the import commands to refresh when updated exports are released.
To add your own terminology (IATE export, corporate glossary, any CSV), use the
same import-terminology command with a custom --column-map instead of --preset.
Style rules are resolved in two layers, with the per-project file taking priority:
- Global default — set
STYLE_RULES_PATHinservice/.envto a file (copyservice/ai_style_rules.example.txt). Applies to every project. - Per-project override — place a file named exactly
ai_style_rules.txtin the OmegaT project's root folder (next toomegat.project). When present it replaces the global rules for translations in that project.
The per-project file must be named exactly ai_style_rules.txt — any other name
(style_rules.txt, ai_style_rules.md, …) is ignored. The plugin logs which path
it checked and whether a file was loaded, so check OmegaT's log if rules don't
seem to apply. Both files use the same format: one rule per line, # for comments.
/translate caches each translation in SQLite, keyed by an exact-match hash of
source_text + source_lang + target_lang + glossary + the model. OmegaT's
MT pane re-queries on every revisit (with repeat-suppression disabled), so
without this, the same segment would trigger a fresh, billable LLM call each
time — the cache returns the stored translation instantly instead, and only
calls the LLM for genuinely new or changed input. Editing the glossary or
switching model busts the cache automatically.
Style rules are deliberately not part of the key (OMP-032). They're
global, so folding them in would re-key every segment on any edit and throw away
a whole warmed cache. Style rules still shape translation and the QA pass on a
cache miss — they just don't invalidate existing hits. The trade-off: after you
edit ai_style_rules.txt, segments already cached keep being served under the
old rules; the new rules only apply to not-yet-cached segments. When a rule
change is significant enough that you want it applied everywhere, clear the cache
manually (DELETE FROM translation_memory in state.db, optionally scoped by
project_id).
The cache is scoped per OmegaT project (project_id) and excludes surrounding
context (fuzzy matches, file summary) from the key — same source text in
different context returns one cached translation, matching how OmegaT's own TM
behaves. There's no eviction; a changed key just orphans the old row, which is
fine at single-user scale. The /translate response includes from_cache: true
when served from the cache. Set TM_CACHE_ENABLED=false to turn the cache off
entirely (every segment then triggers a fresh LLM call).
With QA_ENABLED=true, each new translation gets a second LLM pass that checks
it against the approved glossary and the resolved style rules, then returns a
minimally-corrected version — so the suggestion you see already honours your
terminology and style. It runs only on a cache miss and before the result is
cached, so the cost is paid once per genuinely new segment and never on a repeat.
Segments with neither glossary terms nor style rules are skipped (nothing to
check). Use QA_MODEL to run the pass on a stronger model than AI_MODEL.
Corrections are auditable: each fix is logged to both the service log and OmegaT's
own log (grep for QA correction: in OmegaT's log file). QA is off by default
because it doubles the per-segment LLM cost on cache misses.
When you open a file in OmegaT, a popup offers to pre-translate all untranslated segments in the background:
Pre-translate N untranslated segments in "filename.docx"?
Results will be cached for instant retrieval when you reach each segment.
Clicking Yes sends the whole file to the service in a single request. Each segment goes through the same orchestrator as the live MT pane — glossary terms and project style rules are applied, and the QA pass runs if QA_ENABLED=true. Results land in the TM cache; when you navigate to a pre-translated segment, the MT pane responds instantly with no LLM call.
Key details:
- Non-destructive — nothing is inserted into OmegaT's translation fields. Results only land in the server-side TM cache.
- Fuzzy matches not sent in the batch — OmegaT's matcher is async and tied to active-segment navigation. The cache key excludes fuzzy matches anyway, so the cached result is served on the live request regardless.
- Once per file per session — the popup appears the first time a file is opened in an OmegaT session. Clicking No skips it until the next session.
- Progress — logged to OmegaT's log (
AI Translation Assistant: batch pre-translate...). A completion dialog shows how many segments were cached and how many were already in cache. - Failure is silent — if the service is unreachable or an individual segment fails, the error is logged and live translation continues normally.
The glossary agent (service/glossary/agent.py) is structured so adding your
own PydanticAI tool is ~10 lines alongside
the existing lookup_terminology tool. This works the same whether
AI_MODEL/GLOSSARY_MODEL is Ollama or a cloud model.
Worked example — a DuckDuckGo web-search tool (no API key required):
import httpx2 as httpx # already a transitive dep; add to imports in agent.py
@_glossary_agent.tool
async def fetch_duckduckgo(ctx: RunContext[GlossaryDeps], term: str) -> str:
"""Web-search a term on DuckDuckGo and return stripped result text."""
url = f"https://html.duckduckgo.com/html/?q={term}"
try:
async with httpx.AsyncClient(timeout=10) as client:
resp = await client.get(url, follow_redirects=True)
return _strip_html(resp.text) if resp.status_code == 200 else f"HTTP {resp.status_code}"
except Exception as e:
return f"Error: {e}"Update the Phase 2 prompt in extract_glossary to mention the new tool name so
the LLM knows it can call it. Same idea applies to a corporate glossary API,
Tavily, or anything with an HTTP endpoint — fetch, return text, reference in prompt.
- In OmegaT: Options → Machine Translate → enable "AI Translation Assistant"
- Translate a segment as usual — the plugin calls the local service for each one
- When you open a file, two popups may appear: one offers to extract glossary candidates from the file's content, another offers to pre-translate all untranslated segments in the background (results are cached for instant retrieval)
OmegaT ──▶ Plugin (Java, JAR) ──▶ Service (Python, FastAPI) ──▶ LLM (Ollama / Anthropic / Google)
│
▼
SQLite (glossary + summary + translation memory)
The plugin never reads files directly from the service's perspective — all content is sent in the request payload, not as filesystem paths.
- Built-in terminology presets (Termium, OQLF) are Canadian EN↔FR-focused — import your own CSV for other languages
Issues and pull requests welcome.
MIT — see LICENSE.