Leaderboard

This leaderboard will stay live until 31.12.2026

Find the metric implementation here

Find the full dataset here

Track №1: Document-Level Translation with Explicit Dictionary

scored on the proper mode translations

systemchrF++ (doc)
[enpl, eseu]
chrF++ (para)
[enpl, eseu]
Term Success
[enpl, eseu]
Term Success (lemma)
[enpl, eseu]
Term Success (exclusive)
[enpl, eseu]
COMET
[enpl, eseu]
MetricX ↓
[enpl, eseu]
LLM Judge
[enpl, eseu]
oracle
100.0
[100.0, 100.0]
100.0
[100.0, 100.0]
96.5%
[94.2%, 98.7%]
99.3%
[99.0%, 99.6%]
96.4%
[93.6%, 99.1%]
[△, △]
[△, △]
[△, △]
Cohere CAT+
71.1
[74.9, 67.3]
65.2
[70.7, 59.7]
80.4%
[83.8%, 77.0%]
93.7%
[93.6%, 93.8%]
88.0%
[84.9%, 91.1%]
[△, △]
[△, △]
[△, △]
VNFusion
73.5
[72.8, 74.1]
67.7
[67.4, 68.1]
83.5%
[82.9%, 84.0%]
92.6%
[92.9%, 92.4%]
87.6%
[84.6%, 90.7%]
[△, △]
[△, △]
[△, △]

official competition systems systems added after competition
△ awaiting metric evaluation on an asynchronous worker
Scores average over the language directions in brackets; a value appears once every direction is scored.

Track №2: Document-Level Translation with Sample Bitexts

scored on the sample mode translations

systemchrF++ (doc)
[zhen, enpl, eseu]
chrF++ (para)
[zhen, enpl, eseu]
Term Success
[zhen, enpl, eseu]
Term Success (lemma)
[zhen, enpl, eseu]
Term Success (exclusive)
[zhen, enpl, eseu]
COMET
[zhen, enpl, eseu]
MetricX ↓
[zhen, enpl, eseu]
LLM Judge
[zhen, enpl, eseu]
oracle
100.0
[100.0, 100.0, 100.0]
100.0
[100.0, 100.0, 100.0]
98.6%
[100.0%, 96.6%, 99.3%]
99.8%
[100.0%, 99.4%, 99.9%]
93.4%
[86.3%, 94.9%, 99.1%]
[△, △, △]
[△, △, △]
[△, △, △]
Cohere CAT+
73.4
[71.9, 78.0, 70.4]
69.0
[71.9, 74.6, 60.4]
75.7%
[89.4%, 76.0%, 61.6%]
83.1%
[90.5%, 85.6%, 73.1%]
75.2%
[82.4%, 73.4%, 69.9%]
[△, △, △]
[△, △, △]
[△, △, △]
star-mt
67.3
[70.1, 68.1, 63.6]
62.7
[70.1, 63.6, 54.4]
52.5%
[86.5%, 40.6%, 30.5%]
57.9%
[87.7%, 48.1%, 37.9%]
53.1%
[80.6%, 42.8%, 35.9%]
[△, △, △]
[△, △, △]
[△, △, △]

official competition systems systems added after competition
△ awaiting metric evaluation on an asynchronous worker
Scores average over the language directions in brackets; a value appears once every direction is scored.