We round-tripped 1,268 files through Lexora

Listen to this post12:09
We round-tripped 1,268 files: Trados, memoQ and Office file icons on a violet background

Update, 7 October 2026: after this run we fixed what it found. With the next update, Lexora reads Excel and PowerPoint files saved as Strict Open XML, translates the words of links stored as Word fields, reads text in XML elements named data outside DITA and in CDATA sections, and says so when a file’s only text is in charts, SmartArt or Word fields. We fixed our test too: a file Lexora can’t read back now counts as a failure. Run again with these fixes, 1,159 of the 1,267 files go through and back, 108 are refused with a message, and none fail.

We took 1,268 files in the formats translators are sent, from Trados and memoQ files to Word, Excel, PowerPoint and InDesign, translated every segment in Lexora, exported, and read the result back. Every file Lexora accepted came back with all its translations in place. It refused 122, and for 22 of them the reason it gave was wrong.

The files

Almost all of them come from the test files of the Okapi Framework, a free, open-source set of translation tools whose developers keep test files for every format it reads (Okapi Framework on GitLab, Apache License 2.0). The one memoQ file is a memoQ 10 export someone published on GitHub (the memoQ export). Anyone can download the same files and repeat the test.

They are not clients’ documents. They are what format developers collect: files that once caused a bug, files that show off one awkward feature, empty files and very large ones. Three things are worth knowing about them:

  • 795 Office files. 386 are test files: 348 were last saved by Microsoft Office, 24 by LibreOffice and the rest by other programs. The other 409 are Okapi’s own expected results, translated copies of those test files that Okapi wrote itself. We kept them: they are valid Office files, written by a different program.
  • 274 XML files. Most are small test files for ITS, the W3C standard for marking which parts of an XML file to translate.
  • 3 Trados packages. Two project packages, made in 2017 and 2021, and one return package that Trados made. The return package isn’t round-tripped: we keep it to compare Lexora’s return packages with Trados’s.

The test

For each file, the test does what a translator does, without the typing:

  1. Open the file in Lexora, split into sentences, as the app does by default.
  2. Translate every segment with a marked test translation: the source text with “ÜX” in front, after any formatting tags at the start.
  3. Export the file. A Trados or memoQ package becomes its return package.
  4. Open the exported file again and count the marked translations. Every one must be there. In a bilingual file the number of segments must not change, and a return package must keep every file Lexora doesn’t touch exactly as it was, byte for byte.

A file that fails any of these counts as a failure. A file Lexora refuses counts as refused only when the message tells you what is wrong. Anything else is a failure too. We ran it on 7 October 2026 with Lexora 0.5.2.

The results

Of the 1,267 files tested, 1,145 went through and back with every translation in place, and 122 were refused with a message. None failed. Together the files that went through held more than 32,000 segments.

FormatFilesRound tripRefused
Word (.docx)43340330
XML27420569
Excel (.xlsx)1831767
PowerPoint (.pptx)17916910
InDesign (.idml)90855
Trados (.sdlxliff)43430
HTML39381
PO (.po, .pot)17170
XLIFF (.xlf)660
Trados (.sdlppx)220
memoQ (.mqxliff)110

One correction to our own test belongs here. Our automated run reported 17 of the XML files as refused for “invalid XML”. Lexora had read them correctly; the exported files failed to read back because the test runs in a simulated browser, which writes XML namespaces differently from the real one. We ran those 17 files again in Chrome, the browser engine Lexora’s window is built on, and all 17 came back with every translation, so the table counts them as round trips. That is a flaw in our test: a file Lexora writes and can’t read back should count as a failure, never as a refusal.

The 122 refusals

Lexora is built to refuse a file with a clear message rather than damage it. These are the messages it gave:

  • “No translatable text was found”: 118 files. For 96 of them that is true: they are empty, hold text only in slide templates or equations, or mark all their text as not to be translated. For 22 it is not true; see below.
  • Tracked changes in InDesign: 2 files. Lexora asks you to accept or reject them in InDesign first.
  • Password-protected: 1 Word file.
  • Too large: 1 Excel file with 157,075 cells. Lexora takes up to 50,000 segments per document and asks you to split the file.

Where Lexora got it wrong

In 22 files Lexora said there was nothing to translate, and there was. Nothing was damaged, but the message was wrong, and the text in them went untranslated. Six are Okapi’s translated copies of other files in the list, so there are 16 different files:

  • Strict Open XML (1 Excel and 1 PowerPoint file). Office can save files in a stricter variant of its format, with the same .xlsx and .pptx extensions (Microsoft Learn). Lexora reads Word files saved this way but not Excel or PowerPoint ones.
  • Text only in charts or SmartArt (3 files). Lexora doesn’t translate chart and SmartArt text. When a file has other text as well, it says on import how many charts and SmartArt graphics it leaves untranslated; when that is all the file holds, it says only that it found no text.
  • A link that runs across two paragraphs (1 Word file) and a text box holding a link (1 Word file). Lexora keeps the text of Word fields as it is, and a link written as a field is one. Here that kept text you would want to translate.
  • Text only in a field’s result (1 Word file), a date Word fills in itself. Keeping that is deliberate: Word replaces it when it updates the field.
  • XML with text in elements named data (7 files). In DITA, a common format for manuals, data holds metadata, so Lexora skips it. In these files it holds the text itself.
  • Text in a CDATA section (1 XML file), a WordPress export.

Strict Open XML, links written as fields and the two XML cases are gaps in Lexora. For charts and SmartArt, the message should at least say where the text is and that Lexora won’t translate it. The field result is the one case where leaving the text alone was right, but the message should say that too.

What this test doesn’t show

  • That the exported file looks right. The test checks that every translation is there and that files Lexora doesn’t touch are unchanged. It doesn’t open the result in Word or InDesign and compare the layout.
  • Every segment. In 32 files, some segments were left out because the test translation doesn’t fit their formatting tags. A translator places tags; the test doesn’t.
  • Your files. Test files are smaller and stranger than real work. If you have a file Lexora gets wrong, send it to us, with your client’s permission, and it becomes a test.

Trados, memoQ, Microsoft Office, Word, Excel, PowerPoint, InDesign and the other product names are trademarks of their owners, used here only to name their products. Lexora Studio is independent of these companies and not endorsed by them.

← All posts