Research notes: preserving layout when transforming documents
Translating or anonymizing a document is easy once it becomes plain running text. The hard part is giving the document back looking the same.
- Category
- Research
- Author
- Equipe Devx
- Reading time
- 1 min
One of the questions guiding Lúmen's research sounds simple: after translating or anonymizing a document, does it still look like the same document?
Why it matters
For whoever uses the result, layout is information. The position of a signature, the alignment of a table, the hierarchy of a contract — all of it carries meaning. Returning plain running text forces someone to rebuild the document by hand, which cancels out much of the automation.
What makes the problem hard
- Text changes length. A translated sentence can be 30% longer and no longer fit in the original space.
- Scanned documents have no structure. Before preserving the layout, you have to rebuild it from the image.
- Tables and forms break easily. Merged cells, borderless columns and handwritten fields confuse the reading.
- Anonymizing isn't just deleting. A badly placed redaction box can hide too much or too little.
How we're approaching it
The current line of investigation treats the document as two layers: content and geometry. We first rebuild the geometry — blocks, lines, tables, positions — and only then transform the content, respecting the boundaries of each block.
This is a research post, not a product post: it describes a problem under study, not an available feature. When there's a usable result, it will show up in Lúmen's changelog.
- #research
- #lumen
- #layout