> For the complete documentation index, see [llms.txt](https://ai4commsci.gitbook.io/formosanbank/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://ai4commsci.gitbook.io/formosanbank/the-bank-architecture/corpora/glosbe.md).

# Glosbe

## **Overview**

This corpus contains dictionary and translation-memory entries collected from the crowdsourced online dictionary Glosbe. The current public corpus includes Amis, Atayal, Truku, and Saisiyat data. Most entries are translated into English, with an additional Amis-Mandarin collection.

***

## Corpus Statistics

|                           | <p>Amis<br>unknown</p> | <p>Atayal<br>unknown</p> | Truku | Saisiyat |
| ------------------------- | ---------------------- | ------------------------ | ----- | -------- |
| Word count                | 89,909                 | 577                      | 116   | 600      |
| Total audio               | 0                      | 0                        | 0     | 0        |
| Transcribed               | 0                      | 0                        | 0     | 0        |
| Untranscribed             | 0                      | 0                        | 0     | 0        |
| Translated words          |                        |                          |       |          |
| English                   | 59,287                 | 577                      | 116   | 600      |
| Mandarin                  | 30,622                 | 0                        | 0     | 0        |
| Japanese                  | 0                      | 0                        | 0     | 0        |
| Dutch                     | 0                      | 0                        | 0     | 0        |
| Morphologically segmented | 0                      | 0                        | 0     | 0        |
| Glossed words             | 0                      | 0                        | 0     | 0        |

***

## **Corpus Processing**

The eight published XML files separate lexical entries from translation-memory entries and are organized by ISO 639-3 language code. The original Glosbe crawl is a retained snapshot and cannot be re-run; everything downstream of the published original tier is regenerated by `CodeAndDocs/make_xml.sh`, which runs three steps:

1. **Cleaning and quote correction** (`clean_xml.py`): normalizes punctuation and Unicode debris in the original tier and translations without changing the source spelling, and applies the Amis apostrophe-vs-quotation-mark correction described under Corpus Notes.
2. **Standardization** (`standardize.py`, once per language): rebuilds the `kindOf="standard"` FORM tier in FormosanBank's Ortho113 orthography using per-language conversion tables; stress accents (e.g. on Truku `dálix`) are deleted from the standard tier only, while the original tier keeps them exactly.
3. **Phonology** (`add_phonology.py`, once per language): generates IPA `PHON` tiers for both the original FORM (from the source orthography) and the standard FORM (from Ortho113).

Where one source form has several distinct translations, the additional ones are kept as separate `TRANSL` elements marked `ver="alt"`. This includes the restored Amis-Traditional Chinese translation memory, rebuilt reproducibly from a reviewed conversion contributed by Joseph Lin.

***

## **Corpus Notes**

Because this is a crowdsourced dictionary and translation memory, the linguistic accuracy and the rights attached to individual entries can vary. Treat entries as leads for review rather than as expert-verified dictionary records.

* In the Amis files, the letter `'` (the glottal stop) is also used by the source as a quotation mark. Disambiguating apostrophe vs. quotation mark in the Amis data was done **automatically**: a classifier compares each sentence against its translations and an attested-word list and rewrites apostrophes it judges to be quotation marks to `"`. Occasional false positives are possible and misses are certain — some real quotation marks remain written as `'`, and a few glottal-stop letters may have been rewritten to `"`.
* Every such correction is logged with the full before/after sentence context in `CodeAndDocs/quote_corrections.csv`; the XML as it stood before any automated correction is preserved in `CodeAndDocs/pre_correction_snapshot/`.
* Glosbe entries generally do not identify a dialect, and a single file likely mixes dialects; the standardization assumptions per language are documented in the corpus README.

***

## **Access Details**

* The repo containing the Glosbe corpus in FormosanBank as well as the code to reconstruct the corpus can be found [here](https://github.com/FormosanBank/FormosanBank/tree/main/Corpora/Glosbe).

***

## **Copyright**

The current XML records Glosbe and its contributors or source corpora as rights holders. The material was collected for private research review, and redistribution remains subject to Glosbe's terms and any third-party source rights. This is not a blanket Creative Commons corpus.

***

## Citation

In accordance with our [Terms of Use](/formosanbank/additional-resources/terms-of-use.md), if you use this corpus or any product derived from this corpus in any publication, you must cite both FormosanBank and:

* Glosbe. (2026). *Glosbe \[language]-\[translation language] dictionary and translation memory*. Cite the language pair, URL, and retrieval date recorded in the XML file used.
