> For the complete documentation index, see [llms.txt](https://ai4commsci.gitbook.io/formosanbank/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://ai4commsci.gitbook.io/formosanbank/the-bank-architecture/corpora/ntu-paiwan-asr.md).

# NTU Paiwan ASR

### **Overview**

The NTU Paiwan ASR Corpus contains Paiwan speech data collected by Professor Li-May Sung at National Taiwan University. It covers read text and spontaneous speech collected over two field years.

The current public corpus contains 260 XML files from 16 speakers across Northern, Central, Eastern, and Southern Paiwan. Speaker names are pseudonyms, and the two field years represent distinct participants.

#### **Content Summary**

* **Language Covered**: Paiwan
* **Corpus Content**:
  * Texts with aligned audio for read speech
  * Spontaneous speech recordings
* **Data Contributors**: Sixteen speakers (pseudonyms used to maintain anonymity)

***

### Corpus Statistics

|                           | <p>Paiwan<br>Central</p> | <p>Paiwan<br>Eastern</p> | <p>Paiwan<br>Northern</p> | <p>Paiwan<br>Southern</p> |
| ------------------------- | ------------------------ | ------------------------ | ------------------------- | ------------------------- |
| Word count                | 3,632                    | 20,513                   | 27,652                    | 5,554                     |
| Total audio               | 1.5h                     | 6.6h                     | 8.1h                      | 1.5h                      |
| Transcribed               | 0.9h                     | 4.8h                     | 6.2h                      | 0.9h                      |
| Untranscribed             | 0.5h                     | 1.9h                     | 1.9h                      | 0.6h                      |
| Translated words          |                          |                          |                           |                           |
| English                   | 0                        | 0                        | 0                         | 0                         |
| Mandarin                  | 0                        | 0                        | 0                         | 0                         |
| Japanese                  | 0                        | 0                        | 0                         | 0                         |
| Dutch                     | 0                        | 0                        | 0                         | 0                         |
| Morphologically segmented | 0                        | 0                        | 0                         | 0                         |
| Glossed words             | 0                        | 0                        | 0                         | 0                         |

***

### Corpus Processing

Read-speech XML was built from ELAN sources. The second field year and spontaneous-speech data were built with `CodeAndDocs/build_y2_and_spontaneous.py`, followed by the standard cleaning, orthography, phonology, and validation steps. The public scripts apply pseudonyms and contain no real participant names.

Audio is hosted on Hugging Face rather than committed to Git. Running the corpus's `download_audio_data.sh` downloads the full recordings and uses `CodeAndDocs/extract_audio_clips.py` to recreate the sentence-level clips referenced by the XML. Complete reconstruction also requires the private raw recordings and speaker key, which are intentionally not public.

***

### **Applications**

The NTU Paiwan ASR Corpus is a crucial resource for advancing research and technology in several key areas:

* **Automated Speech Recognition (ASR)**: The aligned text and audio make this corpus particularly valuable for developing ASR systems tailored to Paiwan, enabling future tools like transcription software and voice interfaces in the language.
* **Language Revitalization**: Supports the creation of educational materials and language-learning resources for Paiwan speakers and learners.
* **Speech Technology Development**: Facilitates tools such as text-to-speech systems and pronunciation modeling.
* **Linguistic Analysis**: Enables studies on syntax, phonology, discourse structures, and other linguistic phenomena.
* **Comparative Studies**: Supports research on shared features and differences among Formosan languages.

***

### **Access Details**

* The XML, processing documentation, and audio download script are available [in FormosanBank](https://github.com/FormosanBank/FormosanBank/tree/main/Corpora/NTU_Paiwan_ASR).

***

### **Copyright**

The current public XML records this corpus as CC BY. The central FormosanBank terms and AI-use addendum also apply.

***

### **Acknowledgments**

This corpus was developed through a collaborative effort led by **Dr. Sung** at the NTU Graduate Institute of Linguistics. The project is part of the broader FormosanBank initiative and would not have been possible without the contributions of Paiwan-speaking participants. Special thanks to all collaborators and the broader Paiwan community for their support and engagement in this vital preservation effort.

***

### Citation

In accordance with our [Terms of Use](/formosanbank/additional-resources/terms-of-use.md), if you use this corpus or any product derived from this corpus in any publication, you must cite both FormosanBank and:

* Le Ferrand, É., Prud'hommeaux, E., Hartshorne, J. K., & Sung, L.-M. (2024). *NTU Paiwan ASR Corpus*. *Electronic Resource*.
