For the complete documentation index, see llms.txt. This page is also available as Markdown.

Corpora

Welcome to the FormosanBank Corpora section! Here, you’ll find comprehensive documentation of the corpora used in FormosanBank. Each corpus in this collection represents a unique linguistic dataset, encompassing various types of text and audio recordings. Our corpora are designed to support linguistic research, language education, and revitalization efforts, making these endangered languages accessible and analyzable for researchers, educators, and community members alike. Below is a list of the current corpora of which FormosanBank consists:


Corpus Statistics

By language

Truku
Amis
Bunun
Babuza-Favorlang
Kavalan
Rukai
Siraya
Paiwan
Puyuma
Pazeh
Thao
Saaroa
Sakizaya
Yami
Atayal
Seediq
Tsou
Kanakanavu
Saisiyat

Word count

87,457

2,576,762

249,946

22,112

127,016

275,642

43,173

422,746

230,655

356

165,361

63,050

1,751,644

119,295

887,478

1,009,622

88,015

114,294

101,824

Total audio

84.4h

121.6h

125.4h

0

53.8h

154.3h

0

168.9h

119.9h

0

41.4h

45.9h

49.0h

40.5h

171.2h

92.2h

41.0h

65.1h

47.0h

Transcribed

41.6h

121.2h

125.4h

0

53.8h

154.3h

0

134.1h

119.9h

0

41.4h

45.9h

49.0h

40.5h

160.0h

92.1h

41.0h

65.1h

47.0h

Untranscribed

42.8h

0.4h

0

0

0

0

0

34.8h

0

0

0

0

0

0

11.2h

0.1h

0

0

0

Translated words

English

11,370

126,324

66,402

9,561

27,300

65,067

43,161

74,111

40,780

0

74,694

8,760

19,888

25,478

67,520

50,350

21,413

24,418

22,156

Mandarin

87,325

533,075

249,762

0

126,985

274,819

41,851

210,690

230,491

0

101,796

63,034

101,445

118,409

307,673

217,319

87,978

114,252

101,160

Japanese

0

0

0

0

0

0

0

0

0

351

0

0

0

0

0

0

0

0

0

Dutch

0

0

0

0

0

0

42,913

0

0

0

0

0

0

0

0

0

0

0

0

Morphologically segmented

1,130

7,682

23,463

0

16,286

17,344

2

29,031

0

0

56,656

0

12,968

14,167

5,861

22,803

10,200

22,853

10,849

Glossed words

1,047

7,682

23,463

0

16,286

17,344

0

24,556

0

0

211

0

12,968

14,166

5,861

22,803

10,195

22,846

10,848

By dialect

Amis Coastal

Amis Hengchun

Amis Malan

Amis Southern

Amis Xiuguluan

Amis unknown

Bunun Junqun

Bunun Kaqun

Bunun Luanqun

Bunun Tanqun

Bunun Zhuoqun

Babuza-Favorlang Favorlang

Rukai Dawu

Rukai Dona

Rukai Eastern

Rukai Maolin

Rukai Wanshan

Rukai Wutai

Paiwan Central

Paiwan Eastern

Paiwan Northern

Paiwan Southern

Paiwan unknown

Puyuma Jianhe

Puyuma Nanwang

Puyuma Xiqun

Puyuma Zhiben

Pazeh unknown

Atayal FourSeasons

Atayal Sekolik

Atayal Wanda

Atayal Wenshui

Atayal YilanZeaol

Atayal Zeaol

Atayal unknown

Seediq DeluValley

Seediq Duda

Seediq Tegudaya

Seediq unknown

Word count

310,205

36,726

35,979

36,318

85,137

2,072,397

131,540

29,637

30,240

29,700

28,829

22,112

29,959

28,651

30,129

27,713

24,434

134,756

44,069

60,270

134,637

58,923

124,847

35,216

118,867

37,017

39,555

356

38,123

140,278

30,977

42,774

40,168

39,189

555,969

39,565

41,166

136,271

792,620

Total audio

29.0h

20.4h

16.8h

18.5h

36.6h

0.4h

60.3h

15.9h

15.8h

17.1h

16.2h

0

18.9h

17.8h

18.8h

17.5h

16.8h

64.5h

21.7h

33.4h

62.5h

26.6h

24.7h

18.9h

59.6h

19.8h

21.6h

0

18.3h

72.7h

17.5h

21.2h

20.6h

19.8h

1.1h

23.3h

20.2h

48.6h

0.1h

Transcribed

29.0h

20.4h

16.8h

18.5h

36.6h

0

60.3h

15.9h

15.8h

17.1h

16.2h

0

18.9h

17.8h

18.8h

17.5h

16.8h

64.5h

20.6h

29.6h

58.7h

25.1h

0

18.9h

59.6h

19.8h

21.6h

0

18.3h

62.5h

17.5h

21.2h

20.6h

19.8h

0

23.3h

20.2h

48.6h

0

Untranscribed

0

0

0

0

0

0.4h

0

0

0

0

0

0

0

0

0

0

0

0

1.1h

3.8h

3.8h

1.5h

24.7h

0

0

0

0

0

0

10.1h

0

0

0

0

1.1h

0

0

0

0.1h

Translated words

English

17,367

9,425

9,306

10,323

20,258

59,645

33,027

8,106

8,611

8,121

8,537

9,561

7,576

7,995

8,183

7,320

6,537

27,456

16,251

12,508

19,027

25,203

1,122

9,218

11,536

9,887

10,139

0

9,456

12,665

9,145

16,464

9,497

9,716

577

9,356

10,044

30,814

136

Mandarin

309,016

36,710

35,979

35,968

84,780

30,622

131,442

29,621

30,224

29,684

28,791

0

29,935

28,262

30,061

27,485

24,367

134,709

36,702

38,352

97,721

37,915

0

35,182

118,777

36,993

39,539

0

38,073

116,617

30,913

42,745

40,152

39,173

0

39,428

41,085

136,246

560

Japanese

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

351

0

0

0

0

0

0

0

0

0

0

0

Dutch

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

0

Morphologically segmented

7,682

0

0

0

0

0

23,463

0

0

0

0

0

0

29

0

23

0

17,292

3,719

1,198

9,021

13,971

1,122

0

0

0

0

0

0

0

0

5,861

0

0

0

0

0

22,803

0

Glossed words

7,682

0

0

0

0

0

23,463

0

0

0

0

0

0

29

0

23

0

17,292

3,719

1,198

9,021

9,496

1,122

0

0

0

0

0

0

0

0

5,861

0

0

0

0

0

22,803

0


Coming Soon

In addition to the wide range of corpora already incorporated into FormosanBank, there is a large number of further corpora that permission to include in FormosanBank has been obtained and they are being processed at the moment. Below are some of these corpora:

  • Matthew's Gospel and John's Gospel (Siraya)

  • The Sedik Language of Formosa by Erin Asai (Seediq)

  • Chang's Seediq Reference Grammar (Seediq)

  • Chang's Kavalan reference grammar (Kavalan)

  • Poinsot Amis Dictionary (Amis)

  • Moedict Amis (Amis)

  • Asai's Seediq Language of Formosan (Seediq)

  • ​hala saku la (videos - Atayal)

  • ​hala saku la (text - Atayal)

  • Tung's Descriptive study of Tsou (Tsou)

  • Jeng (1992) Topic and focus in Bunun (Bunun)

  • Blust's Thao Dictionary (Thao)

Last updated