sorry, we can't preview this file

...but you can still download wiki.csv
wiki.csv (314.85 MB)

Bangla Wikipedia dataset

Download (314.85 MB)
dataset
posted on 10.12.2019, 06:41 by Aisha Khatun, Anisur Rahman, Md. Saiful Islam
A subset of the Bangla Wikipedia text. To create the Wikipedia dataset, we collected the Bangla wiki-dump of 10th June, 2019. The files are then merged and each article is selected as a sample text. All HTML tags were removed and the title of the page was stripped from the beginning of the text. This dataset contains 70377 samples with a total number of words being 18229481. The entire dataset has 1289249 unique words, which is 7% of the total vocabulary.

History

Licence

Exports

Logo branding

Licence

Exports