Skip to main content | Skip to Navigation | Text Size : | Language:

logo of Linguistic Data Consortium for Indian Languages (LDC-IL)
Released Datasets | Official Website of Linguistic Data Consortium for Indian Languages

Released Datasets of LDC-IL and their Prices

LDC-IL has so far released a total of 58+ datasets. The list of the datasets released is given below along with their prices for the commercial users.

Sl no. Name of datasets Link Prices
61 A Gold Standard Chhattisgarhi Raw Text Corpus Vol II 13207
62 A Gold Standard Kashmiri Raw Text Corpus Vol II 6932
63 A Gold Standard Maithili Raw Text Corpus Vol II 5208
64 A Gold Standard Telugu Raw Text Corpus Vol II 30249
65 Maithili Raw Speech Corpus Vol II 251520
66 Dogri Sentence Aligned Speech Corpus 39253
67 Manipuri Sentence Aligned Speech Corpus (Bengali Script) 748800
68 Manipuri Sentence Aligned Speech Corpus (Meetei Mayek) 748800
69 Punjabi Sentence Aligned Speech Corpus 335762
70 Telugu Sentence Aligned Speech Corpus 74711
71 Maithili Sentence Aligned Speech Corpus (Tirhuta Script) 274460

These datasets are distributed for both commercial and non-commercial usage.

Please note that for bonafide non-commercial and academic use, the datasets are free of charge. The requester needs to be a bonafide student/faculty/employee of a government funded research Institute or be a government entity.

Additional discounts are available for Startups, MSMEs, entitites from the SAARC countries. For more details about the discount and the procedure to procure the datasets, please login to the Data Distribution portal and see the FAQ page.