Skip to main content | Skip to Navigation | Text Size : | Language:

logo of Linguistic Data Consortium for Indian Languages (LDC-IL)
Released Datasets | Official Website of Linguistic Data Consortium for Indian Languages

Released Datasets of LDC-IL and their Prices

LDC-IL has so far released a total of 240 datasets. The list of the datasets released is given below along with their prices for the commercial users.

Sl no. Name of datasets Link Prices
1 Maithili Parts of Speech Annotated Corpus
2 Bengali Parts of Speech Annotated Corpus
3 Bodo Parts of Speech Annotated Corpus
4 Gujarati Parts of Speech Annotated Corpus
5 Hindi Parts of Speech Annotated Corpus
6 Kannada Parts of Speech Annotated Corpus
7 Kashmiri Parts of Speech Annotated Corpus
8 Konkani Parts of Speech Annotated Corpus
9 Malayalam Parts of Speech Annotated Corpus
10 Manipuri Parts of Speech Annotated Corpus
11 Nepali Parts of Speech Annotated Corpus
12 Assamese Parts of Speech Annotated Corpus
13 Urdu Parts of Speech Annotated Corpus
14 Telugu Parts of Speech Annotated Corpus
15 Punjabi Parts of Speech Annotated Corpus

These datasets are distributed for both commercial and non-commercial usage.

Please note that for bonafide non-commercial and academic use, the datasets are free of charge. The requester needs to be a bonafide student/faculty/employee of a government funded research Institute or be a government entity.

Additional discounts are available for Startups, MSMEs, entitites from the SAARC countries. For more details about the discount and the procedure to procure the datasets, please login to the Data Distribution portal and see the FAQ page.