Skip to main content | Skip to Navigation | Text Size : | Language:

logo of Linguistic Data Consortium for Indian Languages (LDC-IL)
Corpora and Technological Resources for Nepali | Official Website of Linguistic Data Consortium for Indian Languages

Corpora and Technological Resources for Nepali

Maintained by : Umesh Chamling Rai
Last Updated on: 17/01/2025


Introduction

Language resources are essential for effective language study and analysis. Developing high-quality data requires significant time and effort. Innovations like writing systems, the printing press, and computers have aided in storing language information. A comprehensive listing of authentic language data and tools will be beneficial for advancing language studies. These resources are crucial for understanding the evolution and growth of a particular language, especially in enhancing research and improving language technology.

Resource Development Challenges

Language technology presents a profitable and socially beneficial solution to overcome language barriers. However, the digitization of language data poses challenges, particularly in developing accurate language models for complex languages like Nepali. The first step in enabling computers to understand human language is encoding, achieved through the Unicode scheme, which assigns unique codes for Nepali characters.

Language Resources

Numerous materials exist to analyze Nepali data and develop technologies. These resources include text and speech corpora, dictionaries, ontologies, and multimedia databases, alongside software for collection, preparation, and analysis. A linguistic corpus, which represents real-time language usage, is vital for developing various language technologies.

The main application areas of language technology include spell and grammar checking, speech recognition and synthesis, machine translation, and information retrieval. Language resources are foundational for these tools, enhancing communication and interaction between humans and computers.

Prominent Institutions and Initiatives

Linguistic Data Consortium for Indian Languages (LDC-IL)

LDC-IL, housed within the Central Institute of Indian Languages, has developed a comprehensive Nepali text and speech corpus from various sources, including books and newspapers. The data covers multiple domains, with significant contributions to language processing.

AI4Bharat

AI4Bharat, a research lab at IIT Madras, is committed to enhancing AI technology for Indian languages through open-source initiatives. The lab has developed and released an extensive set of datasets, tools, and cutting-edge models. Its focus areas are transliteration, natural language understanding and generation, translation, automatic speech recognition, and speech synthesis.

BHASHINI

BHASHINI aims to transcend language barriers, ensuring that every citizen can effortlessly access digital services in their own language. Using voice as a medium, BHASHINI has the potential to bridge language as well as the digital divide.

Indian Languages Corpora Initiative (ILCI)

The Indian Languages Corpora Initiative (ILCI), launched by TDIL, represents a significant effort to create national corpora based on standardized frameworks. In Phase 1, the initiative successfully developed parallel annotated corpora in 12 major Indian languages, including English, utilizing India's national standards for part-of-speech (POS) annotation. Following Phase 2, the total size of the corpora is now estimated to be around 27 million parallel annotated and chunked words, covering key domains such as: Health and Tourism (HT), Agriculture and Entertainment (AGENT)

Technology Development for Indian Languages (TDIL)

Initiated by the Ministry of Electronics & Information Technology, TDIL focuses on developing multilingual knowledge resources and language technology. This includes balanced corpora and machine translation systems, benefiting numerous Indian languages, including Nepali.

VAANI

Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s spoken languages to propel language AI technologies and content for an inclusive Digital India.

OPUS

A growing collection of translated texts, aligned as a parallel corpus for easy access.

Open SLR

Provides high-quality Nepali multi-speaker speech datasets, facilitating language research.

Mozilla Common Voice

Mozilla Common Voice is a publicly accessible voice dataset. Its goal is to create an open-source, multilingual collection of voices that can be used by anyone to train speech-enabled applications.

ULCA

Universal Language Contribution APIs (ULCA) is an open-source, scalable data platform that supports a variety of datasets for Indic languages. It also provides a user-friendly interface for interacting with these datasets.

HuggingFace

Hugging Face is a machine learning (ML) and data science platform and community that helps users build, deploy and train machine learning models.

OSCAR

The OSCAR project (Open Super-large Crawled Aggregated coRpus) is an Open Source project aiming to provide web-based multilingual resources and datasets for Machine Learning (ML) and Artificial Intelligence (AI) applications.

ELRA

The ELRA Language Resources Association, is a non-profit organisation whose main mission is to make Language Resources (LRs) for Human Language Technologies (HLT) available to the community at large.

CLE

Center for Language Engineering (CLE) is conducting research and development in linguistic and computational aspects of languages, specifically of Pakistan and developing Asia.

Ml Datasets

Ml Datasets an awesome Open Source to help us find open source software

FutureBeeAI

provides 2000+ datasets for AI development process.

Kaggle

The dataset contains the 34 different Nepali characters images present in the number plates of Nepali Vehicles. It contains a total of 26,537 images of characters. Researchers are free to split the set into train, test and validation as they want.



The resources available for Nepali

Resource Resource Centre Specification Source
Text corpus LDC-IL 70,57,524 words View Resource
AI4Bharat 12896.7(in Millions) words View Resource
Bhasini 117427 words View Resource
ILCI 600,000 words View Resource
ELRA 386,879 words View Resource
OSCAR 177,885,116 Words View Resource
Parallel text corpus TDIL 70,000 words View Resource
Bhasini 117427 words View Resource
ELRA 27,060 Words View Resource
CLE 4325 Total Eng Sentences Trnltd View Resource
Handwritten data Kaggle Handwritten images View Resource
Speech corpus ASR LDC-IL 87:14:44 hours View Resource
AI4Bharat 403 hours View Resource
Bhashini 419.517 hrs, grouped by domain View Resource
Open SLR 2000 Sentences View Resource
Mozilla Common Voice 1h 24m View Resource
FutureBeeAI NA View Resource
Speech corpus TTS AI4Bharat 134 hours View Resource
Speech spoken Corpus ELRA 31h 26m View Resource
Synset MeitY 11,713 Words View Resource

Conclusion

The development of language resources for Nepali is a collaborative effort involving various institutions and initiatives. These resources not only support the advancement of language technology but also preserve and promote the rich linguistic heritage of Nepali.