Skip to main content | Skip to Navigation | Text Size : | Language:

logo of Linguistic Data Consortium for Indian Languages (LDC-IL)
Corpora and Technological Resources for Telugu | Official Website of Linguistic Data Consortium for Indian Languages

Corpora and Technological Resources for Telugu

Maintained by : Modugu Kasimbabu
Last Updated on: 17/01/2025


Introduction

Language resources are essential for effective language study and analysis. Developing high-quality data requires significant time and effort. Innovations like writing systems, the printing press, and computers have aided in storing language information. A comprehensive listing of authentic language data and tools will be beneficial for advancing language studies. These resources are crucial for understanding the evolution and growth of a particular language, especially in enhancing research and improving language technology.

Resource Development Challenges

Language technology presents a profitable and socially beneficial solution to overcome language barriers. However, the digitization of language data poses challenges, particularly in developing accurate language models for complex languages like Telugu. The first step in enabling computers to understand human language is encoding, achieved through the Unicode scheme, which assigns unique codes for Telugu characters.

Language Resources

Numerous materials exist to analyze Telugu data and develop technologies. These resources include text and speech corpora, dictionaries, ontologies, and multimedia databases, alongside software for collection, preparation, and analysis. A linguistic corpus, which represents real-time language usage, is vital for developing various language technologies.

The main application areas of language technology include spell and grammar checking, speech recognition and synthesis, machine translation, and information retrieval. Language resources are foundational for these tools, enhancing communication and interaction between humans and computers.

Prominent Institutions and Initiatives

Linguistic Data Consortium for Indian Languages (LDC-IL)

LDC-IL, housed within the Central Institute of Indian Languages, has developed a comprehensive Telugu text and speech corpus from various sources, including books and newspapers. The data covers multiple domains, with significant contributions to language processing.

AI4Bharat

AI4Bharat, a research lab at IIT Madras, is committed to enhancing AI technology for Indian languages through open-source initiatives. The lab has developed and released an extensive set of datasets, tools, and cutting-edge models. Its focus areas are transliteration, natural language understanding and generation, translation, automatic speech recognition, and speech synthesis.

Indian Languages Corpora Initiative (ILCI)

The Indian Languages Corpora Initiative (ILCI), launched by TDIL, represents a significant effort to create national corpora based on standardized frameworks. In Phase 1, the initiative successfully developed parallel annotated corpora in 12 major Indian languages, including English, utilizing India's national standards for part-of-speech (POS) annotation. Following Phase 2, the total size of the corpora is now estimated to be around 27 million parallel annotated and chunked words, covering key domains such as: Health and Tourism (HT), Agriculture and Entertainment (AGENT)

Technology Development for Indian Languages (TDIL)

Initiated by the Ministry of Electronics & Information Technology, TDIL focuses on developing multilingual knowledge resources and language technology. This includes balanced corpora and machine translation systems, benefiting numerous Indian languages, including Telugu.

OPUS

A growing collection of translated texts, aligned as a parallel corpus for easy access.

Open SLR

Provides high-quality Telugu multi-speaker speech datasets, facilitating language research.

Kaggle

A platform for sharing and discovering datasets, including Telugu speech datasets.

FutureBeeAI

provides 2000+ datasets for AI development process.

Elra

The ELRA Language Resources Association, is a non-profit organisation whose main mission is to make Language Resources (LRs) for Human Language Technologies (HLT) available to the community at large. To achieve this goal, ELRA carries out a wide variety of activities around LRs, including Identification & Distribution, Production & Validation, Technology Evaluation, Information Dissemination on HLT.

Common Voice

Common Voice is a publicly accessible voice dataset. Its goal is to create an open-source, multilingual collection of voices that can be used by anyone to train speech-enabled applications.

Sketchengine

Sketch Engine is the ultimate tool to explore how language works. Its algorithms analyze authentic texts of billions of words (text corpora) to identify instantly what is typical in language and what is rare, unusual or emerging usage. It is also designed for text analysis or text mining applications. Sketch Engine is used by linguists, lexicographers, translators, students and teachers. It is a first choice solution for publishers, universities, translation agencies and national language institutes throughout the world.Sketch Engine contains 1 trillion words in 800 ready-to-use corpora in 100+ languages, each having a size of up to 80 billion words to provide a truly representative sample of language.

Open-Speech-EkStep/ULCA- ASR-dataset-corpus

Universal Language Contribution APIs (ULCA) is an open-source, scalable data platform that supports a variety of datasets for Indic languages. It also provides a user-friendly interface for interacting with these datasets.

LERC-UoH

Language Engineering Research Centre at University of Hyderabad is now has a nearly 40 Million word text corpus for Telugu, its own shallow parsing architecture and a wide coverage shallow parser for English, a speaker independent continuous speech recognition system for Telugu and a host of data resources, tools and technologies in many related areas.

Leipzig Corpora Collection

The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs. Leipzig Corpora Collection: Telugu Web text corpus (India) based on material from 2019.

IIIT Hyderabad

The corpus "Sentiraama" was created by G.Rama Rohit Reddy at Language Technologies Research Centre, KCIS, IIIT Hyderabad. The corpus consists of 4 datasets annotated using a 2-value scale, distinguishing between positive and negative sentiment at document level. Corpus consists of datasets from multiple domains such as book reviews, product reviews, movie reviews and song lyrics. Each of them were annotated by the annotators carefully. Telugu Corpus Statistics: Documents-1006; Sentences-46972; Words-298630.

IndicCorp

IndicCorp is a large monolingual corpora with around 9 billion tokens covering 12 of the major Indian languages. It has been developed by discovering and scraping thousands of web sources - primarily news, magazines and books, over a duration of several months. Languages covered: Assamese, Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. Corpus Format: The corpus is a single large text file containing one sentence per line. The publicly released version is randomly shuffled, untokenized and deduplicated.

Oscar

The OSCAR project (Open Super-large Crawled Aggregated coRpus) is an Open Source project aiming to provide web-based multilingual resources and datasets for Machine Learning (ML) and Artificial Intelligence (AI) applications. The project focuses specifically in providing large quantities of unannotated raw data that is commonly used in the pre-training of large deep learning models. The OSCAR project has developed high-performance data pipelines specifically conceived to classify and filter large amounts of web data. The project has also put special attention in improving the data quality of web-based corpora as well as providing data for low-resource languages, so that these new ML/AI technologies are accessible to as many communities as possible.

Bhashini

BHASHINI aims to transcend language barriers, ensuring that every citizen can effortlessly access digital services in their own language. Using voice as a medium, BHASHINI has the potential to bridge language as well as the digital divide. Launched by Honourable PM Shri Narendra Modi in July 2022 under the National Language Technology Mission, BHASHINI aims to provide technology translation services in 22 scheduled Indian languages.

NPLT

National Platform for Language Technology (NPLT) platform for academia, researcher as well as industry to provide access to Indian Language Data, Tools and related web services. This Platform serves as a marketplace of linguistic resources, tools and services developed either by Government, Startups, Industry and other stakeholders. Startups, MSMEs, International Academic Researchers, MNCs and Foreign Entities having an interest in Natural Language Processing can avail Indian Languages resources from this portal.

DATAOCEAN

DATAOCEAN AI (formerly Speechocean) is dedicated to providing training data to improve your AI and machine learning systems. We collect and label images, text, speech, audio, video, and other data used to build and continuously improve the world’s most innovative artificial intelligence systems.

Kaldi

Kaldi is an open-source toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers. For more detailed history and list of contributors see History of the Kaldi project.

IndicASR

IndicASR is built on top of wav2vec2 XLSR-53 and Huggingface's transformers and has pre-trained models for Telugu in the current release. The Telugu model is trained on the train set of MSR Indic corpus + a private corpus of ~94 hours obtained from various telugu interview playlists from Youtube.

Telugu-LLM-Labs

The primary objectives of Telugu LLM Labs involve contributing to open datasets with a specific focus on Telugu, encompassing both native script and romanised versions. For this, the initiative aims to share their experiments, including models and embeddings, concentrating on LLMs and tailored for Telugu. The Romanized Telugu Pretraining Dataset, recognising the prevalent use of romanized Telugu in online conversations, such as WhatsApp or YouTube comments, Telugu LLM Labs unveils the “uonlp_culturaX_telugu_romanized_100k” dataset. Comprising the romanized version of the initial 108,000 rows from the culturaX_telugu dataset, this resource addresses the scarcity of datasets for additional pre-training on models or fine tuning, particularly in romanized Telugu.



The resources available for Telugu

Resource Resource Centre Specification Source
Text corpus LDC-IL 30,10,993 words View Resource
LERC-UoH 40 million words View Resource
MeitY Approx 21,091 Synsets View Resource
Leipzig Corpus Collection 7,50,632 sentences View Resource
IIIT Hyderabad 2,98,630 words View Resource
IndicCorp 47.9M Sentences View Resource
Elra 951,000 words – 43,000 parallel sentences View Resource
Oscar 137,752,065 words View Resource
AI4 Bharat 16,278.7 Millions View Resource
Bhashini English-Telugu parallel sentences 6,408,913 View Resource
IndoWordNet 21091 Words View Resource
Parallel text corpus ILCI 81,000 words View Resource
OPUS English to Telugu 54,363,342 Sentences View Resource
FuturebeeAI English to Telugu 600000 words View Resource
Speech corpus ASR LDC-IL 22:43:59 hours View Resource
NPLT - ASR Consortia 91:49:05 hours View Resource
TDIL 5.4GB corpus View Resource
ELRA 118 hours View Resource
Dataocean.ai ASR 118 hours View Resource
IIIT Hyderabad -(NLTM) 2000.8 hours View Resource
Open SLR NA View Resource
Common Voice 2000 sentences with 168 speakers View Resource
Open-Speech-EkStep/ULCA- ASR-dataset-corpus 1025.93 hours View Resource
AI4 Bharat 136.40 Hours View Resource
Bhashini ASR Dataset 3,053.676 value ASR Unlabeled Dataset 850.116 value View Resource
Kaldi NA View Resource
IndicASR 94 hours View Resource
C-DAC WinRAR ZIP, File Size: 8.22 MB View Resource
Kaggle 117 datasets View Resource
FuturebeeAI 300 above Hours, 15+ datasets View Resource
Telugu-LLM-Labs NA View Resource

Conclusion

The development of language resources for Telugu is a collaborative effort involving various institutions and initiatives. These resources not only support the advancement of language technology but also preserve and promote the rich linguistic heritage of Telugu.