 Corpus Name: Samanantar
     Package: Samanantar.en-as in Moses format
     Website: http://opus.nlpl.eu/Samanantar-v0.2.php
     Release: v0.2
Release date: Fri Dec 23 01:31:35 EET 2022
     License: CC0
   Copyright: Samanantar is released under this licensing scheme: <ul><li>We do not own any of the text from which this data has been extracted.</li> <li>We license the actual packaging of this data under the Creative Commons CC0 license (“no rights reserved”).</li> <li>To the extent possible under law, AI4Bharat has waived all copyright and related or neighboring rights to Samanantar</li> <li>This work is published from: India.</li></ul>

This package is part of OPUS - the open collection of parallel corpora
OPUS Website: http://opus.nlpl.eu

If you are using any of the resources, please cite the following article: <blockquote><pre> @misc{ramesh2021samanantar,<br/> title={Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages}, <br/> author={Gowtham Ramesh and Sumanth Doddapaneni and Aravinth Bheemaraj and Mayank Jobanputra and Raghavan AK and Ajitesh Sharma and Sujit Sahoo and Harshita Diddee and Mahalakshmi J and Divyanshu Kakwani and Navneet Kumar and Aswin Pradeep and Srihari Nagaraj and Kumar Deepak and Vivek Raghavan and Anoop Kunchukuttan and Pratyush Kumar and Mitesh Shantadevi Khapra},<br/> year={2021},<br/> eprint={2104.05596},<br/> archivePrefix={arXiv},<br/> primaryClass={cs.CL}<br/> }<br/> </pre></blockquote>

Samanantar is the largest publicly available parallel corpora collection for Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.6M sentence pairs between English to Indian Languages. The data is distributed by AI4BHARAT

