 Corpus Name: SCB_MT_EN_TH
     Package: SCB_MT_EN_TH/v1.0/raw/en
     Website: http://opus.nlpl.eu/SCB_MT_EN_TH-v1.0.php
     Release: v1.0
Release date: 23 June 2020
   OPUS date: Tue Sep 19 17:11:12 EEST 2023
     License: <a href=https://creativecommons.org/licenses/by-sa/4.0>CC BY-SA 4.0</a>
   Copyright: Information about the sources is available <a href="https://github.com/vistec-AI/dataset-releases/releases/tag/scb-mt-en-th-2020_v1.0">here</a>

This package is part of OPUS - the open collection of parallel corpora
OPUS Website: http://opus.nlpl.eu

Please <a href="http://opus.lingfil.uu.se/LREC2012.txt">cite the following article</a> if you use the OPUS packages and downloads in your own work:<br/> J. Tiedemann, 2012, <a href="http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf"><i>Parallel Data, Tools and Interfaces in OPUS.</i></a> In Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC 2012)<br/>

The AI Research Institute of Thailand (AIResearch), with the collaboration between Vidyasirimedhi Institute of Science and Technology (VISTEC) and Digital Economy Promotion Agency (depa), published this open English-Thai machine translation dataset, with the sponsorship from Siam Commercial Bank (SCB), namely scb-mt-en-th-2020. The dataset contains parallel sentences from various sources such as task-based conversation, organization websites, Wikipedia articles, and government documents. To obtain parallel sentences, they hired professional and crowdsourced translators and build a module to automatically align parallel sentence pairs from documents, articles, and web pages.
Simple example corpus
