Exploring clustering of Philippine languages in multilingual neural machine translation

Clicks: 2
ID: 287695
2022
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #3,724 of 3,757 articles by views in Malay Journal

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 3,757 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Multilingual neural machine translation (MNMT) is a single model capable of translating several language directions. This has been shown to aid in translating low-resource languages such as Philippine languages. Moreover, there is also the empirical observation that clustering linguistically similar languages together can further aid translation performance. Based on these, we propose to develop multilingual Filipino neural machine translation systems wherein Philippine languages are clustered into different groups based on previous computational works and in linguistics studies. As such, several cluster-specific models were built, based around language families or other computational frameworks centered around the language relatedness of English, and eight (8) Philippine languages. Explorations were also made as to the choice of pivot language when performing pivot-based translation. Experiments show that the use of language clusters give comparable to or higher translation scores than using a baseline universal model, such as Tagalog, Cebuano and Hiligaynon being more likely to perform better with each other. Experiments also show that pivot-based translation still scores higher than zero-shot translation, and that English is still the best pivot to be used in a universal translation model setting. Finally, some issues are discussed with regards to the use of conventional automatic metrics on translation outputs concerning Philippine languages.
Reference Key
persistent_1760662380_68f1936c2a7c4 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Coronia, Jeremy Dale O
Journal Malay Journal
Year 2022
DOI
DOI not found
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.