POS Tagger Improvisation with the Addition of Foreign Word Labels on Telkom University News
Abstract
News is a medium of daily information usually obtained by the public. The news consists of a lot of information in it and is composed of sentence structures. Each language is unique with its own sentence structure, like Indonesian and other foreign languages. But nowadays, many media mix Indonesian with foreign languages, making the sentence structure different from Bahasa Indonesia. To classify these words, Part Of Speech Tagging needed to determine the class of words composed of sentences by learning from the Corpus of each language. With the new sentence structure, POS Tagger requires a larger Corpus to learn. The language structure can determine the results of tagging from the POS Tagger. If there are words that are not in the Corpus, it can reduce the accuracy of the POS Tagger. We conducted to enhance the research results by adding data with a different sentence structure from the Indonesian Language Corpus using sentences from online media. Added about 242 sentences with 7,043 tokens on Corpus focused on Foreign Word tags, which total 3819 tags. After doing some testing and scenarios, the results of the accuracy of POS Tagger show an accuracy of 94.7% using the Hidden Markov Model method with the F1-Score tag FW 78%.
Downloads
References
A. Y. Rofiqi, “Clustering Berita Olahraga Berbahasa Indonesia Menggunakan Metode K-Medoid Bersyarat,” J. Simantec, vol. 6, no. 1, pp. 25–32, 2017.
Badan Pengembangan dan Pembinaan Bahasa, Badan pengembangan dan Pembinaan Bahasa Kementerian pendidikan dan kebudayaan. 2017.
A. Z. Amrullah, R. Hartanto, and I. W. Mustika, “A comparison of different part-of-speech tagging technique for text in Bahasa Indonesia,” Proc. - 2017 7th Int. Annu. Eng. Semin. Ina. 2017, 2017, doi: 10.1109/INAES.2017.8068538.
S. K. Nambiar, A. Leons, S. Jose, and Arunsree, “POS Tagger for Malayalam using Hidden Markov Model,” Proc. 2nd Int. Conf. Smart Syst. Inven. Technol. ICSSIT 2019, no. Icssit, pp. 957–960, 2019, doi: 10.1109/ICSSIT46314.2019.8987786.
N. Sabloak, “Part-of-Speech (POS) Tagging Bahasa Indonesia Menggunakan Algoritma Viterbi,” no. x, pp. 1–11, 2016.
Ryan Armiditya Pratama, A. A. Suryani, and W. Maharani, “Part of Speech Tagging for Javanese Language with Hidden Markov Model,” J. Comput. Sci. Informatics Eng., vol. 4, no. 1, pp. 84–91, 2020, doi: 10.29303/jcosine.v4i1.346.
I. G. M. H. Pradiptha and N. A. Sanjaya ER, “Building Balinese Part-of-Speech Tagger Using Hidden Markov Model (HMM),” JELIKU (Jurnal Elektron. Ilmu Komput. Udayana), vol. 9, no. 2, p. 303, 2020, doi: 10.24843/jlk.2020.v09.i02.p18.
A. Dinakaramani, F. Rashel, A. Luthfi, and R. Manurung, “Designing an Indonesian part of speech tagset and manually tagged Indonesian corpus,” Proc. Int. Conf. Asian Lang. Process. 2014, IALP 2014, pp. 66–69, 2014, doi: 10.1109/IALP.2014.6973519.
P. Alva and V. Hegde, “Hidden Markov model for POS tagging in word sense disambiguation,” 2016 Int. Conf. Comput. Syst. Inf. Technol. Sustain. Solut. CSITSS 2016, pp. 279–284, 2016, doi: 10.1109/CSITSS.2016.7779371.
F. M. Hasan, N. Uzzaman, and M. Khan, “Comparison of different POS tagging techniques (n-gram, HMM and brill’s tagger) for Bangla,” Adv. Innov. Syst. Comput. Sci. Softw. Eng., pp. 121–126, 2007, doi: 10.1007/978-1-4020-6264-3_23.
F. Ramadhanti, Y. Wibisono, and R. A. Sukamto, “Analisis Morfologi untuk Menangani Out-of-Vocabulary Words pada Part-of-Speech Tagger Bahasa Indonesia Menggunakan Hidden Markov Model,” J. Linguist. Komputasional, vol. 2, no. 1, p. 6, 2019, doi: 10.26418/jlk.v2i1.13.
S. Rathod and S. Govilkar, “Survey of various POS tagging techniques for Indian regional languages,” Int. J. Comput. Sci. Inf. Technol., vol. 6, no. 3, pp. 2525–2529, 2015, [Online]. Available: www.ijcsit.com.
D. E. Cahyani and M. J. Vindiyanto, “Indonesian part of speech tagging using hidden markov model - Ngram viterbi,” 2019 4th Int. Conf. Inf. Technol. Inf. Syst. Electr. Eng. ICITISEE 2019, pp. 353–358, 2019, doi: 10.1109/ICITISEE48480.2019.9003989.
V. M. Patro and M. Ranjan Patra, “Augmenting Weighted Average with Confusion Matrix to Enhance Classification Accuracy,” Trans. Mach. Learn. Artif. Intell., vol. 2, no. 4, 2014, doi: 10.14738/tmlai.24.328.
F. Rahmad, Y. Suryanto, and K. Ramli, “Performance Comparison of Anti-Spam Technology Using Confusion Matrix Classification,” IOP Conf. Ser. Mater. Sci. Eng., vol. 879, no. 1, 2020, doi: 10.1088/1757-899X/879/1/012076.
Bila bermanfaat silahkan share artikel ini
Berikan Komentar Anda terhadap artikel POS Tagger Improvisation with the Addition of Foreign Word Labels on Telkom University News
Pages: 588-594
Copyright (c) 2022 Winkie Setyono, Donni Richasdy, Mahendra Dwifebri Purbolaksono

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under Creative Commons Attribution 4.0 International License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (Refer to The Effect of Open Access).





















