Nalazite se na CroRIS probnoj okolini. Ovdje evidentirani podaci neće biti pohranjeni u Informacijskom sustavu znanosti RH. Ako je ovo greška, CroRIS produkcijskoj okolini moguće je pristupi putem poveznice www.croris.hr
izvor podataka: crosbi !

Evaluation of Croatian Word Embeddings (CROSBI ID 661764)

Prilog sa skupa u zborniku | izvorni znanstveni rad | međunarodna recenzija

Svoboda, Lukáš ; Beliga, Slobodan Evaluation of Croatian Word Embeddings // Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) / Calzolari, N. ; Choukri, K. ; Cieri, C. et al. (ur.). Pariz: European Language Resources Association (ELRA), 2018. str. 1512-1518

Podaci o odgovornosti

Svoboda, Lukáš ; Beliga, Slobodan

engleski

Evaluation of Croatian Word Embeddings

Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy dataset based on the original English Word2vec word analogy dataset and added some of the specific linguistic aspects from the Croatian language. Next, we created Croatian WordSim353 and RG65 datasets for a basic evaluation of word similarities. We compared created datasets on two popular word representation models, based on Word2Vec tool and fastText tool. Models have been trained on 1.37B tokens training data corpus and tested on a new robust Croatian word analogy dataset. Results show that models are able to create meaningful word representation. This research has shown that free word order and the higher morphological complexity of Croatian language influences the quality of resulting word embeddings.

Croatian word embeddings ; Croatian word analogy ; Croatian language ; Slavic language family ; Word2Vec ; FastText ; Croatian word similarity dataset ; WordSim353 ; RG65

nije evidentirano

nije evidentirano

nije evidentirano

nije evidentirano

nije evidentirano

nije evidentirano

Podaci o prilogu

1512-1518.

2018.

objavljeno

Podaci o matičnoj publikaciji

Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)

Calzolari, N. ; Choukri, K. ; Cieri, C. ; Declerck, T. ; Goggi, S. ; Hasida, K. ; Isahara, H. ; Maegaard, B. ; Mariani, J. ; Mazo, H. ; Moreno, A. ; Odijk, J. ; Piperidis, S. ; Tokunaga, T.

Pariz: European Language Resources Association (ELRA)

979-10-95546-00-9

Podaci o skupu

11th International Conference on Language Resources and Evaluation (LREC 2018)

predavanje

07.05.2018-12.05.2018

Miyazaki, Japan

Povezanost rada

Računarstvo, Informacijske i komunikacijske znanosti

Poveznice