Pretražite po imenu i prezimenu autora, mentora, urednika, prevoditelja

Napredna pretraga

Pregled bibliografske jedinice broj: 938196

Evaluation of Croatian Word Embeddings


Svoboda, Lukáš; Beliga, Slobodan
Evaluation of Croatian Word Embeddings // Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) / Calzolari, N. ; Choukri, K. ; Cieri, C. ; Declerck, T. ; Goggi, S. ; Hasida, K. ; Isahara, H. ; Maegaard, B. ; Mariani, J. ; Mazo, H. ; Moreno, A. ; Odijk, J. ; Piperidis, S. ; Tokunaga, T. (ur.).
Pariz: European Language Resources Association (ELRA), 2018. str. 1512-1518 (predavanje, međunarodna recenzija, cjeloviti rad (in extenso), znanstveni)


CROSBI ID: 938196 Za ispravke kontaktirajte CROSBI podršku putem web obrasca

Naslov
Evaluation of Croatian Word Embeddings

Autori
Svoboda, Lukáš ; Beliga, Slobodan

Vrsta, podvrsta i kategorija rada
Radovi u zbornicima skupova, cjeloviti rad (in extenso), znanstveni

Izvornik
Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) / Calzolari, N. ; Choukri, K. ; Cieri, C. ; Declerck, T. ; Goggi, S. ; Hasida, K. ; Isahara, H. ; Maegaard, B. ; Mariani, J. ; Mazo, H. ; Moreno, A. ; Odijk, J. ; Piperidis, S. ; Tokunaga, T. - Pariz : European Language Resources Association (ELRA), 2018, 1512-1518

ISBN
979-10-95546-00-9

Skup
11th International Conference on Language Resources and Evaluation (LREC 2018)

Mjesto i datum
Miyazaki, Japan, 07.05.2018. - 12.05.2018

Vrsta sudjelovanja
Predavanje

Vrsta recenzije
Međunarodna recenzija

Ključne riječi
Croatian word embeddings ; Croatian word analogy ; Croatian language ; Slavic language family ; Word2Vec ; FastText ; Croatian word similarity dataset ; WordSim353 ; RG65

Sažetak
Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy dataset based on the original English Word2vec word analogy dataset and added some of the specific linguistic aspects from the Croatian language. Next, we created Croatian WordSim353 and RG65 datasets for a basic evaluation of word similarities. We compared created datasets on two popular word representation models, based on Word2Vec tool and fastText tool. Models have been trained on 1.37B tokens training data corpus and tested on a new robust Croatian word analogy dataset. Results show that models are able to create meaningful word representation. This research has shown that free word order and the higher morphological complexity of Croatian language influences the quality of resulting word embeddings.

Izvorni jezik
Engleski

Znanstvena područja
Računarstvo, Informacijske i komunikacijske znanosti



POVEZANOST RADA


Ustanove:
Fakultet informatike i digitalnih tehnologija, Rijeka

Profili:

Avatar Url Slobodan Beliga (autor)

Poveznice na cjeloviti tekst rada:

Pristup cjelovitom tekstu rada www.lrec-conf.org

Citiraj ovu publikaciju:

Svoboda, Lukáš; Beliga, Slobodan
Evaluation of Croatian Word Embeddings // Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) / Calzolari, N. ; Choukri, K. ; Cieri, C. ; Declerck, T. ; Goggi, S. ; Hasida, K. ; Isahara, H. ; Maegaard, B. ; Mariani, J. ; Mazo, H. ; Moreno, A. ; Odijk, J. ; Piperidis, S. ; Tokunaga, T. (ur.).
Pariz: European Language Resources Association (ELRA), 2018. str. 1512-1518 (predavanje, međunarodna recenzija, cjeloviti rad (in extenso), znanstveni)
Svoboda, L. & Beliga, S. (2018) Evaluation of Croatian Word Embeddings. U: Calzolari, N., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Hasida, K., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., Odijk, J., Piperidis, S. & Tokunaga, T. (ur.)Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018).
@article{article, author = {Svoboda, Luk\'{a}\v{s} and Beliga, Slobodan}, year = {2018}, pages = {1512-1518}, keywords = {Croatian word embeddings, Croatian word analogy, Croatian language, Slavic language family, Word2Vec, FastText, Croatian word similarity dataset, WordSim353, RG65}, isbn = {979-10-95546-00-9}, title = {Evaluation of Croatian Word Embeddings}, keyword = {Croatian word embeddings, Croatian word analogy, Croatian language, Slavic language family, Word2Vec, FastText, Croatian word similarity dataset, WordSim353, RG65}, publisher = {European Language Resources Association (ELRA)}, publisherplace = {Miyazaki, Japan} }
@article{article, author = {Svoboda, Luk\'{a}\v{s} and Beliga, Slobodan}, year = {2018}, pages = {1512-1518}, keywords = {Croatian word embeddings, Croatian word analogy, Croatian language, Slavic language family, Word2Vec, FastText, Croatian word similarity dataset, WordSim353, RG65}, isbn = {979-10-95546-00-9}, title = {Evaluation of Croatian Word Embeddings}, keyword = {Croatian word embeddings, Croatian word analogy, Croatian language, Slavic language family, Word2Vec, FastText, Croatian word similarity dataset, WordSim353, RG65}, publisher = {European Language Resources Association (ELRA)}, publisherplace = {Miyazaki, Japan} }

Časopis indeksira:


  • Web of Science Core Collection (WoSCC)
    • Conference Proceedings Citation Index - Science (CPCI-S)
    • Conference Proceedings Citation Index - Social Sciences & Humanities (CPCI-SSH)





Contrast
Increase Font
Decrease Font
Dyslexic Font