René Haas, Leon Derczynski
Automatic language identification is a challenging problem. Discriminating between closely related languages is especially difficult. This paper presents a machine learning approach for automatic language identification for the Nordic languages, which often suffer miscategorisation by existing state-of-the-art tools. Concretely we will focus on discrimination between six Nordic languages: Danish, Swedish, Norwegian (Nynorsk), Norwegian (Bokm{\aa}l), Faroese and Icelandic.
| Task | Dataset | Metric | Value | Model |
|---|---|---|---|---|
| Language Identification | Nordic Language Identification | Accuracy | 0.9711 | FastText |