A survey on Urdu and Urdu like language stemmers and stemming techniques

作者：Abdul Jabbar, Sajid Iqbal, Muhammad Usman Ghani Khan, Shafiq Hussain

摘要

Stemming is one of the basic steps in natural language processing applications such as information retrieval, parts of speech tagging, syntactic parsing and machine translation, etc. It is a morphological process that intends to convert the inflected forms of a word into its root form. Urdu is a morphologically rich language, emerged from different languages, that includes prefix, suffix, infix, co-suffix and circumfixes in inflected and multi-gram words that need to be edited in order to convert them into their stems. This editing (insertion, deletion and substitution) makes the stemming process difficult due to language morphological richness and inclusion of words of foreign languages like Persian and Arabic. In this paper, we present a comprehensive review of different algorithms and techniques of stemming Urdu text and also considering the syntax, morphological similarity and other common features and stemming approaches used in Urdu like languages, i.e. Arabic and Persian analyzed, extract main features, merits and shortcomings of the used stemming approaches. In this paper, we also discuss stemming errors, basic difference between stemming and lemmatization and coin a metric for classification of stemming algorithms. In the final phase, we have presented the future work directions.

论文关键词：Stemming, Natural Language Processing, Information Retrieval, Urdu, Suffixes, Stemming Techniques

论文评审过程：

论文官网地址：https://doi.org/10.1007/s10462-016-9527-1