Afaan Oromo Text Retrieval System

Gezehagn, Gutema Eggi (2012) Afaan Oromo Text Retrieval System. Masters thesis, Addis Ababa University.

[img] PDF (Afaan Oromo Text Retrieval System)
Gezehagn, Gutema Eggi.pdf - Accepted Version
Restricted to Repository staff only

Download (2MB) | Request a copy

Abstract

This study is mainly intended to make possible retrieval of Afan Oromo text documents by applying techniques of modern information retrieval system. Information retrieval is a mechanism that enables finding relevant information material of unstructured nature that satisfies information need of user from large collection. Afaan Oromo text retrieval developed in this study has indexing and searching parts. Vector Space Model of information retrieval system was used to guide searching for relevant document from Oromiffa text corpus. The model is selected since Vector space model is the widely used classic model of information retrieval system. The index file structure used is inverted index file structure. For this study text document corpus is prepared by the researcher encompassing different news article and experiment is made by using 9(nine) different user information need queries. Various techniques of text pre-processing including tokenization, normalization, stop word removal and stemming are used for both document indexing and query text. The experiment shows that the performance is on the average 0.575(57.5%) precision and 0.6264(62.64%) recall. The challenging tasks in the study are handling synonymy and polysemy, inability of the stemmer algorithm to all word variants, and ambiguity of words in the language. The performance the system can be increased if stemming algorithm is improved, standard test corpus is used, and thesaurus is used to handle polysemy and synonymy words in the language.

Item Type: Thesis (Masters)
Subjects: P Language and Literature > P Philology. Linguistics
P Language and Literature > PL Languages and literatures of Eastern Asia, Africa, Oceania
Q Science > QA Mathematics > QA76 Computer software
Divisions: Africana
Depositing User: Selom Ghislain
Date Deposited: 26 Jun 2018 09:02
Last Modified: 26 Jun 2018 09:02
URI: http://thesisbank.jhia.ac.ke/id/eprint/4567

Actions (login required)

View Item View Item