Conjugation-based compression for Hebrew texts

Authors:
Yair Wiseman;Irit Gefner
Affiliations:
The Hebrew University of Jerusalem, Jerusalem, Israel;The Hebrew University of Jerusalem, Jerusalem, Israel
Venue:
ACM Transactions on Asian Language Information Processing (TALIP)
Year:
2007

Citing 6
Cited 0

HUHU: the Hebrew University Hebrew understander

Computer Languages
Arithmetic coding for data compression

Communications of the ACM
KEDMA—Linguistic Tools for Retrieval Systems

Journal of the ACM (JACM)
Full text document retrieval: Hebrew legal texts (report on the first phase of the responsa retrieval project)

SIGIR '71 Proceedings of the 1971 international ACM SIGIR conference on Information storage and retrieval
Can We Do without Ranks in Burrows Wheeler Transform Compression?

DCC '01 Proceedings of the Data Compression Conference
Hebrew Computational Linguistics: Past and Future

Artificial Intelligence Review

Quantified Score

Hi-index	0.00

Visualization

Abstract

Traditional compression techniques do not look deeply into the morphology of languages. This can be less critical in languages like English where most of the sequences are illegal according to the grammatical rules of the language, for example, zx, bv or qe; hence the morphology can add a little information that can be beneficial for the compression algorithm. However, this negligence can be a significant flaw in languages like Hebrew where the grammatical rules allow much more freedom in the sequences of letters and, except tet after gimel, any pair is legal; hence compressing without taking the morphological rules into account can yield a poorer compression ratio. This article suggests a tool that optimizes the Burrows-Wheeler algorithm which is an unaware morphological rules compression method. It first preprocesses a Hebrew text file according to the Hebrew conjugation rules, and, after that, it provides the Burrows-Wheeler algorithm with this preprocessed file so that can be compressed better. Experimental results show a significant improvement.