Arithmetic coding for data compression
Communications of the ACM
Automatic synthesis of compression techniques for heterogeneous files
Software—Practice & Experience
Unbounded length contexts for PPM
DCC '95 Proceedings of the Conference on Data Compression
Evolutionary lossless compression with GP-ZIP*
Proceedings of the 10th annual conference on Genetic and evolutionary computation
A Field Guide to Genetic Programming
A Field Guide to Genetic Programming
A novel approach to design classifiers using genetic programming
IEEE Transactions on Evolutionary Computation
To Zip or not to Zip: effective resource usage for real-time compression
FAST'13 Proceedings of the 11th USENIX conference on File and Storage Technologies
Hi-index | 0.00 |
We use Genetic Programming (GP) to generate programs that predict the data compression ratio for compression algorithms. GP evolves programs with multiple components. One component analyses statistical features extracted from the files' byte frequency distribution to come up with a compression ratio prediction. Another component does the same but by analysing statistical features extracted from the files' raw ASCII representation. A further (evolved) component acts as a decision tree to determine the overall output (compression ratio estimation) returned by an individual. The decision tree produces its result based on a series of comparisons among statistical features extracted from the files and the outputs of the two prediction components. The evolved decision tree has the choice to select either the outputs of the two compression prediction trees or alternatively, to integrate them into an evolved mathematical formula. Experiments with the proposed approach show that GP is able to accurately estimate the compression ratio of unseen files thereby avoiding the need to run multiple compressions on a file to decide which one provide best results.