Data-driven information compression based on machine learning

V.O. Lesyk, A.Yu. Doroshenko

Abstract


The latest data compression methods and systems have come a long way, from the emergence of new artificial intelligence techniques to the widespread use of methods of various types and purposes, which has flooded the research landscape and complicated the process of selecting and identifying candidate models for effective application, further improvement, and research. This article proposes an empirical approach to comparing trained compression methods and models within specific frameworks, highlighting potential application areas, the best models at the current stage of development, and the characteristics, advantages, and disadvantages of current models. An assessment of the general applicability of trained compression and the theoretical limits of entropy is conducted. A series of comparisons is performed, taking into account various subject areas of model application; a set of criteria for comparing and evaluating the implementation effectiveness of each of the considered methods is described; key problem areas that need to be assessed prior to implementing the models into technological processes and as well as their use from the perspective of ordinary users of software applications or data compression packages. An additional assessment of potential technologies for the further development and improvement of learned compression methods is provided, along with an identification of systems and technologies that are losing their potential for progress. This analysis allows researchers to gain a general understanding of the current state of research on the application of machine learning to various compression problems, assess the state of technological development, and—based on the limitations and prob lem areas identified—select the appropriate direction for further research or for scientific literature discovery.

Problems in programming 2026; 3: 102-110


Keywords


machine learning; learned data compression; data compression with neural networks; hybrid com pression methods; artificial intelligence

References


LI, Ming, et al. Understanding is compression. 2024.WELCH, Terry A. A technique for high performance data compression. Computer, 1984, 17.06: 8-19.

WILKENFELD, Daniel A. Understanding as compression: DA Wilkenfeld.Philosophical Studies, 2019, 176.10: 2807-2831

YANG, Yibo; MANDT, Stephan; THEIS, Lucas. An introduction to neural data compres sion. Foundations and Trends in Computer Graphics and Vision, 2023, 15.2: 113-200.

Huang, C.-H. and Wu, J.-L. (2024). Unveiling the Future of Human and Machine Coding: A Survey of End-to-End Learned Image Compression.Entropy, 26.

Hu, Y., Yang, W., Z. and Liu, J. (2020). Learning End-to-End Lossy Image Compression: A Benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44, pp. 4194-4211.

Cao, M., Dai, W., Li, S., Li, C., Zou, J., Chen, Y. and Xiong, H. (2025). End-to-End Optimized Image Compression With Deep Gaussian Process Regression. IEEE Transactions on Circuits and Systems for Video Technology, 35, pp. 3770-3785.

Ehrlich, M. (2022). The First Principles of Deep Learning and Compression.ArXiv, abs/2204.01782.

Porat, D. and Levi, D. (2026). Learned Image and Video Compression: Foundations, Algorithms, and Open Challenges. Proceedings of the 5th Mile-High Video Conference.

Del'etang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau Moya, J., Li, W., Aitchison, M., Orseau, L., Hutter, M. and Veness, J. (2023). Language Modeling Is Compression. ArXiv, abs/2309.10668.

Agustsson, E. and Theis, L. (2020). Universally Quantized Neural Compression. ArXiv, abs/2006.09952.

Blau, Y. and Michaeli, T. (2019). Rethinking Lossy Compression: The Rate-Distortion Perception Tradeoff., pp. 675-685.

Ballé, J., Laparra, V. and Simoncelli, E. P. (2016). End-to-end Optimized Image Compression. ArXiv, abs/1611.01704.

Wang, C., Liang, Z., Niu, K. and Zhang, P. (2026). A Theoretical Framework for Rate Distortion Limits in Learned Image Compression. ArXiv, abs/2601.09254.

Sheng, X., Li, L., Liu, D. and Li, H. (2024). Spatial Decomposition and Temporal Fusion Based Inter Prediction for Learned Video Compression. IEEE Transactions on Circuits and Systems for Video Technology, 34, pp. 6460-6473.

Lesyk, V.O. and Doroshenko, A.Yu. (2023) 'Image compression module based neural network autoencoders', Problems in Programming, 1, pp. 48-57. [in Ukrainian]

Chen, T., Z., Shen, Q., Cao, X. and Wang, Y. (2019). End-to-End Learnt Image Compression via Non-Local Attention Optimization and Improved Context Modeling. IEEE Transactions on Image Processing, 30, pp. 3179-3191.

Jiang, W., Yang, J., Zhai, Y., Gao, F. and Wang, R. (2023). MLIC++: Linear Complexity Multi-Reference Entropy Modeling for Learned Image Compression. ACM Transactions on Multimedia Computing, Communications and Applications, 21, pp. 1 - 25.

Sun, H., H., Ling, F., Xie, H., Sun, Y., Yi, L., Yan, M., Zhong, C., Liu, X. and Wang, G. (2025). A survey and benchmark evaluation for neural network-based lossless universal compressors toward multi-source data. Frontiers of Computer Science, 19.

Hishida, K., Liu, C., Paparrizos, J. and Elmore, A. (2025). Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point Compression. Proc. VLDB Endow., 18, pp. 4396-4409.

Li, Y., Guo, Y., Guerin, F. and Lin, C. (2024). Evaluating Large Language Models for Generalization and Robustness via Data Compression. ArXiv, abs/2402.00861.

Zhang, J., Cheng, Z., Zhao, Y., Wang, S., Zhou, D., Lu, G. and Song, L. (2024). L3TC: Leveraging RWKV for Learned Lossless Low Complexity Text Compression., pp. 13251-13259.

Bellard, F. (2021). NNCP v2: Lossless Data Compression with Transformer.

Mao, Y., Li, J., Cui, Y. and Xue, J. (2023). Faster and Stronger Lossless Compression with Optimized Autoregressive Framework. 2023 60th ACM/IEEE Design Automation Conference (DAC), pp. 1-6.

Mao, Y., Cui, Y., Kuo, T.-W. and Xue, C. (2022). TRACE: A Fast Transformer-based General-Purpose Lossless Compressor. Proceedings of the ACM Web Conference 2022.

Goyal, M., Tatwawadi, K., Chandak, S. and Ochoa, I. (2020). DZip: Improved General Purpose Lossless Compression Based on Novel Neural Network Modeling. 2020 Data Compression Conference (DCC), pp. 372-372.

Li, Z., Huang, C., Wang, X., Hu, H., Wyeth, C., Bu, D., Yu, Q., Gao, W., Liu, X. and Li, M. (2024). Lossless data compression by large models. Nature Machine Intelligence, 7, pp. 794 - 799.

Tacconelli, R. (2026). Nacrith: Neural Lossless Compression via Ensemble Context Modeling and High-Precision Coding. ArXiv, CDF abs/2602.19626.

Jamil, S., Piran, M. and Rahman, M. (2022). Learning-Driven Lossy Image Compression; A Comprehensive Survey. ArXiv, abs/2201.09240.

Di, S., Liu, J., Zhao, K., Liang, X., Underwood, R., Zhang, Z., Shah, M., Huang, Y., Huang, J., Yu, X., Ren, C., Guo, H., Wilkins, G., Tao, D., Tian, J., Jin, S., Jian, Z., Wang, D., Rahman, M. H., Zhang, B., Calhoun, J. C., Li, G., Yoshii, K., Alharthi, K. and Cappello, F. (2024). A Survey on Error Bounded Lossy Compression for Scientific Datasets. ACM Computing Surveys, 57, pp. 1 - 38.

Wondmagegn, A. B., Dinh, T. T. T., Win, T. T., Won, D. and Cho, S.(2026). A Survey on Deep Learning-Based Image Compression: Methods, Efficiency, and Open Challenges. 2026 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), pp. 108-111.

Pinheiro, A. M. G. (2025). JPEG Column: 106th JPEG Meeting. ACM SIGMultimedia Records, 17, pp. 1 - 8.


Refbacks

  • There are currently no refbacks.