Mathematical models of software components for infocommunication service providers
Abstract
The automation of solving intellectual tasks—specifically, the design of software components for information and communication service providers—using artificial intelligence is an important and timely issue in service lifecycle management platforms for IT development companies based on the End-to-End model. This paper addresses the problem of synthesizing service component code using deep learning models. The requirements for service components are analyzed, and possible service component codes are examined. A mathematical model has been developed to describe the requirements for service software components. The paper analyzes existing mathematical models for representing program code in the Java language that can be used for the mathematical description of service component code. A mathematical model of the program code of service components has been developed, which differs from existing ones in control and adaptation to the features of the grammar of a specific language. This allows for the continued use of unimodal language models. The proposed model is a component of the platform for supporting the life cycle of services in information systems of providers. Within an integrated system of tools for automating service lifecycle processes, the proposed mathematical models will enable the training and implementation of deep neural networks for designing the program code of service components.
Problems in programming 2026; 3: 53-68
Keywords
Full Text:
PDF (Українська)References
Zimmermann,T., Zeller,A., Weissgerber,P. and Diehl,S., (2005). Mining version histories to guide software changes. IEEE Transactions on Software Engineering. 31(6), 429–445.
Le, T.H. M., Chen, H. and Babar, M.A., (2020). Deep Learning for Source Code Modeling and Generation: Models, Applications, and Challenges. ACM Comput. Surv. 53(3), article no: 62.
Raychev, V., Vechev, M. and Yahav, E., (2014). Code completion with statistical language models. In: Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation. New York, NY, USA: Association for Computing Machinery. pp.419–428.
Vinyals, O., Fortunato, M. & Jaitly, N., (2017). Pointer-Networks. [Preprint].
Li,J., Wang,Y., Lyu, M.R. & King, I., (2018). Code Completion with Neural Attention and Pointer Networks. In: Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, 13–19July 2018, Stockholm, Sweden. International Joint Conferences on Artificial Intelligence Organization. pp.4159–4165.
Balog, M., Gaunt, A.L., Brockschmidt,M., Nowozin,S. and Tarlow,D., (2017). DeepCoder: Learning to Write Programs. [Preprint].
Wang, Z., Chu, Z., Doan, T.V., Ni,S., Yang, M. and Zhang, W., (2024). History, Development and Principles of Large Language Models-An Introductory Survey. [Preprint].
Guo,D., Ren,S., Lu,S., Feng,Z., Tang,D., Liu,S., Zhou,L., Duan,N., Svyatkovskiy,A., Fu,S., Tufano,M., Deng,S.K., Clement,C., Drain,D., Sundaresan,N., Yin,J., Jiang,D. & Zhou,M., (2021). GraphCodeBERT: Pretraining Code Representations with Data Flow. [Preprint].
Jiang,X., Zheng,Z., Lyu,C., Li,L. & Lyu,L., (2021). TreeBERT: A tree-based pretrained model for programming language. In: C.de Campos and M.H.Maathuis, eds. Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence. PMLR. pp.54–63.
Wang, X., Wang, Y., Mi, F., Zhou, P., Wan, Y., Liu, X., Li, L., Wu, H., Liu, J. & Jiang, X., (2021). SynCoBERT: Syntax Guided Multi-Modal Contrastive Pre Training for Code Representation. [Preprint].
Ahmad, W., Chakraborty, S., Ray, B. & Chang, K.-W., (2021). Unified Pre-training for Program Understanding and Generation. In: .Toutanova, A.Rumshisky, L.Zettlemoyer, D.Hakkani-Tur, I.Beltagy, S.Bethard, R.Cotterell, T.Chakraborty and Y.Zhou, eds. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics. pp. 2655–2668.
Wang, Y., Wang, W., Joty ,S. & Hoi, S.C. H., (2021). CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. [Preprint].
Niu, C., Li, C., Ng, V., Ge, J., Huang, L. & Luo, B., (2022). SPT-code: sequence-to sequence pre-training for learning source code representations. In: Proceedings of the 44th International Conference on Software Engineering, 21–29May 2022, Pittsburgh, Pennsylvania. New York, NY, United States: Association for Computing Machinery. pp.2006–2018.
Guo, D., Lu, S., Duan, N., Wang, Y., Zhou, M. & Yin, J., (2022). UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In: S.Muresan, P.Nakov and A.Villavicencio, eds. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 22–27May 2022, Dublin, Ireland. Kerrville, TX, USA: Association for Computational Linguistics. pp.7212–7225.
Chakraborty, S., Ahmed, T., Ding, Y., Devanbu, P.T. & Ray, B., (2022). NatGen: generative pre-training by "naturalizing" source code. In: Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. New York, NY, USA: Association for Computing Machinery. pp.18–30.
Tipirneni, S., Zhu, M. & Reddy, C.K., (2024). StructCoder: Structure-Aware Transformer for Code Generation. ACM Trans. Knowl. Discov. Data. 18(3), 1–20.
Gong,L., Elhoushi,M. and Cheung,A., (2024). AST-T5: Structure-Aware Pretraining for Code Generation and Understanding. In: R.Salakhutdinov, Z.Kolter, K.Heller, A.Weller, N.Oliver, J.Scarlett & F. Berkenkamp, eds. Proceedings of the 41st International Conference on Machine Learning, PMLR 235, 21–27July 2024, Vienna, Austria. MLResearchPress. pp.15839–15853.
Bartkowiak,P. and Graliński,F., (2025). Seamlessly Integrating Tree-Based Positional Embeddings into Transformer Models for Source Code Representation. In: H. Fei, K. Tu, Y. Zhang, X. Hu, W. Han, Z. Jia, Z. Zheng, Y. Cao, M. Zhang, W. Lu, N. Siddharth, N.Xue and Y.Zhang, eds. Proceedings of the 1st Joint Workshop on Large Language Models and Structure Modeling (XLLM 2025), Vienna, Austria. Kerrville, TX, USA: Association for Computational Linguistics. pp.91–98.
Ren, S., Guo, D., Lu, S., Zhou, L., Liu, S., Tang, D., Sundaresan, N., Zhou, M., Blanco, A. and Ma, S., (2020). CodeBLEU: a Method for Automatic Evaluation of Code Synthesis. [Preprint].
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J. and Tian, Y., (2025). Training Large Language Models to Reason in a Continuous Latent Space. [Preprint].
Brunsfeld, M. et al., (2026). tree-sitter/tree sitter [online]. Version 0.26.9. [computer software].
Tree-sitter Grammars, (2025). tree-sitter properties.Version 0.3.0. [computer software].
PlantUML, (2026). PlantUML [online]. Version 1.2026.2. [computer software].
mermaid-js, (2026). mermaid. Version 11.14.0. [computer software].
Spring Boot, (2026). Version 4.0.5. [computer software].
Tree-sitter Grammars, (2026). tree-sitter markdown. Version 0.5.3. [computer software].
Refbacks
- There are currently no refbacks.








