Mapping the Landscape of AI-driven Quranic Recitation Recognition: A Bibliometric and Thematic Analysis (2016–2026)
DOI:
https://doi.org/10.53840/e-jpi.v13i2.405Keywords:
Arabic speech recognition, recitation recognition, artificial intelligence, deep learning, computer-assisted language learningAbstract
This study aims to explore the development and application of artificial intelligence techniques in Arabic speech recognition, with a specific focus on recitation accuracy, pronunciation analysis, and language learning support. It seeks to identify trends, methods, and challenges in AI-based Arabic recitation recognition systems. A bibliometric and systematic review approach was employed using the Scopus database. Relevant publications from 2016 to 2026 were retrieved using a structured query combining keywords related to speech recognition, machine learning, and Arabic language processing. The selected studies were analyzed based on research trends, methodologies, and application domains. The results indicate a growing interest in deep learning approaches such as neural networks, recurrent neural networks (RNN), convolutional neural networks (CNN), and transformer-based models for Arabic speech and recitation recognition. Applications are primarily focused on pronunciation assessment, computer-assisted language learning, and Quranic recitation systems. However, challenges remain in handling dialectal variations, limited annotated datasets, and the complexity of Arabic phonetics. This study is limited to publications indexed in Scopus and may exclude relevant works from other databases. Future research should focus on improving dataset availability, incorporating tajweed rules, and enhancing model robustness for diverse Arabic dialects and recitation styles. This paper provides a comprehensive overview of AI-driven Arabic recitation recognition, highlighting current trends and gaps while offering insights for future research in educational technology and speech processing.
Keywords: Arabic speech recognition, recitation recognition, artificial intelligence, deep learning, pronunciation assessment, computer-assisted language learning, bibliometric review.
Downloads
References
Aria, M., & Cuccurullo, C. (2017). bibliometrix: An R-tool for comprehensive science mapping analysis. Journal of Informetrics, 11(4), 959–975. https://doi.org/10.1016/j.joi.2017.08.007
Baevski, A., Schneider, S., & Auli, M. (2020). VQ-Wav2Vec: Self-supervised learning of discrete speech representations. Advances in Neural Information Processing Systems, 33. https://doi.org/10.48550/arXiv.1910.05453
Chan, W., Jaitly, N., Le, Q. V., & Vinyals, O. (2016). Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://doi.org/10.1109/ICASSP.2016.7472621
Chiu, C.-C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, E., Jaitly, N., Li, B., Chorowski, J., & Bacchiani, M. (2018). State-of-the-art speech recognition with sequence-to-sequence models. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://doi.org/10.1109/ICASSP.2018.8462105
Cobo, M. J., López-Herrera, A. G., Herrera-Viedma, E., & Herrera, F. (2011). Science mapping software tools: Review, analysis, and cooperative study among tools. Journal of the American Society for Information Science and Technology, 62(7), 1382–1402. https://doi.org/10.1002/asi.21525
Donthu, N., Kumar, S., Mukherjee, D., Pandey, N., & Lim, W. M. (2021). How to conduct a bibliometric analysis: An overview and guidelines. Journal of Business Research, 133, 285–296. https://doi.org/10.1016/j.jbusres.2021.04.070
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
Graves, A., Mohamed, A., & Hinton, G. (2013). Speech recognition with deep recurrent neural networks. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6645–6649). IEEE. https://doi.org/10.48550/arXiv.1303.5778
Hinton, G., Deng, L., Yu, D., Dahl, G. E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T. N., & Kingsbury, B. (2012). Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Processing Magazine, 29(6), 82–97. https://doi.org/10.1109/MSP.2012.2205597
Huang, X., Zou, D., Cheng, G., Chen, X., & Xie, H. (2023). Trends, research issues and applications of artificial intelligence in language education. Educational Technology & Society, 26(1), 1–15. https://doi.org/10.30191/ETS.202301_26(1).0001
Knott, B., Venkataraman, S., Hannun, A., Sengupta, S., Ibrahim, M., & van der Maaten, L. (2021). CrypTen: Secure multi-party computation meets machine learning. Advances in Neural Information Processing Systems, 34. https://doi.org/10.48550/arXiv.2109.00984
Mahmod, M. A., & Zeki, A. (2019). Automated Quranic Tajweed checking rules system through recitation recognition: A review. In 4th International Conference on Islamic Applications in Computer Science and Technologies.
Miao, Y., Gowayyed, M., & Metze, F. (2016). EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding. In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). https://doi.org/10.1109/ASRU.2015.7404828
Prabhavalkar, R., Rao, K., Sainath, T. N., Li, B., Johnson, L., & Jaitly, N. (2017). A comparison of sequence-to-sequence models for speech recognition. In Proceedings of Interspeech 2017. https://doi.org/10.21437/Interspeech.2017-233
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2022). Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356. https://doi.org/10.48550/arXiv.2212.04356
Rahman, A., Kabir, M. M., Mridha, M. F., & Alatiyyah, M. (2024). Arabic speech recognition: Advancement and challenges. IEEE Access. https://doi.org/10.1109/ACCESS.2024.3376237
Rao, K., Sak, H., & Prabhavalkar, R. (2017). Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer. In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). https://doi.org/10.1109/ASRU.2017.8268945
Shaalan, K. (2011). [Review of the book Introduction to Arabic natural language processing, by N. Y. Habash]. Machine Translation, 24(3), 285–289. https://doi.org/10.1007/s10590-011-9087-8
Soliman, A. B., Eissa, K., & El-Beltagy, S. R. (2017). AraVec: A set of Arabic word embedding models for use in Arabic NLP. Procedia Computer Science, 117, 256–265. https://doi.org/10.1016/j.procs.2017.10.117
Van Eck, N. J., & Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84(2), 523–538. https://doi.org/10.1007/s11192-009-0146-3
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (pp. 5998–6008). https://doi.org/10.48550/arXiv.1706.03762
Watanabe, S., Hori, T., Kim, S., Hershey, J. R., & Hayashi, T. (2017). Hybrid CTC/attention architecture for end-to-end speech recognition. IEEE Journal of Selected Topics in Signal Processing, 11(8), 1240–1253. https://doi.org/10.1109/JSTSP.2017.2763455
Yousfi, B., & Zeki, A. M. (2016). Automatic speech recognition for the Holy Qur'an: A review. In Proceedings of the International Conference on Data Mining, Multimedia, Image Processing and their Applications (ICDMMIPA). Kuala Lumpur, Malaysia.
Zupic, I., & Čater, T. (2015). Bibliometric methods in management and organization. Organizational Research Methods, 18(3), 429–472. https://doi.org/10.1177/1094428114562629
Downloads
Published
Issue
Section
License
Copyright (c) 2026 e-Jurnal Penyelidikan dan Inovasi

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.










