Search by keyword or author

Optimization-Driven Deep Learning Models for 3D Human Action Recognition: A Survey

Published: May 31, 2026 | Views: 107

Authors

Kamalpreet Kaur, Ankit Bansal, and Baljit Singh Khehra

Keywords
Human action recognition, 3D skeletal data, Deep learning, Graph convolutional networks, Metaheuristic optimization, Spatiotemporal modeling, Hybrid neural networks, Transformer-based models

Abstract

Background: Due to its numerous applications in healthcare, surveillance, human-computer interaction, sports analytics, and smart environments, Human Action Recognition (HAR) using 3D skeletal data has grown in importance as a field of study. Conventional HAR techniques mainly relied on manually created features, which were unreliable and data dependent. The ability to model intricate spatiotemporal patterns in human motion has greatly improved with the development of deep learning models like CNNs, RNNs, GCNs, and hybrid architectures. The integration of optimization techniques into deep learning-based HAR systems is motivated by unresolved issues like high computational cost, sensitivity to hyperparameters, poor cross-dataset generalization, and lack of interpretability.

Purpose: This survey’s goal is to offer a thorough analysis of optimization-driven deep learning models for 3D human action recognition. The research seeks to examine cutting-edge deep learning architectures for 3D HAR and analyze how metaheuristic optimization techniques can enhance the precision, effectiveness, and generalization of models.

Methods: Using a methodical survey approach, this study covers deep learning architectures such as Transformer-based, CNN-based, RNN-based, GCN-based, and Hybrid-DNN models. Optimization techniques include Chaos Game Optimization, Whale Optimization Algorithm, Grey Wolf Optimizer, Rao-3, Wild Horse Optimization, and Sea Horse Optimization. Model comparison is conducted using benchmark datasets such as NTU RGB+D, NTU RGB+D 120, Kinetics-Skeleton, SYSU-3D, and N-UCLA. Assessment uses standard performance metrics, mainly accuracy, along with reported robustness and computational efficiency.

Results: According to the survey, optimization-driven deep learning models routinely perform better in 3D HAR than conventional and non-optimized methods. Higher recognition accuracy is attained by optimized models, frequently surpassing 90–95% on benchmark datasets. Hyperparameter tuning is greatly enhanced by metaheuristic optimization, which lowers overfitting and computational inefficiency. Despite advancements, problems like cross-dataset generalization, model explainability, and real-time processing still exist.

Conclusion: Recent developments in optimization-driven deep learning models for 3D Human Action Recognition (HAR) were examined in this survey. The analysis demonstrates that combining metaheuristic optimization techniques with deep learning architectures greatly increases recognition accuracy, generalization, and computational efficiency.

References

  • Ahmad, A. Y. A. B., Alzubi, J., James, S., Nyangaresi, V. O., Kutralakani, C., & Krishnan, A. (2024). Enhancing human action recognition with adaptive hybrid deep attentive networks and Archerfish optimization. Computers, Materials & Continua, 80(3). https://doi.org/10.32604/cmc.2024.052771
  • Ahmad, T., Jin, L., Zhang, X., Lai, S., Tang, G., & Lin, L. (2021). Graph convolutional neural network for human action recognition: A comprehensive survey. IEEE Transactions on Artificial Intelligence, 2(2), 128–145. https://doi.org/10.1109/TAI.2021.3076974
  • Al-Wesabi, F. N., Albraikan, A. A., Hilal, A. M., Al-Shargabi, A. A., Alhazbi, S., Al Duhayyim, M., et al. (2021). Design of optimal deep learning based human activity recognition on sensor enabled internet of things environment. IEEE Access, 9, 143988–143996. https://doi.org/10.1109/ACCESS.2021.3112973
  • Amirthalingam, P., Alatawi, Y., Chellamani, N., Shanmuganathan, M., Ali, M. A. S., Alqifari, S. F., et al. (2024). Sea horse optimization–deep neural network: A medication adherence monitoring system based on hand gesture recognition. Sensors, 24(16), 5224. https://doi.org/10.3390/s24165224
  • Anbazhagan, K., Swamy, G., Janani, R., & Farakte, A. (2024). Deep learning based human activity recognition in smart home. In 2024 4th International Conference on Data Engineering and Communication Systems (ICDECS) (pp. 1–6). IEEE. https://doi.org/10.1109/ICDECS59733.2023.10502527
  • Arshad, M. H., Bilal, M., & Gani, A. (2022). Human activity recognition: Review, taxonomy and open challenges. Sensors, 22(17), 6463. https://doi.org/10.3390/s22176463
  • Basak, H., Kundu, R., Singh, P. K., Ijaz, M. F., Woźniak, M., & Sarkar, R. (2022). A union of deep learning and swarm-based optimization for 3D human action recognition. Scientific Reports, 12(1), 5494. https://doi.org/10.1038/s41598-022-09293-8
  • Bavil, A. F., Damirchi, H., & Taghirad, H. D. (2023). Action capsules: Human skeleton action recognition. Computer Vision and Image Understanding, 233, 103722. https://doi.org/10.1016/j.cviu.2023.103722
  • Beddiar, D. R., Nini, B., Sabokrou, M., & Hadid, A. (2020). Vision-based human activity recognition: A survey. Multimedia Tools and Applications, 79(41), 30509–30555. https://doi.org/10.1007/s11042-020-09004-3
  • Bian, S., Liu, M., Zhou, B., Lukowicz, P., & Magno, M. (2024). Body-area capacitive or electric field sensing for human activity recognition and human-computer interaction: A comprehensive survey. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 8(1), 1–49. https://doi.org/10.1145/3643555
  • Bulling, A., Blanke, U., & Schiele, B. (2014). A tutorial on human activity recognition using body-worn inertial sensors. ACM Computing Surveys, 46(3), 1–33. https://doi.org/10.1145/2499621
  • Bulugu, I. (2024). Adaptive shift graph convolutional neural network for hand gesture recognition based on 3D skeletal similarity. Signal, Image and Video Processing, 18(11), 7583–7595. https://doi.org/10.1007/s11760-024-03412-w
  • Caetano, C., Brémond, F., & Schwartz, W. R. (2019). Skeleton image representation for 3D action recognition based on tree structure and reference joints. In 2019 32nd SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI) (pp. 16–23). IEEE. https://doi.org/10.1109/SIBGRAPI.2019.00011
  • Cao, Z., Simon, T., Wei, S.-E., & Sheikh, Y. (2017). Realtime multi-person 2D pose estimation using part affinity fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 7291–7299). https://doi.org/10.48550/arXiv.1611.08050
  • Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In European Conference on Computer Vision (pp. 213–229). Springer. https://doi.org/10.1007/978-3-030-58452-8_13
  • Challa, S. K., Kumar, A., Semwal, V. B., & Dua, N. (2023). An optimized deep learning model for human activity recognition using inertial measurement units. Expert Systems, 40(10), e13457. https://doi.org/10.1111/exsy.13457
  • Chaquet, J. M., Carmona, E. J., & Fernández-Caballero, A. (2013). A survey of video datasets for human action and activity recognition. Computer Vision and Image Understanding, 117(6), 633–659. https://doi.org/10.1016/j.cviu.2013.01.013
  • Chen, D., Chen, M., Wu, P., Wu, M., & Zhang, T. (2025). Two-stream spatio-temporal GCN-transformer networks for skeleton-based action recognition. Scientific Reports, 15(1), 87752. https://doi.org/10.1038/s41598-025-87752-8
  • Chen, J., Yang, W., Liu, C., & Yao, L. (2021). A data augmentation method for skeleton-based action recognition with relative features. Applied Sciences, 11(23), 11481. https://doi.org/10.3390/app112311481
  • Chen, L., Wei, H., & Ferryman, J. (2013). A survey of human motion analysis using depth imagery. Pattern Recognition Letters, 34(15), 1995–2006. https://doi.org/10.1016/j.patrec.2013.02.006
  • Chen, Y., Zhang, Z., Yuan, C., Li, B., Deng, Y., & Hu, W. (2021). Channel-wise topology refinement graph convolution for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 13359–13368). https://doi.org/10.48550/arXiv.2107.12213
  • Chen, Z., Zhang, L., Jiang, C., Cao, Z., & Cui, W. (2019). WiFi CSI based passive human activity recognition using attention based BLSTM. IEEE Transactions on Mobile Computing, 18(11), 2714–2724. https://doi.org/10.1109/TMC.2018.2878233
  • Cherla, S., Kulkarni, K., Kale, A., & Ramasubramanian, V. (2008). Towards fast, view-invariant human action recognition. In 2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops (pp. 1–8). IEEE. https://doi.org/10.1109/CVPRW.2008.4563179
  • Cho, S., Maqbool, M., Liu, F., & Foroosh, H. (2020). Self-attention network for skeleton-based human action recognition. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 635–644). https://doi.org/10.48550/arXiv.1912.08435
  • Coronado, E., Itadera, S., & Ramirez-Alpizar, I. G. (2023). Integrating virtual, mixed, and augmented reality to human–robot interaction applications using game engines: A brief review of accessible software tools and frameworks. Applied Sciences, 13(3), 1292. https://doi.org/10.3390/app13031292
  • D’Sa, A. G., & Prasad, B. G. (2019). A survey on vision based activity recognition, its applications and challenges. In 2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP) (pp. 1–8). IEEE. https://doi.org/10.1109/ICACCP.2019.8882896
  • Daga, Y., & Meena, S. (2022). Applications of human activity recognition in different fields: A review. In 2022 IEEE 9th Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON) (pp. 1–6). IEEE. https://doi.org/10.1109/UPCON56432.2022.9986408
  • Dhiman, C., & Vishwakarma, D. K. (2019). A review of state-of-the-art techniques for abnormal human activity recognition. Engineering Applications of Artificial Intelligence, 77, 21–45. https://doi.org/10.1016/j.engappai.2018.08.014
  • Ding, X., Yang, K., & Chen, W. (2019). An attention-enhanced recurrent graph convolutional network for skeleton-based action recognition. In Proceedings of the 2019 2nd International Conference on Signal Processing and Machine Learning (pp. 79–84). https://doi.org/10.1145/3372806.3372814
  • Ding, X., Yang, K., & Chen, W. (2020). A semantics-guided graph convolutional network for skeleton-based action recognition. In Proceedings of the 2020 4th International Conference on Innovation in Artificial Intelligence (pp. 130–136). https://doi.org/10.1145/3390557.3394129
  • Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al. (2020). An image is worth 16×16 words: Transformers for image recognition at scale. arXiv. https://doi.org/10.48550/arXiv.2010.11929
  • Duan, H., Zhao, Y., Chen, K., Lin, D., & Dai, B. (2022). Revisiting skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2969–2978). https://github.com/kennymckormick/pyskl
  • Edel, M., & Köppe, E. (2016). Binarized-BLSTM-RNN based human activity recognition. In 2016 International Conference on Indoor Positioning and Indoor Navigation (IPIN) (pp. 1–7). IEEE. https://doi.org/10.1109/IPIN.2016.7743581
  • El-Assal, M., Tirilly, P., & Bilasco, I. M. (2024). S3TC: Spiking separated spatial and temporal convolutions with unsupervised STDP-based learning for action recognition. In International Conference on Pattern Recognition (pp. 299–314). Springer. https://doi.org/10.1007/978-3-031-78395-1_20
  • Gao, J., He, T., Zhou, X., & Ge, S. (2019). Focusing and diffusion: Bidirectional attentive graph convolutional networks for skeleton-based action recognition. arXiv. https://doi.org/10.48550/arXiv.1912.11521
  • Gao, X., Jin, Y., Dou, Q., Fu, C.-W., & Heng, P.-A. (2021). Accurate grid keypoint learning for efficient video prediction. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 5908–5915). IEEE. https://doi.org/10.1109/IROS51168.2021.9636874
  • Gaur, D., & Kumar Dubey, S. (2022). Development of activity recognition model using LSTM-RNN deep learning algorithm. Journal of Information and Organizational Sciences, 46(2), 277–291. https://doi.org/10.31341/jios.46.2.1
  • Gavrilyuk, K., Sanford, R., Javan, M., & Snoek, C. G. M. (2020). Actor-transformers for group activity recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 839–848). https://doi.org/10.48550/arXiv.2003.12737
  • Ghadiyaram, D., Tran, D., & Mahajan, D. (2019). Large-scale weakly-supervised pre-training for video action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12046–12055). https://doi.org/10.1109/CVPR.2019.01232
  • Gilroy, S., Glavin, M., Jones, E., & Mullins, D. (2021). Pedestrian occlusion level classification using keypoint detection and 2D body surface area estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (pp. 3833–3839). https://doi.org/10.1109/ICCVW54120.2021.00427
  • Guha, R., Khan, A. H., Singh, P. K., Sarkar, R., & Bhattacharjee, D. (2021). CGA: A new feature selection model for visual human action recognition. Neural Computing and Applications, 33, 5267–5286. https://doi.org/10.1007/s00521-020-05297-5
  • Gupta, S. (2021). Deep learning based human activity recognition (HAR) using wearable sensor data. International Journal of Information Management Data Insights, 1(2), 100046. https://doi.org/10.1016/j.jjimei.2021.100046
  • Ha, S., Yun, J.-M., & Choi, S. (2015). Multi-modal convolutional neural networks for activity recognition. In 2015 IEEE International Conference on Systems, Man, and Cybernetics (pp. 3017–3022). IEEE. https://doi.org/10.1109/SMC.2015.525
  • Hameed, S., Qolomany, B., Belhaouari, S. B., Abdallah, M., Qadir, J., & Al-Fuqaha, A. (2025). Large language model enhanced particle swarm optimization for hyperparameter tuning for deep learning models. IEEE Open Journal of the Computer Society, 6, 574–585. https://doi.org/10.1109/OJCS.2025.3564493
  • He, J., Zhang, C., He, X., & Dong, R. (2020). Visual recognition of traffic police gestures with convolutional pose machine and handcrafted features. Neurocomputing, 390, 248–259. https://doi.org/10.1016/j.neucom.2019.07.103
  • Hoang, V.-H., Lee, J. W., Piran, M. J., & Park, C.-S. (2023). Advances in skeleton-based fall detection in RGB videos: From handcrafted to deep learning approaches. IEEE Access, 11, 92322–92352. https://doi.org/10.1109/ACCESS.2023.3307138
  • Hossain, H. M. S., Al Haiz Khan, M. D. A., & Roy, N. (2018). DeActive: Scaling activity recognition with active deep learning. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2(2), 1–23. https://doi.org/10.1145/3214269
  • Hu, J.-F., Zheng, W.-S., Lai, J., & Zhang, J. (2015). Jointly learning heterogeneous features for RGB-D activity recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 5344–5352). https://doi.org/10.1109/CVPR.2015.7299172
  • Huang, X., Mei, G., & Zhang, J. (2023). Cross-source point cloud registration: Challenges, progress and prospects. Neurocomputing, 548, 126383. https://doi.org/10.1016/j.neucom.2023.126383
  • Huang, X., Zhou, H., Wang, J., Feng, H., Han, J., Ding, E., et al. (2023). Graph contrastive learning for skeleton-based action recognition. arXiv. https://doi.org/10.48550/arXiv.2301.10900
  • Huynh-The, T., Hua, C.-H., & Kim, D.-S. (2019). Encoding pose features to images with data augmentation for 3-D action recognition. IEEE Transactions on Industrial Informatics, 16(5), 3100–3111. https://doi.org/10.1109/TII.2019.2910876
  • Ibh, M., Grasshof, S., Witzner, D., & Madeleine, P. (2023). TemPose: A new skeleton-based transformer model designed for fine-grained motion recognition in badminton. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (pp. 5199–5208). https://doi.org/10.1109/CVPRW59228.2023.00548
  • Inoue, M., Inoue, S., & Nishida, T. (2018). Deep recurrent neural network for mobile human activity recognition with high throughput. Artificial Life and Robotics, 23, 173–185. https://doi.org/10.1007/s10015-017-0422-x
  • Islam, M. M., Nooruddin, S., Karray, F., & Muhammad, G. (2022). Human activity recognition using tools of convolutional neural networks: A state-of-the-art review, data sets, challenges, and future prospects. Computers in Biology and Medicine, 149, 106060. https://doi.org/10.1016/j.compbiomed.2022.106060
  • Jain, P., Kulis, B., & Grauman, K. (2008). Fast image search for learned metrics. In 2008 IEEE Conference on Computer Vision and Pattern Recognition (pp. 1–8). IEEE. https://doi.org/10.1109/CVPR.2008.4587841
  • Jaiswal, S., & Jaiswal, T. (2020). Remarkable skeleton based human action recognition. Artificial Intelligence Evolution, 108–121. https://doi.org/10.37256/aie.122020562
  • Jegham, I., Khalifa, A. B., Alouani, I., & Mahjoub, M. A. (2020). Vision-based human action recognition: An overview and real world challenges. Forensic Science International: Digital Investigation, 32, 200901. https://doi.org/10.1016/j.fsidi.2019.200901
  • Jhuang, H., Gall, J., Zuffi, S., Schmid, C., & Black, M. J. (2013). Towards understanding action recognition. In Proceedings of the IEEE International Conference on Computer Vision (pp. 3192–3199). https://doi.org/10.1109/ICCV.2013.396
  • Kahatapitiya, K., & Ryoo, M. S. (2021). Coarse-fine networks for temporal activity detection in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8385–8394). https://doi.org/10.1109/CVPR46437.2021.00828
  • Karthickkumar, S., & Kumar, K. (2020). A survey on deep learning techniques for human action recognition. In 2020 International Conference on Computer Communication and Informatics (ICCCI) (pp. 1–6). IEEE. https://doi.org/10.1109/ICCCI48352.2020.9104135
  • Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., et al. (2017). The Kinetics human action video dataset. arXiv. https://doi.org/10.48550/arXiv.1705.06950
  • Khan, M. A., Javed, K., Khan, S. A., Saba, T., Habib, U., Khan, J. A., et al. (2024). Human action recognition using fusion of multiview and deep features: An application to video surveillance. Multimedia Tools and Applications, 83(5), 14885–14911. https://doi.org/10.1007/s11042-020-08806-9
  • Khan, S., Naseer, M., Hayat, M., Zamir, S. W., Khan, F. S., & Shah, M. (2022). Transformers in vision: A survey. ACM Computing Surveys, 54(10s), 1–41. https://doi.org/10.1145/3505244
  • Kishore, P. V. V., Kumar, D. A., Sastry, A. S. C. S., & Kumar, E. K. (2018). Motionlets matching with adaptive kernels for 3-D Indian sign language recognition. IEEE Sensors Journal, 18(8), 3327–3337. https://doi.org/10.1109/JSEN.2018.2810449
  • Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., & Serre, T. (2011). HMDB: A large video database for human motion recognition. In 2011 International Conference on Computer Vision (pp. 2556–2563). IEEE. https://doi.org/10.1109/ICCV.2011.6126543
  • Kumar, M. T. K., Kishore, P. V. V., Madhav, B. T. P., Kumar, D. A., Kala, N. S., Rao, K. P. K., et al. (2021). Can skeletal joint positional ordering influence action recognition on spectrally graded CNNs: A perspective on achieving joint order independent learning. IEEE Access, 9, 139611–139626. https://doi.org/10.1109/ACCESS.2021.3119455
  • Lale, T., Yüksek, G., & Çınar, R. F. (2026). A novel hybrid metaheuristic for optimizing deep CNN hyperparameters to enhance heart disease prediction on a comprehensive merged dataset. Cluster Computing, 29(2), 117. https://doi.org/10.1007/s10586-025-05892-y
  • Lara, O. D., & Labrador, M. A. (2013). A survey on human activity recognition using wearable sensors. IEEE Communications Surveys & Tutorials, 15(3), 1192–1209. https://doi.org/10.1109/SURV.2012.110112.00192
  • Li, M., & Sun, Q. (2021). 3D skeletal human action recognition using a CNN fusion model. Mathematical Problems in Engineering, 2021(1), 6650632. https://doi.org/10.1155/2021/6650632
  • Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., & Tian, Q. (2019). Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3595–3603). https://doi.org/10.48550/arXiv.1904.12659
  • Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., & Tian, Q. (2021). Symbiotic graph neural networks for 3D skeleton-based human action recognition and motion prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6), 3316–3333. https://doi.org/10.1109/TPAMI.2021.3053765
  • Li, S., Cao, Q., Liu, L., Yang, K., Liu, S., Hou, J., et al. (2021). GroupFormer: Group activity recognition with clustered spatial-temporal transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 13668–13677). https://doi.org/10.48550/arXiv.2108.12630
  • Li, S., Li, W., Cook, C., & Gao, Y. (2019). Deep independently recurrent neural network (IndRNN). arXiv. https://doi.org/10.48550/arXiv.1910.06251
  • Li, Y., Fan, Y., Xiang, X., Demandolx, D., Ranjan, R., Timofte, R., et al. (2023). Efficient and explicit modelling of image hierarchies for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 18278–18289). https://doi.org/10.48550/arXiv.2303.00748
  • Li, Y., Wang, Z., Wang, L., & Wu, G. (2018). Actions as moving points. In European Conference on Computer Vision (pp. 68–84). Springer. https://doi.org/10.1007/978-3-030-01216-8_5
  • Lin, J., Keogh, E., Lonardi, S., & Chiu, B. (2003). A symbolic representation of time series, with implications for streaming algorithms. In Proceedings of the 8th ACM SIGMOD Workshop on Research Issues in Data Mining and Knowledge Discovery (pp. 2–11). ACM. https://doi.org/10.1145/882082.882086
  • Liu, M., & Yuan, J. (2018). Recognizing human actions as the evolution of pose estimation maps. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1159–1168). https://doi.org/10.1109/CVPR.2018.00127
  • Liu, Y., Yang, J., Perera, M., Ji, P., Kim, D., Xu, M., et al. (2022). Representation-centric survey of skeletal action recognition and the ANUBIS benchmark. arXiv. https://doi.org/10.2139/ssrn.6465233
  • Mei, G., Poiesi, F., Saltori, C., Zhang, J., Ricci, E., & Sebe, N. (2023). Overlap-guided Gaussian mixture models for point cloud registration. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 4511–4520). https://doi.org/10.48550/arXiv.2210.09836
  • Meng, Z., Zhang, M., Guo, C., Fan, Q., Zhang, H., Gao, N., et al. (2020). Recent progress in sensing and computing techniques for human activity recognition and motion analysis. Electronics, 9(9), 1357. https://doi.org/10.3390/electronics9091357
  • Nadeem, A., Jalal, A., & Kim, K. (2021). Automatic human posture estimation for sport activity recognition with robust body parts detection and entropy Markov model. Multimedia Tools and Applications, 80(14), 21465–21498. https://doi.org/10.1007/s11042-021-10687-5
  • Nadeem, M. S. (2024). Hybrid architecture for human action recognition using skeleton data [Master’s thesis]. https://hdl.handle.net/10155/1901
  • Nguyen, H.-C., Nguyen, T.-H., Scherer, R., & Le, V.-H. (2023). Deep learning for human activity recognition on 3D human skeleton: Survey and comparative study. Sensors, 23(11), 5121. https://doi.org/10.3390/s23115121
  • Ozcan, T., & Basturk, A. (2020). Human action recognition with deep learning and structural optimization using a hybrid heuristic algorithm. Cluster Computing, 23(4), 2847–2860. https://doi.org/10.1007/s10586-020-03050-0
  • Pareek, P., & Thakkar, A. (2021). A survey on video-based human action recognition: Recent updates, datasets, challenges, and applications. Artificial Intelligence Review, 54(3), 2259–2322. https://doi.org/10.1007/s10462-020-09904-8
  • Paul, M., Haque, S. M. E., & Chakraborty, S. (2013). Human detection in surveillance videos and its applications—A review. EURASIP Journal on Advances in Signal Processing, 2013(1), 176. https://doi.org/10.1186/1687-6180-2013-176
  • Plizzari, C., Cannici, M., & Matteucci, M. (2021). Spatial temporal transformer network for skeleton-based action recognition. In International Conference on Pattern Recognition (pp. 694–701). Springer. https://doi.org/10.1007/978-3-030-68796-0_50
  • Poppe, R. (2010). A survey on vision-based human action recognition. Image and Vision Computing, 28(6), 976–990. https://doi.org/10.1016/j.imavis.2009.11.014
  • Qi, M., Wang, Y., Qin, J., Li, A., Luo, J., & Van Gool, L. (2020). StagNet: An attentive semantic RNN for group activity and individual action recognition. IEEE Transactions on Circuits and Systems for Video Technology, 30(2), 549–565. https://doi.org/10.1109/TCSVT.2019.2894161
  • Ramasamy Ramamurthy, S., & Roy, N. (2018). Recent trends in machine learning for human activity recognition—A survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 8(4), e1254. https://doi.org/10.1002/widm.1254
  • Raziani, S., & Azimbagirad, M. (2022). Deep CNN hyperparameter optimization algorithms for sensor-based human activity recognition. Neuroscience Informatics, 2(3), 100078. https://doi.org/10.1016/j.neuri.2022.100078
  • Ren, B., Liu, M., Ding, R., & Liu, H. (2024). A survey on 3D skeleton-based action recognition using learning methods. Cyborg and Bionic Systems, 5, 100. https://doi.org/10.34133/cbsystems.0100
  • Ren, B., Liu, Y., Song, Y., Bi, W., Cucchiara, R., Sebe, N., et al. (2023). Masked jigsaw puzzle: A versatile position embedding for vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 20382–20391). https://doi.org/10.48550/arXiv.2205.12551
  • Seidenari, L., Varano, V., Berretti, S., Del Bimbo, A., & Pala, P. (2013). Recognizing actions from depth cameras as weakly aligned multi-part bag-of-poses. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (pp. 479–485). https://doi.org/10.1109/CVPRW.2013.77
  • Shahroudy, A., Liu, J., Ng, T.-T., & Wang, G. (2016). NTU RGB+D: A large scale dataset for 3D human activity analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1010–1019). https://doi.org/10.48550/arXiv.1604.02808
  • Shi, L., Zhang, Y., Cheng, J., & Lu, H. (2019). Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12026–12035). https://doi.org/10.48550/arXiv.1805.07694
  • Shi, L., Zhang, Y., Cheng, J., & Lu, H. (2020). Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition. In Proceedings of the Asian Conference on Computer Vision. https://link.springer.com/conference/accv
  • Shreyas, D. G., Raksha, S., & Prasad, B. G. (2020). Implementation of an anomalous human activity recognition system. SN Computer Science, 1(3), 168. https://doi.org/10.1007/s42979-020-00169-0
  • Shu, X., Zhang, L., Sun, Y., & Tang, J. (2021). Host–parasite: Graph LSTM-in-LSTM for group activity recognition. IEEE Transactions on Neural Networks and Learning Systems, 32(2), 663–674. https://doi.org/10.1109/TNNLS.2020.2978942
  • Si, C., Chen, W., Wang, W., Wang, L., & Tan, T. (2019). An attention enhanced graph convolutional LSTM network for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1227–1236).
  • Sighencea, B. I., Stanciu, I. R., & Căleanu, C. D. (2023). D-STGCN: Dynamic pedestrian trajectory prediction using spatio-temporal graph convolutional networks. Electronics, 12(3), 611. https://doi.org/10.3390/electronics12030611
  • Singh, R., Sonawane, A., & Srivastava, R. (2020). Recent evolution of modern datasets for human activity recognition: A deep survey. Multimedia Systems, 26(2), 83–106. https://doi.org/10.1007/s00530-019-00635-7
  • Subetha, T., & Chitrakala, S. (2016). A survey on human activity recognition from videos. In 2016 International Conference on Information Communication and Embedded Systems (ICICES) (pp. 1–7). IEEE. https://doi.org/10.1109/ICICES.2016.7518920
  • Sundaresan, K., Jayarajan, P., & Nallakumar, R. (2025). Exploring human activity recognition systems: Insights from computer vision approaches. Research Square. https://doi.org/10.21203/rs.3.rs-5923670/v1
  • Tahir, B. S., Ageed, Z. S., Hasan, S. S., & Zeebaree, S. R. M. (2023). Modified wild horse optimization with deep learning enabled symmetric human activity recognition model. Computers, Materials & Continua, 75(2), 4009–4024. https://doi.org/10.32604/cmc.2023.037433
  • Tien, P. W., Wei, S., Calautit, J. K., Darkwa, J., & Wood, C. (2021). Vision-based human activity recognition for reducing building energy demand. Building Services Engineering Research and Technology, 42(6), 691–713. https://doi.org/10.1177/01436244211026120
  • Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., & Jégou, H. (2021). Training data-efficient image transformers and distillation through attention. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 10347–10357). PMLR. https://proceedings.mlr.press/v139/touvron21a.html
  • Trivedi, N., & Sarvadevabhatla, R. K. (2022). Psumnet: Unified modality part streams are all you need for efficient pose-based action recognition. In European Conference on Computer Vision (pp. 211–227). Springer. https://github.com/skelemoa/psumnet
  • Vaswani, A. (2017). Attention is all you need. arXiv. https://cir.nii.ac.jp/crid/1370580229800306054
  • Wan, S., Qi, L., Xu, X., Tong, C., & Gu, Z. (2020). Deep learning models for real-time human activity recognition with smartphones. Mobile Networks and Applications, 25(2), 743–755. https://doi.org/10.1007/s11036-019-01445-x
  • Wang, B., Huang, L., & Hoai, M. (2020). Active vision for early recognition of human actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1081–1091). https://doi.org/10.1109/CVPR42600.2020.00116
  • Wang, H., Kläser, A., Schmid, C., & Liu, C.-L. (2013). Dense trajectories and motion boundary descriptors for action recognition. International Journal of Computer Vision, 103(1), 60–79. https://doi.org/10.1007/s11263-012-0594-8
  • Wang, J., Liu, Z., Wu, Y., & Yuan, J. (2012). Mining actionlet ensemble for action recognition with depth cameras. In 2012 IEEE Conference on Computer Vision and Pattern Recognition (pp. 1290–1297). IEEE. https://doi.org/10.1109/CVPR.2012.6247813
  • Wang, J., Nie, X., Xia, Y., Wu, Y., & Zhu, S.-C. (2014). Cross-view action modeling, learning and recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2649–2656). https://doi.org/10.1109/CVPR.2014.339
  • Wang, L., Ding, Z., Tao, Z., Liu, Y., & Fu, Y. (2019). Generative multi-view human action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 6212–6221). https://doi.org/10.1109/ICCV.2019.00631
  • Wang, W., Mei, G., Ren, B., Huang, X., Poiesi, F., Van Gool, L., et al. (2023). Zero-shot point cloud registration. arXiv. https://doi.org/10.48550/arXiv.2312.06063
  • Wu, C., Wu, X.-J., & Kittler, J. (2019). Spatial residual layer and dense connection block enhanced spatial temporal graph convolutional network for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops.
  • Wu, J., Wang, L., Wang, L., Guo, J., & Wu, G. (2019). Learning actor relation graphs for group activity recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9964–9974). https://doi.org/10.48550/arXiv.1904.10117
  • Xia, L., Chen, C.-C., & Aggarwal, J. K. (2012). View invariant human action recognition using histograms of 3D joints. In 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops (pp. 20–27). IEEE. https://doi.org/10.1109/CVPRW.2012.6239233
  • Xiang, W., Li, C., Zhou, Y., Wang, B., & Zhang, L. (2023). Generative action description prompts for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 10276–10285). https://doi.org/10.48550/arXiv.2208.05318
  • Xing, Y., & Zhu, J. (2021). Deep learning-based action recognition with 3D skeleton: A survey. CAAI Transactions on Intelligence Technology, 6, 80–92. https://doi.org/10.1049/cit2.12014
  • Yan, S., Xiong, Y., & Lin, D. (2018). Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence. https://doi.org/10.1609/aaai.v32i1.12328
  • Yang, D., Li, M. M., Fu, H., Fan, J., Zhang, Z., & Leung, H. (2020). Unifying graph embedding features with graph convolutional networks for skeleton-based action recognition. arXiv. https://doi.org/10.48550/arXiv.2003.03007
  • Yang, X., Li, S., Niu, S., & Yue, X. (2025). Graph network learning for human skeleton modeling: A survey. Artificial Intelligence Review, 59(1), 31. https://doi.org/10.1007/s10462-025-11442-0
  • Yang, Z., Li, Y., Yang, J., & Luo, J. (2018). Action recognition with spatio–temporal visual attention on skeleton image sequences. IEEE Transactions on Circuits and Systems for Video Technology, 29(8), 2405–2415. https://doi.org/10.1109/TCSVT.2018.2864148
  • Ye, F., Pu, S., Zhong, Q., Li, C., Xie, D., & Tang, H. (2020). Dynamic GCN: Context-enriched topology learning for skeleton-based action recognition. In Proceedings of the 28th ACM International Conference on Multimedia (pp. 55–63). ACM. https://doi.org/10.1145/3394171.3413941
  • Ye, L., Rochan, M., Liu, Z., & Wang, Y. (2019). Cross-modal self-attention network for referring image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10502–10511). https://doi.org/10.1109/CVPR.2019.01075
  • Yu, S., Xie, L., Liu, L., & Xia, D. (2019). Learning long-term temporal features with deep neural networks for human action recognition. IEEE Access, 8, 1840–1850. https://doi.org/10.1109/ACCESS.2019.2962284
  • Yuan, L., He, Z., Wang, Q., Xu, L., & Ma, X. (2022). Spatial transformer network with transfer learning for small-scale fine-grained skeleton-based Tai Chi action recognition. In IECON 2022–48th Annual Conference of the IEEE Industrial Electronics Society (pp. 1–6). IEEE. https://doi.org/10.1109/IECON49645.2022.9968668
  • Yun, K., Honorio, J., Chattopadhyay, D., Berg, T. L., & Samaras, D. (2012). Two-person interaction detection using body-pose features and multiple instance learning. In 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops (pp. 28–35). IEEE. https://doi.org/10.1109/CVPRW.2012.6239234
  • Zahin, A., Tan, L. T., & Hu, R. Q. (2019). Sensor-based human activity recognition for smart healthcare: A semi-supervised machine learning approach. In Artificial intelligence for communications and networks (pp. 450–472). Springer. https://doi.org/10.1007/978-3-642-29336-8_15
  • Zappardino, F., Uricchio, T., Seidenari, L., & Del Bimbo, A. (2021). Learning group activities from skeletons without individual action labels. In 2020 25th International Conference on Pattern Recognition (ICPR) (pp. 10412–10417). IEEE. https://doi.org/10.1109/ICPR48806.2021.9413195
  • Zhang, J., & Li, Y. (2024). SL-GCNN: A graph convolutional neural network for granular human motion recognition. IEEE Access. https://doi.org/10.1109/ACCESS.2024.3514082
  • Zhang, J., Jia, Y., Xie, W., & Tu, Z. (2022). Zoom transformer for skeleton-based group activity recognition. IEEE Transactions on Circuits and Systems for Video Technology, 32(12), 8646–8659. https://doi.org/10.1109/TCSVT.2022.3193574
  • Zhang, P., Lan, C., Xing, J., Zeng, W., Xue, J., & Zheng, N. (2017). View adaptive recurrent neural networks for high performance human action recognition from skeleton data. In Proceedings of the IEEE International Conference on Computer Vision (pp. 2117–2126). https://doi.org/10.48550/arXiv.1703.08274
  • Zhang, P., Lan, C., Zeng, W., Xing, J., Xue, J., & Zheng, N. (2020). Semantics-guided neural networks for efficient skeleton-based human action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 1112–1121). https://doi.org/10.48550/arXiv.1904.01189
  • Zhang, P., Xue, J., Lan, C., Zeng, W., Gao, Z., & Zheng, N. (2019). EleAtt-RNN: Adding attentiveness to neurons in recurrent neural networks. IEEE Transactions on Image Processing, 29, 1061–1073. https://doi.org/10.1109/TIP.2019.2937724
  • Zhang, W., Liu, Z., Zhou, L., Leung, H., & Chan, A. B. (2017). Martial arts, dancing and sports dataset: A challenging stereo and multi-view dataset for 3D human pose estimation. Image and Vision Computing, 61, 22–39. https://doi.org/10.1016/j.imavis.2017.02.002
  • Zhao, R., Wang, K., Su, H., & Ji, Q. (2019). Bayesian graph convolution LSTM for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 6881–6890). https://doi.org/10.1109/ICCV.2019.00698
  • Zhou, X., Liang, W., Kevin, I., Wang, K., Wang, H., Yang, L. T., et al. (2020). Deep-learning-enhanced human activity recognition for internet of healthcare things. IEEE Internet of Things Journal, 7(7), 6429–6438. https://doi.org/10.1109/JIOT.2020.2985082
  • Zhu, W., Ma, X., Liu, Z., Liu, L., Wu, W., & Wang, Y. (2023). MotionBERT: A unified perspective on learning human motion representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15085–15099). https://doi.org/10.48550/arXiv.2210.06551

How to Cite

Kamalpreet Kaur, Ankit Bansal, and Baljit Singh Khehra. Optimization-Driven Deep Learning Models for 3D Human Action Recognition: A Survey. J. Multidiscip. Res. Healthcare. 2026, 12, 10-36
Optimization-Driven Deep Learning Models for 3D Human Action Recognition: A Survey

Current Issue

PeriodicityBiannually
Issue-1June
Issue-2December
ISSN Print2393-8536
ISSN Online2393-8544
RNI No.CHAENG/2014/57978

This work is licensed under a Creative Commons Attribution 4.0 International License.

Articles in Journal of Multidisciplinary Research in Healthcare by Chitkara University Publications are Open Access articles that are published with licensed under a Creative Commons Attribution- CC-BY 4.0 International License. Based on a work at https://jmrh.chitkara.edu.in/. This license permits one to use, remix, tweak and reproduction in any medium, even commercially provided one give credit for the original creation.

View Legal Code of the above-mentioned license, https://creativecommons.org/licenses/by/4.0/legalcode

View Licence Deed here https://creativecommons.org/licenses/by/4.0/

Creative Commons License

Journal of Multidisciplinary Research in Healthcare by Chitkara University Publications is licensed under a Creative Commons Attribution 4.0 International License.
Based on a work at https://jmrh.chitkara.edu.in/

Visibility, Memberships and Ethics