Background: Due to its numerous applications in healthcare, surveillance, human-computer interaction, sports analytics, and smart environments, Human Action Recognition (HAR) using 3D skeletal data has grown in importance as a field of study. Conventional HAR techniques mainly relied on manually created features, which were unreliable and data dependent. The ability to model intricate spatiotemporal patterns in human motion has greatly improved with the development of deep learning models like CNNs, RNNs, GCNs, and hybrid architectures. The integration of optimization techniques into deep learning-based HAR systems is motivated by unresolved issues like high computational cost, sensitivity to hyperparameters, poor cross-dataset generalization, and lack of interpretability.
Purpose: This survey’s goal is to offer a thorough analysis of optimization-driven deep learning models for 3D human action recognition. The research seeks to examine cutting-edge deep learning architectures for 3D HAR and analyze how metaheuristic optimization techniques can enhance the precision, effectiveness, and generalization of models.
Methods: Using a methodical survey approach, this study covers deep learning architectures such as Transformer-based, CNN-based, RNN-based, GCN-based, and Hybrid-DNN models. Optimization techniques include Chaos Game Optimization, Whale Optimization Algorithm, Grey Wolf Optimizer, Rao-3, Wild Horse Optimization, and Sea Horse Optimization. Model comparison is conducted using benchmark datasets such as NTU RGB+D, NTU RGB+D 120, Kinetics-Skeleton, SYSU-3D, and N-UCLA. Assessment uses standard performance metrics, mainly accuracy, along with reported robustness and computational efficiency.
Results: According to the survey, optimization-driven deep learning models routinely perform better in 3D HAR than conventional and non-optimized methods. Higher recognition accuracy is attained by optimized models, frequently surpassing 90–95% on benchmark datasets. Hyperparameter tuning is greatly enhanced by metaheuristic optimization, which lowers overfitting and computational inefficiency. Despite advancements, problems like cross-dataset generalization, model explainability, and real-time processing still exist.
Conclusion: Recent developments in optimization-driven deep learning models for 3D Human Action Recognition (HAR) were examined in this survey. The analysis demonstrates that combining metaheuristic optimization techniques with deep learning architectures greatly increases recognition accuracy, generalization, and computational efficiency.
Kamalpreet Kaur, Ankit Bansal, and Baljit Singh Khehra. Optimization-Driven Deep Learning Models for 3D Human Action Recognition: A Survey.
. 2026, 12, 10-36