Volume 13
Issue 6
IEEE/CAA Journal of Automatica Sinica
| Citation: | D. Su, J. Wang, C. Yang, and W. Gui, “A two-timescale neurodynamic approach to sharpness-aware minimization in deep learning,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1530–1532, Jun. 2026. doi: 10.1109/JAS.2026.126020 |
| [1] |
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro, “Exploring generalization in deep learning,” in Proc. Advances in NeurIPS, Long Beach, USA, 2017, pp. 5947–5956.
|
| [2] |
X. Wen and M. Zhou, “Evolution and role of optimizers in training deep learning models,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 10, pp. 2039–2042, 2024. doi: 10.1109/JAS.2024.124806
|
| [3] |
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in Proc. ICLR, Toulon, France, 2017.
|
| [4] |
D. Su, L. Jin, and J. Wang, “Noise-resistant sharpness-aware minimization in deep learning,” Neural Networks, vol. 181, Art. no. 106829, 2025. doi: 10.1016/j.neunet.2024.106829
|
| [5] |
D. Su, J. Han, C. Yang, and W. Gui, “Optimization algorithms based on double-integral coevolutionary neurodynamics in deep learning,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 6, pp. 1236–1245, 2025. doi: 10.1109/JAS.2025.125210
|
| [6] |
Z. X. Li, Y. L. Wang, and F. Wang, “DI-YOLOv5: An improved dual-wavelet-based YOLOv5 for dense small object detection,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 2, pp. 457–459, 2025. doi: 10.1109/JAS.2024.124368
|
| [7] |
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio, “Fantastic generalization measures and where to find them,” in Proc. ICLR, Addis Ababa, Ethiopia, 2020.
|
| [8] |
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, “Entropy-SGD: Biasing gradient descent into wide valleys,” J. Stat. Mech. Theor. Exp., Art. no. 124018, 2019.
|
| [9] |
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, “Sharpness-aware minimization for efficiently improving generalization,” in Proc. ICLR, 2021.
|
| [10] |
B. Li and G. Giannakis, “Enhancing sharpness-aware optimization through variance suppression,” in Proc. Advances in Neural Information Processing Systems, New Orleans, LA, USA, 2023, pp. 70861–70879.
|
| [11] |
P. Khanh, H. C. Luong, B. Mordukhovich, and D. Tran, “Fundamental convergence analysis of sharpness-aware minimization” in Proc. Advances in NeurIPS, Vancouver, Canada, 2024, pp. 13149–13182.
|
| [12] |
S. Li and C. C. Cheah, “Learning laws for deep convolutional neural networks with guaranteed convergence,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 1, pp. 170–185, 2026. doi: 10.1109/JAS.2025.125171
|
| [13] |
J. Kwon, J. Kim, H. Park, and I. K. Choi, “ASAM: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks,” in Proc. ICML, 2021, pp. 5905–5914.
|
| [14] |
X. Le and J. Wang, “A two-time-scale neurodynamic approach to constrained minimax optimization,” IEEE Trans. Neural Netw. Learn. Syst., vol. 28, no. 3, pp. 620–629, 2017. doi: 10.1109/TNNLS.2016.2538288
|
| [15] |
Z. Zuo, C. Liu, Q.-L. Han, and J. Song, “Unmanned aerial vehicles: Control methods and future challenges,” IEEE/CAA J. Autom. Sinica, vol. 9, no. 4, pp. 601–614, 2022. doi: 10.1109/JAS.2022.105410
|
| [16] |
J. Wang and J. Wang, “Two-timescale multilayer recurrent neural networks for nonlinear programming,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 1, pp. 37–47, 2022. doi: 10.1109/TNNLS.2020.3027471
|
| [17] |
H. Che and J. Wang, “A two-timescale duplex neurodynamic approach to mixed-integer optimization,” IEEE Trans. Neural Netw. Learn. Syst., vol. 32, no. 1, pp. 36–48, 2021. doi: 10.1109/TNNLS.2020.2973760
|
| [18] |
Z. Xia, Y. Liu, J. Wang, and J. Wang, “Two-timescale recurrent neural networks for distributed minimax optimization,” Neural Networks, vol. 165, pp. 527–539, 2023. doi: 10.1016/j.neunet.2023.06.003
|
| [19] |
S. Zeng and T. Doan, “Fast two-time-scale stochastic gradient method with applications in reinforcement learning,” in Proc. Conf. on Learning Theory, PMLR, Edmonton, Canada, 2024, pp. 5166–5212.
|
| [20] |
J. Chae, K. Kim, and D. Kim, “Two-timescale extragradient for finding local minimax points,” in Proc. ICLR, Vienna, Austria, 2024.
|
| [21] |
T. Lin, C. Jin, and M. I. Jordan, “Two-timescale gradient descent ascent algorithms for nonconvex minimax optimization,” J. Mach. Learn. Res., vol. 26, no. 11, pp. 1–45, 2025.
|