A journal of IEEE and CAA , publishes high-quality papers in English on original theoretical/experimental research and development in all areas of automation
Volume 13 Issue 6
Jun.  2026

IEEE/CAA Journal of Automatica Sinica

  • JCR Impact Factor: 18.3, Top 1 (SCI Q1)
    CiteScore: 28.2, Top 1% (Q1)
    Google Scholar h5-index: 95, TOP 5
Turn off MathJax
Article Contents
Y.-S. Ma, J. Sun, and Y. Xu, “H∞ optimal output regulation of unknown linear systems via an adaptive dynamic programming and internal model,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1314–1324, Jun. 2026. doi: 10.1109/JAS.2026.125777
Citation: Y.-S. Ma, J. Sun, and Y. Xu, “H optimal output regulation of unknown linear systems via an adaptive dynamic programming and internal model,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1314–1324, Jun. 2026. doi: 10.1109/JAS.2026.125777

H Optimal Output Regulation of Unknown Linear Systems via an Adaptive Dynamic Programming and Internal Model

doi: 10.1109/JAS.2026.125777
Funds:  This work was supported by the National Natural Science Foundation of China (62322305, 62495090, 62495095)
More Information
  • This paper delves into the $ H_\infty $ optimal output regulation problem for continuous-time linear systems with an unknown system model. By integrating the internal model principle with optimal control, we derive an optimal control policy and a worst-case disturbance policy through the formulation and solution of a zero-sum game problem. Subsequently, leveraging adaptive dynamic programming, we propose a policy iteration learning algorithm capable of learning both the optimal control policy and the worst-case disturbance policy directly from system data. The existing algorithms necessitate an initial stabilizing policy, a full-rank condition, and the storage of historical data to guarantee algorithm convergence. In contrast, we design a dual policy iteration algorithm equipped with an online learning mechanism, thereby eliminating these additional prerequisites. Simulation results with an autonomous ground vehicle underscore the effectiveness of our proposed algorithm, and its superiority is further demonstrated through comparisons with existing methodologies.

     

  • loading
  • [1]
    C. Deng, W. Gao, C. Wen, Z. Chen, and W. Wang, “Data-driven practical cooperative output regulation under actuator faults and DoS attacks,” IEEE Trans. Cybern., vol. 53, no. 11, pp. 7417–7428, Nov. 2023. doi: 10.1109/TCYB.2023.3263480
    [2]
    Y. Xu and Z.-G. Wu, “Online learning algorithm design for adaptive output regulation with initial excitation,” IEEE Trans. Automat. Contr., vol. 70, no. 9, pp. 6300–6307, Sep. 2025. doi: 10.1109/TAC.2025.3558612
    [3]
    Y.-S. Ma, J. Sun, Y. Xu, S.-S. Cui, and Z.-G. Wu, “Adaptive dynamic programming for optimal control of unknown LTI system via interval excitation,” IEEE Trans. Autom. Control, vol. 70, no. 7, pp. 4896–4903, Jul. 2025. doi: 10.1109/TAC.2025.3542328
    [4]
    W. Gao, Z.-P. Jiang, and T. Chai, “Resilient control under denial-of-service and uncertainty: An adaptive dynamic programming approach,” IEEE Trans. Autom. Control, vol. 70, no. 6, pp. 4085–4092, Jun. 2025. doi: 10.1109/TAC.2025.3527305
    [5]
    D. Wang, N. Gao, D. Liu, N. Li, and F. L. Lewis, “Recent progress in reinforcement learning and adaptive dynamic programming for advanced control applications,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 1, pp. 18–36, Jan. 2024. doi: 10.1109/JAS.2023.123843
    [6]
    S. Fan, D. Yue, B. Wang, C. Deng, and H. Yan, “Distributed optimization for uncertain nonlinear MASs under event-triggered communication,” Automatica, vol. 177, Art. no. 112134, Jul. 2025. doi: 10.1016/j.automatica.2025.112134
    [7]
    F. Zhao, S. Luo, W. Gao, and C. Wen, “Event-triggered cooperative adaptive optimal output regulation for multiagent systems under switching network: An adaptive dynamic programming approach,” IEEE Trans. Syst. Man Cybern. Syst., vol. 55, no. 3, pp. 1707–1721, Mar. 2025. doi: 10.1109/TSMC.2024.3514202
    [8]
    Y. Jiang and Z.-P. Jiang, “Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics,” Automatica, vol. 48, no. 10, pp. 2699–2704, Oct. 2012. doi: 10.1016/j.automatica.2012.06.096
    [9]
    Y. Xu, J. Sun, Y.-J. Pan, and Z.-G. Wu, “Optimal tracking control of heterogeneous MASs using event-driven adaptive observer and reinforcement learning,” IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 4, pp. 5577–5587, Apr. 2024. doi: 10.1109/TNNLS.2022.3208237
    [10]
    Y. Xu, Z.-G. Wu, W.-W. Che, and D. Meng, “Reinforcement learning-based unknown reference tracking control of HMASs with nonidentical communication delays,” Sci. China Inf. Sci., vol. 66, no. 7, Art. no. 170203, Jul. 2023. doi: 10.1007/s11432-022-3729-7
    [11]
    C. Chen, L. Xie, K. Xie, F. L. Lewis, and S. Xie, “Adaptive optimal output tracking of continuous-time systems via output-feedback-based reinforcement learning,” Automatica, vol. 146, Art. no. 110581, Dec. 2022. doi: 10.1016/j.automatica.2022.110581
    [12]
    J. Zhao, C. Yang, W. Gao, L. Zhou, and X. Liu, “Adaptive optimal output regulation of interconnected singularly perturbed systems with application to power systems,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 3, pp. 595–607, Mar. 2024. doi: 10.1109/JAS.2023.123651
    [13]
    Z. Wang, Y. Wang, and Z. Kowalczuk, “Adaptive optimal discrete-time output-feedback using an internal model principle and adaptive dynamic programming,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 1, pp. 131–140, Jan. 2024. doi: 10.1109/JAS.2023.123759
    [14]
    T. Bian and Z.-P. Jiang, “Value iteration and adaptive dynamic programming for data-driven adaptive optimal control design,” Automatica, vol. 71, pp. 348–360, Sep. 2016. doi: 10.1016/j.automatica.2016.05.003
    [15]
    Y. Jiang, W. Gao, J. Wu, T. Chai, and F. L. Lewis, “Reinforcement learning and cooperative H output regulation of linear continuous-time multi-agent systems,” Automatica, vol. 148, Art. no. 110768, Feb. 2023. doi: 10.1016/j.automatica.2022.110768
    [16]
    W. Gao, M. Mynuddin, D. C. Wunsch, and Z.-P. Jiang, “Reinforcement learning-based cooperative optimal output regulation via distributed adaptive internal model,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 10, pp. 5229–5240, Oct. 2022. doi: 10.1109/TNNLS.2021.3069728
    [17]
    W. Gao, C. Deng, Y. Jiang, and Z.-P. Jiang, “Resilient reinforcement learning and robust output regulation under denial-of-service attacks,” Automatica, vol. 142, Art. no. 110366, Aug. 2022. doi: 10.1016/j.automatica.2022.110366
    [18]
    B. Zhang, C. Deng, and B. Wang, “Resilient optimal virtual synchronous generator control under DoS attacks,” IEEE Trans. Circuits Syst. Ⅱ: Express Briefs, vol. 72, no. 7, pp. 913–917, Jul. 2025. doi: 10.1109/TCSII.2025.3574800
    [19]
    H. Jiang and B. Zhou, “Bias-policy iteration based adaptive dynamic programming for unknown continuous-time linear systems,” Automatica, vol. 136, Art. no. 110058, Feb. 2022. doi: 10.1016/j.automatica.2021.110058
    [20]
    B. Luo, Y. Yang, H.-N. Wu, and T. Huang, “Balancing value iteration and policy iteration for discrete-time control,” IEEE Trans. Syst. Man Cybern. Syst., vol. 50, no. 11, pp. 3948–3958, Nov. 2020. doi: 10.1109/TSMC.2019.2898389
    [21]
    Y. Yang, B. Kiumarsi, H. Modares, and C. Xu, “Model-free λ-policy iteration for discrete-time linear quadratic regulation,” IEEE Trans. Neural Netw. Learn. Syst., vol. 34, no. 2, pp. 635–649, Feb. 2023. doi: 10.1109/TNNLS.2021.3098985
    [22]
    H. Jiang, B. Zhou, and G.-R. Duan, “Modified λ-policy iteration based adaptive dynamic programming for unknown discrete-time linear systems,” IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 3, pp. 3291–3301, Mar. 2024. doi: 10.1109/TNNLS.2023.3244934
    [23]
    C. Chen, F. L. Lewis, and B. Li, “Homotopic policy iteration-based learning design for unknown linear continuous-time systems,” Automatica, vol. 138, Art. no. 110153, Apr. 2022. doi: 10.1016/j.automatica.2021.110153
    [24]
    D. Liu, Q. Wei, and P. Yan, “Generalized policy iteration adaptive dynamic programming for discrete-time nonlinear systems,” IEEE Trans. Syst. Man Cybern. Syst., vol. 45, no. 12, pp. 1577–1591, Dec. 2015. doi: 10.1109/TSMC.2015.2417510
    [25]
    B. Luo, Y. Yang, and D. Liu, “Policy iteration Q-learning for data-based two-player zero-sum game of linear discrete-time systems,” IEEE Trans. Cybern., vol. 51, no. 7, pp. 3630–3640, Jul. 2021. doi: 10.1109/TCYB.2020.2970969
    [26]
    Q. Liu, H. Yan, H. Zhang, M. Wang, and Y. Tian, “Data-driven H output consensus for heterogeneous multiagent systems under switching topology via reinforcement learning,” IEEE Trans. Cybern., vol. 54, no. 12, pp. 7865–7876, Dec. 2024. doi: 10.1109/TCYB.2024.3419056
    [27]
    S. Xue, B. Luo, D. Liu, and Y. Yang, “Constrained event-triggered H control based on adaptive dynamic programming with concurrent learning,” IEEE Trans. Syst. Man Cybern. Syst., vol. 52, no. 1, pp. 357–369, Jan. 2022. doi: 10.1109/TSMC.2020.2997559
    [28]
    R. Song, L. Liu, L. Xia, and F. L. Lewis, “Online optimal event-triggered H control for nonlinear systems with constrained state and input,” IEEE Trans. Cybern., vol. 53, no. 1, pp. 131–141, Jan. 2023.
    [29]
    Z. Ming, H. Zhang, Y. Li, and Y. Liang, “Mixed H2/H control for nonlinear closed-loop Stackelberg games with application to power systems,” IEEE Trans. Autom. Sci. Eng., vol. 21, no. 1, pp. 69–77, Jan. 2024. doi: 10.1109/TASE.2022.3216733
    [30]
    S. Hu, Y. Luo, X. Xie, and H. Zhang, “H optimal load frequency control of power system: A novel model-free approach,” IEEE Trans. Circuits Syst. Ⅱ: Express Briefs, vol. 72, no. 1, pp. 228–232, Jan. 2025. doi: 10.1109/TCSⅡ.2024.3495679
    [31]
    S. K. Jha, S. B. Roy, and S. Bhasin, “Initial excitation-based iterative algorithm for approximate optimal control of completely unknown LTI systems,” IEEE Trans. Autom. Control, vol. 64, no. 12, pp. 5230–5237, Dec. 2019. doi: 10.1109/TAC.2019.2912828
    [32]
    S. K. Jha, S. B. Roy, and S. Bhasin, “Memory-efficient filter-based approximate optimal regulation of unknown LTI systems using initial excitation,” in Proc. IEEE Conf. Decision and Control, Miami Beach, USA, 2018, pp. 1638−1643.
    [33]
    S. K. Jha, S. B. Roy, and S. Bhasin, “Memory-efficient filter based novel policy iteration technique for adaptive LQR,” in Proc. Annu. American Control Conf., Milwaukee, USA, 2018, 4963−4968.
    [34]
    K. Xie, X. Yu, and W. Lan, “Optimal output regulation for unknown continuous-time linear systems by internal model and adaptive dynamic programming,” Automatica, vol. 146, Art. no. 110564, Dec. 2022. doi: 10.1016/j.automatica.2022.110564
    [35]
    H. Li and Q. Wei, “Initial excitation-based optimal control for continuous-time linear nonzero-sum games,” IEEE Trans. Syst. Man Cybern. Syst., vol. 54, no. 9, pp. 5444–5455, Sep. 2024. doi: 10.1109/TSMC.2024.3405023
    [36]
    D. Kleinman, “On an iterative technique for Riccati equation computations,” IEEE Trans. Automat. Contr., vol. 13, no. 1, pp. 114–115, Feb. 1968. doi: 10.1109/tac.1968.1098829
    [37]
    F. L. Lewis and V. L. Syrmos, Optimal Control. Hoboken, USA: John Wiley & Sons, Inc., 1995.
    [38]
    X. Hu, L. Xie, L. Xie, S. Lu, W. Xu, and H. Su, “Distributed model predictive control for vehicle platoon with mixed disturbances and model uncertainties,” IEEE Trans. Intell. Transp. Syst., vol. 23, no. 10, pp. 17354–17365, Oct. 2022. doi: 10.1109/TITS.2022.3153307

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)  / Tables(2)

    Article Metrics

    Article views (285) PDF downloads(19) Cited by()

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return