A journal of IEEE and CAA , publishes high-quality papers in English on original theoretical/experimental research and development in all areas of automation
Volume 13 Issue 6
Jun.  2026

IEEE/CAA Journal of Automatica Sinica

  • JCR Impact Factor: 18.3, Top 1 (SCI Q1)
    CiteScore: 28.2, Top 1% (Q1)
    Google Scholar h5-index: 95, TOP 5
Turn off MathJax
Article Contents
G. Wang and H. Zhang, “Data-driven algorithms for finite-horizon and infinite-horizon indefinite linear quadratic stochastic optimal control problems,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1459–1469, Jun. 2026. doi: 10.1109/JAS.2026.125747
Citation: G. Wang and H. Zhang, “Data-driven algorithms for finite-horizon and infinite-horizon indefinite linear quadratic stochastic optimal control problems,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1459–1469, Jun. 2026. doi: 10.1109/JAS.2026.125747

Data-Driven Algorithms for Finite-Horizon and Infinite-Horizon Indefinite Linear Quadratic Stochastic Optimal Control Problems

doi: 10.1109/JAS.2026.125747
Funds:  This work was supported in part by the National Key Research and Development Program of China (2022YFA1006100), the National Natural Science Foundation of China (61925306), and the Natural Science Foundation of Shandong Province (ZR2019ZD42)
More Information
  • This paper is devoted to devising data-driven algorithms for finite-horizon and infinite-horizon linear quadratic stochastic optimal control (LQSOC) problems. In our study, the diffusion terms of system dynamics are permitted to hinge upon both control and state variables, and the weighting matrices of cost functionals are allowed to be indefinite. It is acknowledged that the optimal controls of finite-horizon and infinite-horizon indefinite LQSOC problems are correlated with a generalized differential Riccati equation (GDRE) and a generalized algebraic Riccati equation (GARE). Herein, we propose two data-driven algorithms to approximate the solutions of these Riccati equations, and thereby determine optimal controls, without leveraging the information of all system parameters. Additionally, we prove the convergence of these algorithms and examine the impact of computational errors. Finally, we validate the performance of these data-driven algorithms via three simulation examples.

     

  • loading
  • [1]
    A. M. Letov, “Analytical design of control systems,” Autom. Remote Control, vol. 22, no. 4, pp. 363–372, Apr. 1961.
    [2]
    R. E. Bellman, I. Glicksberg, and O. A. Gross, Some Aspects of the Mathematical Theory of Control Processes. Santa Monica, USA: RAND Corporation, 1958.
    [3]
    R. E. Kalman, “Contributions to the theory of optimal control,” Bol. Soc. Mat. Mexicana, vol. 5, pp. 102−119, 1960.
    [4]
    W. M. Wonham, “On a matrix Riccati equation of stochastic control,” SIAM J. Control, vol. 6, no. 4, pp. 681–697, Nov. 1968. doi: 10.1137/0306044
    [5]
    J. Yong and X. Zhou, Stochastic Controls: Hamiltonian Systems and HJB Equations. New York, USA: Springer, 1999.
    [6]
    P. Huang, G. Wang, S. Wang, and H. Xiao, “A mean-field game for a forward-backward stochastic system with partial observation and common noise,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 3, pp. 746–759, Mar. 2024. doi: 10.1109/JAS.2023.124047
    [7]
    B. D. O. Anderson and J. B. Moore, Optimal Control: Linear Quadratic Methods. Mineola, USA: Dover Publications, 2007.
    [8]
    S. Chen, X. Li, and X. Zhou, “Stochastic linear quadratic regulators with indefinite control weight costs,” SIAM J. Control Optim., vol. 36, no. 5, pp. 1685–1702, Jan. 1998. doi: 10.1137/S0363012996310478
    [9]
    S. Chen and X. Zhou, “Stochastic linear quadratic regulators with indefinite control weight costs. II,” SIAM J. Control Optim., vol. 39, no. 4, pp. 1065–1081, Jan. 2000. doi: 10.1137/S0363012998346578
    [10]
    M. A. Rami and X. Zhou, “Linear matrix inequalities, Riccati equations, and indefinite stochastic linear quadratic controls,” IEEE Trans. Autom. Control, vol. 45, no. 6, pp. 1131–1143, Jun. 2000. doi: 10.1109/9.863597
    [11]
    M. A. Rami, X. Chen, J. B. Moore, and X. Zhou, “Solvability and asymptotic behavior of generalized Riccati equations arising in indefinite stochastic LQ controls,” IEEE Trans. Autom. Control, vol. 46, no. 3, pp. 428–440, Mar. 2001. doi: 10.1109/9.911419
    [12]
    J. Sun, X. Li, and J. Yong, “Open-loop and closed-loop solvabilities for stochastic linear quadratic optimal control problems,” SIAM J. Control Optim., vol. 54, no. 5, pp. 2274–2308, Jan. 2016. doi: 10.1137/15M103532X
    [13]
    R. C. Merton, “On estimating the expected return on the market: An exploratory investigation,” J. Financ. Econ., vol. 8, no. 4, pp. 323−361, Dec. 1980.
    [14]
    D. G. Luenberger, Investment Science. Oxford, UK: Oxford University Press, 1998.
    [15]
    K. Du, Q. Meng, and F. Zhang, “A Q-learning algorithm for discrete-time linear-quadratic control with random parameters of unknown distribution: Convergence and stabilization,” SIAM J. Control Optim., vol. 60, no. 4, pp. 1991–2015, Aug. 2022. doi: 10.1137/20M1379605
    [16]
    Y. Wang, Y. Ni, Z. Chen, and J. Zhang, “Probabilistic framework of Howard’s policy iteration: BML evaluation and robust convergence analysis,” IEEE Trans. Autom. Control, vol. 69, no. 8, pp. 5200–5215, Aug. 2024. doi: 10.1109/TAC.2023.3344870
    [17]
    B. Pang and Z. Jiang, “Reinforcement learning for adaptive optimal stationary control of linear stochastic systems,” IEEE Trans. Autom. Control, vol. 68, no. 4, pp. 2383–2390, Apr. 2023. doi: 10.1109/TAC.2022.3172250
    [18]
    W. Zhang, J. Guo, and X. Jiang, “Model-free H control of Itô stochastic system via off-policy reinforcement learning,” Automatica, vol. 174, Art. no. 112144, Apr. 2025. doi: 10.1016/j.automatica.2025.112144
    [19]
    M. Basei, X. Guo, A. Hu, and Y. Zhang, “Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon,” J. Mach. Learn. Res., vol. 23, no. 1, Art. no. 178, Jan. 2022.
    [20]
    H. Zhang, “An adaptive dynamic programming-based algorithm for infinite-horizon linear quadratic stochastic optimal control problems,” J. Appl. Math. Comput., vol. 69, no. 3, pp. 2741–2760, Jun. 2023. doi: 10.1007/s12190-023-01857-9
    [21]
    G. Wang and H. Zhang, “System transformation and model-free value iteration algorithms for continuous-time linear quadratic stochastic optimal control problems,” Int. J. Syst. Sci., vol. 56, no. 2, pp. 293−302, 2025.
    [22]
    H. Wang, T. Zariphopoulou, and X. Zhou, “Reinforcement learning in continuous time and space: A stochastic control approach,” J. Mach. Learn. Res., vol. 21, no. 1, Art. no. 198, Jan. 2020.
    [23]
    Y. Jia and X. Zhou, “Policy evaluation and temporal-difference learning in continuous time and space: A martingale approach,” J. Mach. Learn. Res., vol. 23, no. 1, Art. no. 154, Jan. 2022. doi: 10.2139/ssrn.3905379
    [24]
    Y. Jia and X. Zhou, “Q-learning in continuous time,” J. Mach. Learn. Res., vol. 24, no. 1, Art. no. 161, Jan. 2023.
    [25]
    Y. Jia and X. Zhou, “Policy gradient and actor-critic learning in continuous time and space: Theory and algorithms,” J. Mach. Learn. Res., vol. 23, no. 1, Art. no. 275, Jan. 2022. doi: 10.2139/ssrn.3969101
    [26]
    Y. Huang, Y. Jia, and X. Zhou, “Sublinear regret for an actor-critic algorithm in continuous-time linear-quadratic reinforcement learning,” arXiv preprint arXiv: 2407.17226, 2024.
    [27]
    Z. Chen and Q. Zhang, “Backward stochastic control system with entropy regularization,” SIAM J. Control Optim., vol. 63, no. 3, pp. 1981–2006, Jun. 2025. doi: 10.1137/24M1684700
    [28]
    H. Wang and X. Zhou, “Continuous-time mean-variance portfolio selection: A reinforcement learning framework,” Math. Finance, vol. 30, no. 4, pp. 1273–1308, Oct. 2020. doi: 10.1111/mafi.12281
    [29]
    M. Dai, Y. Dong, and Y. Jia, “Learning equilibrium mean-variance strategy,” Math. Finance, vol. 33, no. 4, pp. 1166–1212, Oct. 2023. doi: 10.1111/mafi.12402
    [30]
    L. Cui, B. Pang, and Z. Jiang, “Reinforcement-learning-based risk-sensitive optimal feedback mechanisms of biological motor control,” in Proc. 62nd IEEE Conf. Decision and Control, Singapore, 2023, pp. 7944−7949.
    [31]
    G. Wang and H. Zhang, “Finite-horizon and infinite-horizon linear quadratic optimal control problems: A data-driven Euler scheme,” J. Franklin Inst., vol. 361, no. 13, Art. no. 107054, Sep. 2024. doi: 10.1016/j.jfranklin.2024.107054
    [32]
    K. J. Åström and B. Wittenmark, Adaptive Control. 2nd ed. Reading, USA: Addison-Wesley, 1995.
    [33]
    G. Tao, “Multivariable adaptive control: A survey,” Automatica, vol. 50, no. 11, pp. 2737–2764, Nov. 2014. doi: 10.1016/j.automatica.2014.10.015
    [34]
    Y. Jiang and Z. Jiang, “Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics,” Automatica, vol. 48, no. 10, pp. 2699–2704, Oct. 2012. doi: 10.1016/j.automatica.2012.06.096
    [35]
    K. Xie, X. Yu, and W. Lan, “Optimal output regulation for unknown continuous-time linear systems by internal model and adaptive dynamic programming,” Automatica, vol. 146, Art. no. 110564, Dec. 2022. doi: 10.1016/j.automatica.2022.110564
    [36]
    G. Wang and H. Zhang, “Value iteration algorithm for continuous-time linear quadratic stochastic optimal control problems,” Sci. China Inf. Sci., vol. 67, no. 2, Art. no. 122204, Feb. 2024. doi: 10.1007/s11432-023-3820-3

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(9)

    Article Metrics

    Article views (368) PDF downloads(18) Cited by()

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return