A journal of IEEE and CAA , publishes high-quality papers in English on original theoretical/experimental research and development in all areas of automation
Volume 13 Issue 7
Jul.  2026

IEEE/CAA Journal of Automatica Sinica

  • JCR Impact Factor: 18.3, Top 1 (SCI Q1)
    CiteScore: 28.2, Top 1% (Q1)
    Google Scholar h5-index: 95, TOP 5
Turn off MathJax
Article Contents
J. Li, X. Wang, X. Meng, and Frank L. Lewis, “Optimal sensor selection of linear quadratic regulation with unknown sensor noise covariances,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 7, pp. 1747–1754, Jul. 2026. doi: 10.1109/JAS.2025.125915
Citation: J. Li, X. Wang, X. Meng, and Frank L. Lewis, “Optimal sensor selection of linear quadratic regulation with unknown sensor noise covariances,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 7, pp. 1747–1754, Jul. 2026. doi: 10.1109/JAS.2025.125915

Optimal Sensor Selection of Linear Quadratic Regulation With Unknown Sensor Noise Covariances

doi: 10.1109/JAS.2025.125915
Funds:  This work was supported in part by the National Natural Science Foundation of China (62073158), the Key Science and Technology Research Project of the Education Department of Liaoning Province (LJ222410148037), and the “Xingliao Talent Program” of Liaoning Province (XLYC2402025, XLYC2203160)
More Information
  • This paper addresses an optimal sensor selection problem under the framework of linear quadratic regulation. Unlike prior work on optimal sensor scheduling, we assume that the sensor noise covariance matrices are comparable but unknown. Then, the optimal sensor selection problem is formulated as finding an optimal policy of selecting a sensor from a set of sensors to minimize the expected quadratic performance of a linear system given the number of trials. An action value method from reinforcement learning is adopted for estimating the values of selections and making selection decisions based on the estimates. Several ways of balancing exploration and exploitation are presented and compared for efficacy. Numerical simulations are conducted to demonstrate the effectiveness of the proposed algorithms.

     

  • loading
  • [1]
    C. Chen, S. Zhu, X. Guan, and X. S. Shen, Wireless Sensor Networks: Distributed Consensus Estimation. Cham, Switzerland: Springer, 2014.
    [2]
    X. Meng and T. Chen, “Optimality and stability of event triggered consensus state estimation for wireless sensor networks,” in Proc. American Control Conf., Portland, USA, 2014, pp. 3565−3570.
    [3]
    J. Zhong, Z. Huang, L. Feng, W. Du, and Y. Li, “A hyper-heuristic framework for lifetime maximization in wireless sensor networks with a mobile sink,” IEEE/CAA J. Autom. Sinica, vol. 7, no. 1, pp. 223–236, Jan. 2020. doi: 10.26686/wgtn.13200152
    [4]
    D. Bajovic, B. Sinopoli, and J. Xavier, “Sensor selection for event detection in wireless sensor networks,” IEEE Trans. Signal Process., vol. 59, no. 10, pp. 4938–4953, Oct. 2011. doi: 10.1109/TSP.2011.2160630
    [5]
    M. P. Vitus, W. Zhang, A. Abate, J. Hu, and C. J. Tomlin, “On efficient sensor scheduling for linear dynamical systems,” Automatica, vol. 48, no. 10, pp. 2482–2493, Oct. 2012. doi: 10.1016/j.automatica.2012.06.092
    [6]
    L. Zhao, W. Zhang, J. Hu, A. Abate, and C. J. Tomlin, “On the optimal solutions of the infinite-horizon linear sensor scheduling problem,” IEEE Trans. Autom. Control, vol. 59, no. 10, pp. 2825–2830, Oct. 2014. doi: 10.1109/TAC.2014.2314222
    [7]
    A. B. Asghar, S. T. Jawaid, and S. L. Smith, “A complete greedy algorithm for infinite-horizon sensor scheduling,” Automatica, vol. 81, pp. 335–341, Jul. 2017. doi: 10.1016/j.automatica.2017.04.018
    [8]
    C. Yang, J. Wu, X. Ren, W. Yang, H. Shi, and L. Shi, “Deterministic sensor selection for centralized state estimation under limited communication resource,” IEEE Trans. Signal Process., vol. 63, no. 9, pp. 2336–2348, May 2015. doi: 10.1109/TSP.2015.2412916
    [9]
    L. Ye, N. Woodford, S. Roy, and S. Sundaram, “On the complexity and approximability of optimal sensor selection and attack for Kalman filtering,” IEEE Trans. Autom. Control, vol. 66, no. 5, pp. 2146–2161, May 2021. doi: 10.1109/TAC.2020.3007383
    [10]
    K. Manohar, J. N. Kutz, and S. L. Brunton, “Optimal sensor and actuator selection using balanced model reduction,” IEEE Trans. Autom. Control, vol. 67, no. 4, pp. 2108–2115, Apr. 2022. doi: 10.1109/TAC.2021.3082502
    [11]
    L. Huang, J. Wu, Y. Mo, and L. Shi, “Joint sensor and actuator placement for infinite-horizon LQG control,” IEEE Trans. Autom. Contr., vol. 67, no. 1, pp. 398–405, Jan. 2022. doi: 10.1109/TAC.2021.3055194
    [12]
    L. Zheng, M. Liu, S. Zhang, and J. Lan, “A novel sensor scheduling algorithm based on deep reinforcement learning for bearing-only target tracking in UWSNs,” IEEE/CAA J. Autom. Sinica, vol. 10, no. 4, pp. 1077–1079, Apr. 2023. doi: 10.1109/JAS.2023.123159
    [13]
    B. Feng, M. Fu, H. Ma, Y. Xia, and B. Wang, “Kalman filter with recursive covariance estimation—sequentially estimating process noise covariance,” IEEE Trans. Ind. Electron., vol. 61, no. 11, pp. 6253–6263, Nov. 2014. doi: 10.1109/TIE.2014.2301756
    [14]
    C. Pang and S. Sun, “Fusion predictors for multisensor stochastic uncertain systems with missing measurements and unknown measurement disturbances,” IEEE Sens. J., vol. 15, no. 8, pp. 4346–4354, Aug. 2015. doi: 10.1109/JSEN.2015.2416511
    [15]
    N. Davari and A. Gholami, “An asynchronous adaptive direct Kalman filter algorithm to improve underwater navigation system performance,” IEEE Sens. J., vol. 17, no. 4, pp. 1061–1068, Feb. 2017. doi: 10.1109/JSEN.2016.2637402
    [16]
    S. Zhao, Y. S. Shmaliy, C. K. Ahn, and F. Liu, “Self-tuning unbiased finite impulse response filtering algorithm for processes with unknown measurement noise covariance,” IEEE Trans. Contr. Syst. Technol., vol. 29, no. 3, pp. 1372–1379, May 2021. doi: 10.1109/TCST.2020.2991609
    [17]
    E. Javanfar and M. Rahmani, “Data-based filters for non-Gaussian dynamic systems with unknown output noise covariance,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 4, pp. 866–877, Apr. 2024. doi: 10.1109/JAS.2023.124164
    [18]
    K. Myers and B. Tapley, “Adaptive sequential estimation with unknown noise statistics,” IEEE Trans. Autom. Control, vol. 21, no. 4, pp. 520–523, Aug. 1976. doi: 10.1109/TAC.1976.1101260
    [19]
    Z. M. Durovic and B. D. Kovacevic, “Robust estimation with unknown noise statistics,” IEEE Trans. Autom. Control, vol. 44, no. 6, pp. 1292–1296, Jun. 1999. doi: 10.1109/9.769393
    [20]
    H. Deng and M. Krstić, “Output-feedback stabilization of stochastic nonlinear systems driven by noise of unknown covariance,” Syst. Control Lett., vol. 39, no. 3, pp. 173–182, Mar. 2000. doi: 10.1016/S0167-6911(99)00084-5
    [21]
    P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,” Mach. Learn., vol. 47, no. 2, pp. 235–256, May 2002. doi: 10.1023/a:1013689704352
    [22]
    M. M. Fouda, S. Hashima, S. Sakib, Z. M. Fadlullah, K. Hatano, and X. Shen, “Optimal channel selection in hybrid RF/VLC networks: A multi-armed bandit approach,” IEEE Trans. Veh. Technol., vol. 71, no. 6, pp. 6853–6858, Jun. 2022. doi: 10.1109/TVT.2022.3163078
    [23]
    R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction. Cambridge, USA: MIT Press, 2018.
    [24]
    J. Li, J. Ding, T. Chai, F. L. Lewis, and S. Jagannathan, “Adaptive interleaved reinforcement learning: Robust stability of affine nonlinear systems with unknown uncertainty,” IEEE Trans. Neural Networks Learn. Syst., vol. 33, no. 1, pp. 270–280, Jan. 2022. doi: 10.1109/TNNLS.2020.3027653
    [25]
    J. Li, M. Yang, F. L. Lewis, and M. Zheng, “Compensator-based self-learning: Optimal operational control for two-time-scale systems with input constraints,” IEEE Trans. Ind. Inform., vol. 20, no. 7, pp. 9465–9475, Jul. 2024. doi: 10.1109/TII.2024.3384621
    [26]
    A. Ferdowsi, S. Ali, W. Saad, and N. B. Mandayam, “Cyber-physical security and safety of autonomous connected vehicles: Optimal control meets multi-armed bandit learning,” IEEE Trans. Commun., vol. 67, no. 10, pp. 7228–7244, Oct. 2019. doi: 10.1109/TCOMM.2019.2927570
    [27]
    S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu, “On the sample complexity of the linear quadratic regulator,” Found. Comput. Math., vol. 20, no. 4, pp. 633–679, Aug. 2020. doi: 10.1007/s10208-019-09426-y
    [28]
    D. Wang, N. Gao, D. Liu, J. Li, and F. L. Lewis, “Recent progress in reinforcement learning and adaptive dynamic programming for advanced control applications,” IEEE/CAA J. Autom. Sinica, vol. 11, no. 1, pp. 18–36, Jan. 2024. doi: 10.1109/JAS.2023.123843
    [29]
    Y. Abbasi-Yadkori and C. Szepesvári, “Regret bounds for the adaptive control of linear quadratic systems,” in Proc. 24th Annu. Conf. Learning Theory, 2011, pp. 1−26.
    [30]
    J. A. Chekan and C. Langbort, “Regret bounds for online-learning-based linear quadratic control under database attacks,” Automatica, vol. 151, Art. no. 110876, May 2023. doi: 10.1016/j.automatica.2023.110876
    [31]
    S. K. Jha and S. Bhasin, “Adaptive linear quadratic regulator for continuous-time systems with uncertain dynamics,” IEEE/CAA J. Autom. Sinica, vol. 7, no. 3, pp. 833–841, May 2020. doi: 10.1109/jas.2019.1911438
    [32]
    B. Kiumarsi, F. L. Lewis, M. B. Naghibi-Sistani, and A. Karimpour, “Optimal tracking control of unknown discrete-time linear systems using input-output measured data,” IEEE Trans. Cybern., vol. 45, no. 12, pp. 2770–2779, Dec. 2015. doi: 10.1109/TCYB.2014.2384016
    [33]
    H. Modares, F. L. Lewis, and Z. P. Jiang, “Optimal output-feedback control of unknown continuous-time linear systems using off-policy reinforcement learning,” IEEE Trans. Cybern., vol. 46, no. 11, pp. 2401–2410, Nov. 2016. doi: 10.1109/TCYB.2015.2477810
    [34]
    S. A. A. Rizvi and Z. Lin, “Reinforcement learning-based linear quadratic regulation of continuous-time systems using dynamic output feedback,” IEEE Trans. Cybern., vol. 50, no. 11, pp. 4670–4679, Nov. 2020. doi: 10.1109/TCYB.2018.2886735
    [35]
    J. Li, X. Wang, and X. Meng, “Learning based optimal sensor selection for linear quadratic control with unknown sensor noise covariance,” in Proc. Am. Contr. Conf., San Diego, USA, 2023, pp. 4173−4178.
    [36]
    D. S. Naidu, Optimal Control Systems. Boca Raton, USA: CRC Press, 2003.
    [37]
    F. L. Lewis, D. L. Vrabie, and V. L. Syrmos, Optimal Control. Hoboken, USA: John Wiley & Sons, 2012.
    [38]
    K. J. ÅAstrőm, Introduction to Stochastic Control Theory. Amsterdam, Netherlands: Academic Press, 1970.
    [39]
    J. Edmonds, “Matroids and the greedy algorithm,” Math. Programm., vol. 1, no. 1, pp. 127–136, 1971. doi: 10.1007/BF01584082
    [40]
    C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, Cambridge University, Cambridge, UK, 1989.

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(4)  / Tables(1)

    Article Metrics

    Article views (17) PDF downloads(0) Cited by()

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return