Volume 13
Issue 6
IEEE/CAA Journal of Automatica Sinica
| Citation: | W. Ma, Q.-L. Han, X. Zhu, W. Zhou, J. Xiong, Z. Ren, S. Wen, and Y. Xiang, “OpenClaw in the wild: Security analysis of autonomous agents,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 6, pp. 1257–1273, Jun. 2026. doi: 10.1109/JAS.2026.126209 |
| [1] |
V. Raval, M. Zeid, and P. Enjeti, “Circuit-AI: A self-hosted AI-agent language model framework for control loop implementation and simulation,” in Proc. IEEE Int. Communications Energy Conf., Houston, USA, 2025, pp. 91−96.
|
| [2] |
X. Zhu, W. Zhou, Q.-L. Han, W. Ma, S. Wen, and Y. Xiang, “When software security meets large language models: A survey,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 2, pp. 317–334, Feb. 2025. doi: 10.1109/JAS.2024.124971
|
| [3] |
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al., “A survey on large language model based autonomous agents,” Front. Comput. Sci., vol. 18, no. 6, Art. no. 186345, 2024. doi: 10.1007/s11704-024-40231-1
|
| [4] |
W. Zhou, X. Zhu, Q.-L. Han, L. Li, X. Chen, S. Wen, and Y. Xiang, “The security of using large language models: A survey with emphasis on ChatGPT,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 1, pp. 1–26, Jan. 2025. doi: 10.1109/jas.2024.124983
|
| [5] |
Z. Deng, R. Sun, M. Xue, W. Ma, S. Wen, S. Nepal, and Y. Xiang, “Hardening LLM fine-tuning: From differentially private data selection to trustworthy model quantization,” IEEE Trans. Inform. Foren. Sec., vol. 20, pp. 7211–7226, Jun. 2025. doi: 10.1109/TIFS.2025.3581103
|
| [6] |
Z. Deng, Y. Guo, C. Han, W. Ma, J. Xiong, S. Wen, and Y. Xiang, “AI agents under threat: A survey of key security challenges and future pathways,” ACM Comput. Surv., vol. 57, no. 7, Art. no. 182, Feb. 2025.
|
| [7] |
Z. Li, W. Wu, Y. Guo, J. Sun, and Q.-L. Han, “Embodied multi-agent systems: A review,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 6, pp. 1095–1116, Jun. 2025. doi: 10.1109/JAS.2025.125552
|
| [8] |
P. Steinberger, “OpenClaw: Your own personal AI assistant,” 2026. [Online]. Available: https://github.com/openclaw/openclaw.
|
| [9] |
LangChain AI, “LangGraph: Build resilient language agents as graphs,” 2024. [Online]. Available: https://github.com/langchain-ai/langgraph
|
| [10] |
CrewAI Inc., “CrewAI,” 2025. [Online]. Available: https://github.com/crewAIInc/crewAI
|
| [11] |
LlamaIndex, “LlamaIndex,” 2025. [Online]. Available: https://www.llamaindex.ai/.
|
| [12] |
Z. Ji, D. Wu, W. Jiang, P. Ma, Z. Li, Y. Gao, S. Wang, and Y. Li, “Taming various privilege escalation in LLM-based agent systems: A mandatory access control framework,” arXiv preprint arXiv: 2601.11893, 2026.
|
| [13] |
OWASP Cheat Sheets Series Team, “AI agent security cheat sheet,” 2026. [Online]. Available: https://cheatsheetseries.owasp.org/cheatsheets/AI_AgentSecurityCheatSheet.html. Accessed on: Mar. 22, 2026.
|
| [14] |
Z. Deng, W. Ma, Q.-L. Han, W. Zhou, X. Zhu, S. Wen, and Y. Xiang, “Exploring DeepSeek: A survey on advances, applications, challenges and future directions,” IEEE/CAA J. Autom. Sinica, vol. 12, no. 5, pp. 872–893, May 2025. doi: 10.1109/JAS.2025.125498
|
| [15] |
Y. Liu, Z. Chen, Y. Zhang, G. Deng, Y. Li, J. Ning, and L. Y. Zhang, “Malicious agent skills in the wild: A large-scale security empirical study,” arXiv preprint arXiv: 2602.06547, 2026.
|
| [16] |
Y. Jiang, D. Li, H. Deng, B. Ma, X. Wang, Q. Wang, and G. Yu, “SoK: Agentic skills-beyond tool use in LLM agents,” arXiv preprint arXiv: 2602.20867, 2026.
|
| [17] |
H. Zhang, J. Huang, K. Mei, Y. Yao, Z. Wang, C. Zhan, H. Wang, and Y. Zhang, “Agent security bench (ASB): Formalizing and benchmarking attacks and defenses in LLM-based agents,” in Proc. 13th Int. Conf. Learning Representations, Singapore, Singapore, 2025.
|
| [18] |
B. D. Sunil, I. Sinha, P. Maheshwari, S. Todmal, S. Mallik, and S. Mishra, “Memory poisoning attack and defense on memory based LLM-agents,” arXiv preprint arXiv: 2601.05504, 2026.
|
| [19] |
S. Dong, S. Xu, P. He, Y. Li, J. Tang, T. Liu, H. Liu, and Z. Xiang, “Memory injection attacks on LLM agents via query-only interaction,” in Proc. 39th Annu. Conf. Neural Information Processing Systems, 2025.
|
| [20] |
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection,” in Proc. 16th ACM Workshop on Artificial Intelligence and Security, Copenhagen, Denmark, 2023, pp. 79−90.
|
| [21] |
Y. Liu, W. Wang, R. Feng, Y. Zhang, G. Xu, G. Deng, Y. Li, and L. Zhang, “Agent skills in the wild: An empirical study of security vulnerabilities at scale,” arXiv preprint arXiv: 2601.10338, 2026.
|
| [22] |
A. Doshi, Y. Hong, C. Xu, E. Kang, A. Kapravelos, and C. Kästner, “Towards verifiably safe tool use for LLM agents,” in Proc. Int. Conf. Software Engineering, New Ideas and Emerging Results Track, 2026.
|
| [23] |
G. Singh and V. K. Madisetti, “MCP-secure: A runtime access control layer for privilege-aware LLM agent tooling,” IEEE Open J. Comput. Soc., vol. 7, pp. 504–513, Feb. 2026. doi: 10.1109/OJCS.2026.3664314
|
| [24] |
Z. Li, W. Li, and X. Li, “Defensible design for OpenClaw: Securing autonomous tool-invoking agents,” arXiv preprint arXiv: 2603.13151, 2026.
|
| [25] |
S. Chen, J. Piet, C. Sitawarin, and D. Wagner, “StruQ: Defending against prompt injection with structured queries,” in Proc. 34th USENIX Conf. Security Symp., Seattle, USA, 2025, Art. no. 123.
|
| [26] |
W. Ma, Y. Guo, Q.-L. Han, W. Zhou, X. Zhu, J. Xiong, S. Wen, and Y. Xiang, “Understanding agentic AI: Algorithms and infrastructure,” IEEE/CAA J. Autom. Sinica, vol. 13, no. 4, pp. 776–795, 2026. doi: 10.1109/JAS.2026.125993
|
| [27] |
S. Robertson and H. Zaragoza, “The probabilistic relevance framework: BM25 and beyond,” Found. Trends Inform. Retr., vol. 3, no. 4, pp. 333–389, Apr. 2009. doi: 10.1561/1500000019
|
| [28] |
P. Steinberger, “ClawHub,” [Online]. Available: https://clawhub.ai/.
|
| [29] |
badlogic, “Pi monorepo,” 2026. [Online]. Available: https://github.com/badlogic/pi-mono. Accessed on: Mar. 22, 2026.
|
| [30] |
IBM BeeAI, “Agent communication protocol,” 2026. [Online]. Available: https://agentcommunicationprotocol.dev. Accessed on: Mar. 22, 2026.
|
| [31] |
Autobot, “Ultra-efficient personal AI assistant powered by crystal,” [Online]. Available: https://github.com/crystal-autobot/autobot. Accessed on: Mar. 22, 2026.
|
| [32] |
Pickle-Bot, “Your own AI assistant,” 2026. [Online]. Available: https://github.com/czl9707/pickle-bot. Accessed on: Mar. 22, 2026.
|
| [33] |
Moxxy, “Your everyday AI assistant agents,” 2026. [Online]. Available: https://github.com/moxxy-ai/moxxy. Accessed on: Mar. 22, 2026.
|
| [34] |
TinyAGI, “TinyAGI is the agent teams orchestrator for one person company (fka TinyClaw),” 2026. [Online]. Available: https://github.com/TinyAGI/tinyagi. Accessed on: Mar. 22, 2026.
|
| [35] |
HermitClaw, “A tiny AI creature that lives in a folder on your computer,” [Online]. Available: https://github.com/brendanhogan/hermitclaw. Accessed on: Mar. 22, 2026.
|
| [36] |
Nanobot, “Ultra-lightweight personal AI assistant,” [Online]. Available: https://github.com/HKUDS/nanobot. Accessed on: Mar. 22, 2026.
|
| [37] |
OpenFang, “The agent operating system,” 2026. [Online]. Available: https://github.com/RightNow-AI/openfang. Accessed on: Mar. 22, 2026.
|
| [38] |
AstrBot, “Agentic IM chatbot infrastructure,” [Online]. Available: https://github.com/AstrBotDevs/AstrBot. Accessed on: Mar. 22, 2026.
|
| [39] |
ZeroClaw, “Personal AI assistant,” [Online]. Available: https://github.com/zeroclaw-labs/zeroclaw. Accessed on: Mar. 22, 2026.
|
| [40] |
SupaClaw, “Built entirely on Supabase built-in features,” [Online]. Available: https://github.com/vincenzodomina/supaclaw. Accessed on: Mar. 22, 2026.
|
| [41] |
LettaBot, “Your personal AI assistant that remembers everything,” [Online]. Available: https://github.com/letta-ai/lettabot. Accessed on: Mar. 22, 2026.
|
| [42] |
Flowly, “Your personal AI assistant, one click away,” [Online]. Available: https://www.useflowlyapp.com/en. Accessed on: Mar. 22, 2026.
|
| [43] |
DroidClaw, “An AI agent that controls your android phone,” 2026. [Online]. Available: https://github.com/unitedbyai/droidclaw. Accessed on: Mar. 22, 2026.
|
| [44] |
BabyClaw, “Claude code on a VPS, controlled from telegram,” [Online]. Available: https://github.com/yogesharc/babyclaw. Accessed on: Mar. 22, 2026.
|
| [45] |
Zclaw, “Your personal AI assistant,” 2026. [Online]. Available: https://github.com/tnm/zclaw. Accessed on: Mar. 22, 2026.
|
| [46] |
MimiClaw, “MimiClaw: Pocket AI assistant on a $5 chip,” 2026. [Online]. Available: https://github.com/memovai/mimiclaw. Accessed on: Mar. 22, 2026.
|
| [47] |
Picobot, “The AI agent that runs anywhere — even on a $5 VPS,” 2026. [Online]. Available: https://github.com/louisho5/picobot. Accessed on: Mar. 22, 2026.
|
| [48] |
PicoClaw, “PicoClaw: Ultra-efficient AI assistant in go,” 2026. [Online]. Available: https://github.com/sipeed/picoclaw. Accessed on: Mar. 22, 2026.
|
| [49] |
AngelClaw, “The multi-tenant agent operating system,” 2026. [Online]. Available: https://github.com/Abdur-rahmaanJ/angel-claw. Accessed on: Mar. 22, 2026.
|
| [50] |
SafeClaw, “The zero-cost alternative to OpenClaw,” 2026. [Online]. Available: https://github.com/princezuda/safeclaw. Accessed on: Mar. 22, 2026.
|
| [51] |
Shrew, “Shrew: Ultra-lightweight agent,” 2026. [Online]. Available: https://github.com/Masmedeam/shrew. Accessed on: Mar. 22, 2026.
|
| [52] |
IronClaw, “Your secure personal AI assistant, always on your side,” 2026. [Online]. Available: https://github.com/nearai/ironclaw. Accessed on: Mar. 22, 2026.
|
| [53] |
Moltis, “A rust-native claw you can trust,” 2026. [Online]. Available: https://github.com/moltis-org/moltis. Accessed on: Mar. 22, 2026.
|
| [54] |
N. Shapira, C. Wendler, A. Yen, G. Sarti, K. Pal, O. Floody, A. Belfki, A. Loftus, A. R. Jannali, N. Prakash, et al., “Agents of chaos,” arXiv preprint arXiv: 2602.20021, 2026.
|
| [55] |
E. Zverev, S. Abdelnabi, S. Tabesh, M. Fritz, and C. H. Lampert, “Can LLMs separate instructions from data? And what do we even mean by that?” in Proc. 13th Int. Conf. Learning Representations, Singapore, Singapore, 2025.
|
| [56] |
Q. Zhan, Z. Liang, Z. Ying, and D. Kang, “InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents,” in Proc. Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 2024, pp. 10471−10506.
|
| [57] |
X. Yang, Y. He, S. Ji, B. Hooi, and J. S. Dong, “Zombie agents: Persistent control of self-evolving LLM agents via self-reinforcing injections,” arXiv preprint arXiv: 2602.15654, 2026.
|
| [58] |
S. S. Srivastava and H. He, “MemoryGraft: Persistent compromise of LLM agents via poisoned experience retrieval,” arXiv preprint arXiv: 2512.16962, 2025.
|
| [59] |
R. Wen, H. Li, C. Xiao, and N. Zhang, “AgentSys: Secure and dynamic LLM agents through explicit hierarchical memory management,” arXiv preprint arXiv: 2602.07398, 2026.
|
| [60] |
Y. Liu, Y. Cui, and H. Zhang, “RRTL: Red teaming reasoning large language models in tool learning,” IEEE Internet Things J., vol. 13, no. 6, pp. 11868–11880, Mar. 2026. doi: 10.1109/JIOT.2025.3642164
|
| [61] |
J. Shi, Z. Yuan, G. Tie, P. Zhou, N. Z. Gong, and L. Sun, “Prompt injection attack to tool selection in LLM agents,” in Proc. 33rd Annu. Network and Distributed System Security Symp., San Diego, USA, 2026.
|
| [62] |
B. Dong, H. Feng, and Q. Wang, “Clawdrain: Exploiting tool-calling chains for stealthy token exhaustion in OpenClaw agents,” arXiv preprint arXiv: 2603.00902, 2026.
|
| [63] |
K. Zhou, Y. Zheng, Y. He, M. Xue, X. Gong, Y. Wang, and K.-Y. Lam, “Beyond max tokens: Stealthy resource amplification via tool calling chains in LLM agents,” arXiv preprint arXiv: 2601.10955, 2026.
|
| [64] |
H. Chang, E. Bao, X. Luo, and T. Yu, “Overcoming the retrieval barrier: Indirect prompt injection in the wild for LLM systems,” in Proc. ACM on Web Conf., 2026.
|
| [65] |
X. Wang, J. Bloch, Z. Shao, Y. Hu, S. Zhou, and N. Z. Gong, “Webinject: Prompt injection attack to web agents,” in Proc. Conf. Empirical Methods in Natural Language Processing, Suzhou, China, 2025, pp. 2010−2030.
|
| [66] |
E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr, “AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents,” Proc. 38th Int. Conf. Neural Information Processing Systems, Vancouver, Canada, 2024, Art. no. 2636.
|
| [67] |
F. Jia, T. Wu, X. Qin, and A. Squicciarini, “The task shield: Enforcing task alignment to defend against indirect prompt injection in LLM agents,” in Proc. 63rd Annu. Meeting Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 2025, pp. 29680−29697.
|
| [68] |
K. Zhu, X. Yang, J. Wang, W. Guo, and W. Y. Wang, “Melon: Provable defense against indirect prompt injection attacks in AI agents,” in Proc. 42nd Int. Conf. Machine Learning, Vancouver, Canada, 2025, pp. 80310−80329.
|
| [69] |
F. Liu, Z. Chen, T. Lan, H. Tan, Z. Xu, X. Li, G. Chen, Y. Meng, and H. Zhu, “Trojan’s whisper: Stealthy manipulation of OpenClaw through injected bootstrapped guidance,” arXiv preprint arXiv: 2603.19974, 2026.
|
| [70] |
Z. Guo, Z. Chen, X. Nie, J. Lin, Y. Zhou, and W. Zhang, “SkillProbe: Security auditing for emerging agent skill marketplaces via multi-agent collaboration,” arXiv preprint arXiv: 2603.21019, 2026.
|
| [71] |
C. McCauley, K. Schulz, R. Tracey, and J. Martin, “Exploring the security risks of AI assistants like OpenClaw,” [Online]. Available: https://www.hiddenlayer.com/research/exploring-the-security-risks-of-ai-assistants-like-openclaw. Accessed on: Apr. 26, 2026.
|
| [72] |
A. Oliveira, B. Tancio, D. Fiser, P. Lin, and R. Reyes, “Malicious OpenClaw skills used to distribute atomic MACOs stealer,” [Online]. Available: https://www.trendmicro.com/en_us/research/26/b/openclaw-skills-used-to-distribute-atomic-macos-stealer.html. Accessed on: Apr. 26, 2026.
|
| [73] |
X. Deng, Y. Zhang, J. Wu, J. Bai, S. Yi, Z. Zou, Y. Xiao, R. Qiu, J. Ma, J. Chen, et al., “Taming OpenClaw: Security analysis and mitigation of autonomous LLM agent threats,” arXiv preprint arXiv: 2603.11619, 2026.
|
| [74] |
X. Jin, M. Duan, Q. Lin, A. Chan, Z. Chen, J. Du, and X. Ren, “Proof-of-guardrail in AI agents and what (not) to trust from it,” arXiv preprint arXiv: 2603.05786, 2026.
|
| [75] |
H. Li, X. Liu, C. H. Chun, D. Li, N. Zhang, and C. Xiao, “DRIFT: Dynamic rule-based defense with injection isolation for securing LLM agents,” in Proc. 39th Annu. Conf. Neural Information Processing Systems, 2025.
|
| [76] |
Y. Wen, Y. Zhang, J. Lian, X. Yi, X. Xie, and D. Yang, “Contextualized privacy defense for LLM agents,” arXiv preprint arXiv: 2603.02983, 2026.
|
| [77] |
H. An, J. Zhang, T. Du, C. Zhou, Q. Li, T. Lin, and S. Ji, “IPIGuard: A novel tool dependency graph-based defense against indirect prompt injection in LLM agents,” in Proc. Conf. Empirical Methods in Natural Language Processing, Suzhou, China, 2025, pp. 1023−1039.
|
| [78] |
T. Shi, J. He, Z. Wang, H. Li, L. Wu, W. Guo, and D. Song, “ Progent: Securing AI agents with privilege control,” arXiv preprint arXiv: 2504.11703, 2025.
|
| [79] |
R. K. Sharma and D. Grossman, “AC4A: Access control for agents,” arXiv preprint arXiv: 2603.20933, 2026.
|
| [80] |
Z. Shan, J. Xin, Y. Zhang, and M. Xu, “Don’t let the claw grip your hand: A security analysis and defense framework for OpenClaw,” arXiv preprint arXiv: 2603.10387, 2026.
|
| [81] |
B. Rombaut, S. Masoumzadeh, K. Vasilevski, D. Lin, and A. E. Hassan, “Watson: A cognitive observability framework for the reasoning of LLM-powered agents,” in Proc. 40th IEEE/ACM Int. Conf. Automated Software Engineering, Seoul, Republic of Korea, 2025, pp. 739−751.
|
| [82] |
A. Kaplunovich, “Advancing LLM agents for code generation: Observability, orchestration, reliable performance,” in Proc. Int. Conf. Intelligent Computing, Communication, Networking and Services, Varna, Bulgaria, 2025, pp. 108−116.
|
| [83] |
Q. Lan, A. Kaul, and S. Jones, “Prompt injection detection in LLM integrated applications,” Int. J. Network Dynam. Intell., vol. 4, no. 2, Art. no. 100013, 2025. doi: 10.53941/ijndi.2025.100013
|
| [84] |
Y. Wang, F. Xu, Z. Lin, G. He, Y. Huang, H. Gao, Z. Niu, S. Lian, and Z. Liu, “From assistant to double agent: Formalizing and benchmarking attacks on OpenClaw for personalized local AI agent,” arXiv preprint arXiv: 2602.08412, 2026.
|
| [85] |
Y. Zhao, B. Yuan, J. Huang, H. Yuan, Z. Yu, L. Hu, H. Xu, A. Shankarampeta, Z. Huang, W. Ni, et al., “AMA-Bench: Evaluating long-horizon memory for agentic applications,” in Proc. ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems, 2026.
|
| [86] |
A. AlSayyad, K. Y. Huang, and R. Pal, “AgentTrace: A structured logging framework for agent system observability,” in Proc. LLM-based Multi-Agent Systems: Towards Responsible, Reliable, and Scalable Agentic Systems, 2026.
|