文章总结: 汇总了AI安全和agent安全的所有外刊文章,包括威胁框架与标准、攻击研究、防御策略等内容,提供了丰富的参考资料。
综合评分: 85
文章分类: AI安全,代码审计,漏洞分析,威胁情报,恶意软件
AI安全、agent安全所有外刊文章汇总(全网最全)
原创
小猫信安
小猫信安
小猫信安
2026年10月3日 07:30
天津
在小说阅读器读本章
去阅读
在公众号小说中沉浸阅读
温馨提示:文章字体最小号阅读最佳
请善用curl+f搜索
目录
威胁框架与标准综述与系统化研究攻击研究通过工具进行提示词注入工具投毒与供应链权限提升与过度授权数据外泄与隐私间接提示词注入跨插件攻击针对 Agent 的后门攻击Agent 欺骗与操控越狱与护栏绕过防御研究权限与访问控制运行时监控与沙箱输入/输出校验形式化验证与分析评估与红队测试基准测试与数据集工具与框架Agent Skill 规范行业报告与博客文章
威胁框架与标准
- “`
OWASP Agentic AI Threats and Mitigations(OWASP 代理式 AI 威胁与缓解): https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/OWASP Top 10 for LLM Applications(OWASP LLM 应用前 10 大风险): https://owasp.org/www-project-top-10-for-large-language-model-applications/OWASP Agentic Skills Top 10 (AST10)(OWASP Agentic Skills 前 10 大风险): https://owasp.org/www-project-agentic-skills-top-10/MITRE ATLAS™(ATLAS 威胁地图): https://atlas.mitre.org/NIST AI Risk Management Framework(NIST AI 风险管理框架): https://www.nist.gov/artificial-intelligence/ai-risk-management-frameworkNIST SP 800-218A: Secure Software Development for AI(NIST AI 安全软件开发): https://csrc.nist.gov/pubs/sp/800/218/a/finalEU AI Act(欧盟 AI 法案): https://artificialintelligenceact.eu/Anthropic Responsible Scaling Policy(Anthropic 负责任扩展政策): https://www.anthropic.com/index/anthropics-responsible-scaling-policyIETF draft-klrc-aiagent-auth-01: AI Agent Authentication and Authorization(AI Agent 身份认证与授权草案): https://datatracker.ietf.org/doc/draft-klrc-aiagent-auth/IETF draft-niyikiza-oauth-attenuating-agent-tokens-00: Attenuating Authorization Tokens for Agentic Delegation Chains(Agent 委托链授权令牌削弱草案): https://datatracker.ietf.org/doc/draft-niyikiza-oauth-attenuating-agent-tokens-00/Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act(网络安全 AI Agent 如何突破 EU Cyber Resilience Act): https://arxiv.org/abs/2607.07109综述与系统化研究Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges(连接各类 Agentic AI 安全研究点): https://arxiv.org/abs/2609.23894Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems(可信 Agentic AI:失效模式、缓解策略与生命周期框架): https://arxiv.org/abs/2609.22712Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation(攻击成功率并非数字:Agentic AI 安全评估的有效性研究): https://arxiv.org/abs/2609.25173A Survey on LLM-based Autonomous Agents: Common Attacks and Defenses(LLM 自主 Agent 攻击与防御综述): https://arxiv.org/abs/2402.09283Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents(Agent 安全 benchmark): https://arxiv.org/abs/2410.02644Security of AI Agents(AI Agent 安全): https://arxiv.org/abs/2406.08689Not All Agents Are Created Equal: A Survey on Software-use Agent Security(并非所有 Agent 都一样:软件使用型 Agent 安全综述): https://arxiv.org/abs/2502.02761A Survey on the Honesty of Large Language Models(大语言模型诚实性综述): https://arxiv.org/abs/2409.18786A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models(大语言模型越狱攻击与防御综述): https://arxiv.org/abs/2402.13457Prompt Injection Attacks and Defenses in LLM-Integrated Applications(LLM 集成应用中的 Prompt Injection 攻击与防御): https://arxiv.org/abs/2310.12815The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies(LLM Agent 的安全与隐私综述与案例研究): https://arxiv.org/abs/2407.19354Self-Evolving Agents: A Survey(自进化 Agent 调研): https://arxiv.org/abs/2504.01641Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies(自进化 LLM Agent 系统安全): https://arxiv.org/abs/2606.23075From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents(从思考者到社会:分层自治 AI Agent 安全): https://arxiv.org/abs/2603.07496Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes(Agentic AI 故障分类): https://arxiv.org/abs/2603.06847Security Considerations for Multi-agent Systems(多 Agent 系统安全考量): https://arxiv.org/abs/2603.09002The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey(Agentic AI 攻击与防御全景综述): https://arxiv.org/abs/2603.11088Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats(OpenClaw 安全分析与缓解): https://arxiv.org/abs/2603.11619OpenClaw as Language Infrastructure: A Case-Centered Survey of a Public Agent Ecosystem in the Wild(OpenClaw 作为语言基础设施的案例调查): https://www.preprints.org/manuscript/202603.1060AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations(企业网络运营中的多 Agent 安全框架): https://arxiv.org/abs/2603.09134MCP-in-SoS: Risk Assessment Framework for Open-Source MCP Servers(开源 MCP Server 的系统风险评估框架): https://arxiv.org/abs/2603.10194SoK: The Attack Surface of Agentic AI — Tools, and Autonomy(Agentic AI 攻击面综述): https://arxiv.org/abs/2603.22928Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation(面向安全 LLM Agent 的威胁面、攻击、防御与评估): https://arxiv.org/abs/2606.10749Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems(数据 Agent 遭遇攻击): https://arxiv.org/abs/2606.08661Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents(知道太多的 Agent:面向数据中心的隐私调查): https://arxiv.org/abs/2606.26627Security Engineering of OpenClaw: Analyzing Attack Surface Expansion and Trust-Boundary Violations(OpenClaw 的安全工程分析): https://arxiv.org/abs/2606.15008LLM Agents Security Duality: A Comprehensive Survey of Self-Security and Empowered Cybersecurity(LLM Agent 安全二象性综述): https://arxiv.org/abs/2606.28450Agent Security Meets Regulatory Reality: A Practitioner Systematization of Autonomous-Agent Threats and Controls in Regulated Financial Systems(Agent 安全与监管现实结合): https://arxiv.org/abs/2606.29142Security and Privacy in Agentic AI: Grand Challenges and Future Directions(Agentic AI 安全与隐私的重大挑战与未来方向): https://arxiv.org/abs/2607.06608Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents(信任但验证?揭示自治编码 Agent 的安全债务): https://arxiv.org/abs/2607.12428Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation(Agent Skill 安全:威胁模型、攻击、防御与评估): https://arxiv.org/abs/2607.13987The Ethics of Autonomous AI Agents for Offensive Security(用于攻击性安全的自治 AI Agent 的伦理问题): https://arxiv.org/abs/2607.20255The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI(Chronos 漏洞:Agentic AI 中时间持久化与记忆欺骗分类): https://arxiv.org/abs/2607.19433Engineering Trustworthy Agentic AI for Critical Systems(关键系统中的可信 Agentic AI 工程): https://arxiv.org/abs/2607.18548Agent Security Needs Redefinition through a Holistic Framework(Agent 安全需要整体框架重定义): https://arxiv.org/abs/2607.22024An Empirical Study of Model Context Protocol Applications(MCP 应用实证研究): https://arxiv.org/abs/2607.25635Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation(基于世界模型的具身 AI 安全): https://arxiv.org/abs/2607.28226From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems(从单体到群体:多 Agent Web 系统攻击面演进研究): https://arxiv.org/abs/2608.00202The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents(没有 CVE 的漏洞:管理 AI 编码 Agent 的授权与职责差异): https://arxiv.org/abs/2608.05884On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models(理解、识别与缓解 Agentic LLM 漏洞): https://arxiv.org/abs/2608.10530When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling(Agent 参与 Web3 时的攻击面研究): https://arxiv.org/abs/2608.17275A2ABreak: Systematic Security Analysis of the A2A Protocol(A2A 协议系统安全分析): https://arxiv.org/abs/2609.10871When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems(测试通过掩盖漏洞:Agentic 系统沉默失败实证研究): https://arxiv.org/abs/2609.10548Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges(可信 Agentic AI:威胁、架构与开放挑战的全面综述): https://arxiv.org/abs/2609.13731Authorization Architectures for Tool-Using AI Agents(工具使用型 AI Agent 的授权架构): https://arxiv.org/abs/2609.15906SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration(Agentic AI 时代越狱再思考): https://arxiv.org/abs/2609.12413After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind(聚会之后:治理大规模 Agent-Skill 生态遗留问题): https://arxiv.org/abs/2609.17274SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes(交易 Agent 还是市场崩盘者?学术金融 LLM 交易方案的稳健性与安全失效分析): https://arxiv.org/abs/2609.19705Characterizing Network Centralization and Observability in the Remote MCP Ecosystem(远程 MCP 生态中的网络中心化与可观测性): https://arxiv.org/abs/2609.19100PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs(PentestChain:低成本的 MCP 驱动自动渗透测试框架): https://arxiv.org/abs/2609.18120Toward Secure AI-Powered Penetration Testing Agents: Security Threats, Guardrails, and Architectural Perspectives(安全 AI 渗透测试 Agent:威胁、护栏与架构视角): https://arxiv.org/abs/2609.16694
---
## 攻击研究
通过工具进行提示词注入Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning(攀山:利用课程式强化学习进行前沿模型 Prompt Injection 红队测试): https://arxiv.org/abs/2609.33628Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection(不是你签约的那回事:真实世界 LLM 集成应用中的间接 Prompt Injection): https://arxiv.org/abs/2302.12173Inject My PDF: Prompt Injection for Your Resume(注入我的 PDF:简历中的 Prompt Injection): https://arxiv.org/abs/2403.14381InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents(InjecAgent:工具集成 LLM Agent 中间接 Prompt Injection 标准评测): https://arxiv.org/abs/2403.02691Automatic and Universal Prompt Injection Attacks against Large Language Models(针对大语言模型的自动化通用 Prompt Injection 攻击): https://arxiv.org/abs/2403.04957WIPI: A New Web Threat for LLM-Driven AI Agents(WIPI:LLM 驱动 AI Agent 的新型 Web 威胁): https://arxiv.org/abs/2402.16965GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines(GitInject:AI 驱动 CI/CD 流水线中的真实 Prompt Injection 攻击): https://arxiv.org/abs/2606.09935Prompt Injection as Role Confusion(Prompt Injection 作为角色混淆): https://arxiv.org/abs/2603.12277On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models(共享嵌入序列模型中指令与数据的不可分离性): https://arxiv.org/abs/2606.27567They’ll Verify. They Just Won’t Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface(他们会核验,但不会行动:权限框架与洗白代码如何把可信 Agentic CI/CD 变成攻击面): https://arxiv.org/abs/2607.19267Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents(搭乘感知之路:对多模态 LLM Agent 的隐蔽同步音频 Prompt Injection): https://arxiv.org/abs/2607.28165The Anatomy of a Prompt Injection: A Component Model for Structured Analysis(Prompt Injection 的剖析:结构化分析组件模型): https://arxiv.org/abs/2608.07808From Prompt Injection to Web Exploitation: Revisiting Classic Vulnerabilities in LLM-Integrated Applications(从 Prompt Injection 到 Web 漏洞利用:重温 LLM 集成应用中的经典漏洞): https://arxiv.org/abs/2608.10281Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents(控制 token 注入抑制思维链并绕过推理监督): https://arxiv.org/abs/2609.27542Decision Hijacking: Prompt Injection Attacks on Jev’s Typed Probabilistic Decisions(决策劫持:针对 Jev 类型化概率决策的 Prompt Injection 攻击): https://arxiv.org/abs/2609.28613工具投毒与供应链ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP(ShareLock:针对 MCP 的隐蔽多工具阈值投毒攻击): https://arxiv.org/abs/2606.27027PhantomSkill: Malicious Code Injection in Agent Skill Ecosystems(PhantomSkill:Agent Skill 生态中的恶意代码注入): https://arxiv.org/abs/2606.19191Dynamic Malicious Skills in Agentic AI(Agentic AI 中的动态恶意 Skill): https://arxiv.org/abs/2606.16287Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents(收敛性绕行劫持:基于 Skill 的 LLM Agent 中的任务保留型资源放大): https://arxiv.org/abs/2608.12273Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners(看见不等于筛查:Agent Skill 扫描器中的多模态隐藏指令攻击): https://arxiv.org/abs/2606.18198″Do Not Mention This to the User”: Detecting and Understanding Malicious Agent Skills in the Wild(不要告诉用户:检测与理解真实世界中的恶意 Agent Skill): https://arxiv.org/abs/2602.06547Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills(表面无害,实则有害:通过中性提示实现隐蔽幻觉引导): https://arxiv.org/abs/2605.29354Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search(MCP 的隐形威胁:通过树搜索生成隐蔽注入载荷): https://arxiv.org/abs/2603.24203Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning(MCP 威胁建模与工具投毒导致的 Prompt Injection 漏洞分析): https://arxiv.org/abs/2603.22489Are AI-assisted Development Tools Immune to Prompt Injection?(AI 辅助开发工具是否免疫于 Prompt Injection?): https://arxiv.org/abs/2603.21642Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks(Skill-Inject:测量 Agent 对 Skill 文件攻击的脆弱性): https://arxiv.org/abs/2602.20156ToolSword: Unveiling Safety Issues of LLMs in Tool Learning Across Three Stages(ToolSword:揭示 LLM 工具学习在三阶段中的安全问题): https://arxiv.org/abs/2402.10753Compromising Agents via MCP(通过 MCP 攻破 Agent): https://arxiv.org/abs/2504.03767Osmosis Distillation: Model Hijacking with the Fewest Samples(渗透蒸馏:用最少样本劫持模型): https://arxiv.org/abs/2603.04859Personality Self-Replicators(人格自复制体): https://arxiv.org/abs/2603.XXXXXPoisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents(毒化的操作手册:揭示知识投毒对 AI 安全 Agent 的影响): https://arxiv.org/abs/2606.24402Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains(Skill 不只是孤岛:测量 Agent Skill 供应链中的依赖与风险): https://arxiv.org/abs/2607.01136KidnapRAG: A Black-Box Attack for Hijacking Reasoning in Agentic Retrieval-Augmented Generation Systems(KidnapRAG:劫持 Agentic RAG 推理的黑盒攻击): https://arxiv.org/abs/2607.00422Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol(MCP 中工具元数据载荷的 Unicode TAG block 隐蔽): https://arxiv.org/abs/2607.05744Beware of Agentic Botnets: Scalable Untargeted Promptware Attacks via Universal and Transferable Adversarial HalluSquatting(注意 Agentic Botnet:泛化式对抗性误报促成的大规模无目标 Promptware 攻击): https://arxiv.org/abs/2607.07433Skills That Don’t Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents(并不存在的 Skill:LLM Agent 中幻觉 Skill 推荐的大规模研究): https://arxiv.org/abs/2607.12340Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware(隐蔽与引爆:绕过扫描器与动态检测 Agent Skill 恶意软件): https://arxiv.org/abs/2607.02357Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents(安装完成,现在你已被攻破:利用 setup instructions 攻击 AI 编码 Agent): https://arxiv.org/abs/2607.15143When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems(当经验变成指令:自进化 Agent Skill 系统中的轨迹投毒): https://arxiv.org/abs/2608.05563Towards a Risk Assessment of Malicious Skill Files in Coding Agents(恶意 Skill 文件风险评估): https://arxiv.org/abs/2608.05223SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains(SynChain:让计算机使用 Agent 自己构建攻击链): https://arxiv.org/abs/2608.06862ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners(ColluSkill:跨 Skill 组合规避扫描器的对抗性攻击): https://arxiv.org/abs/2608.09732CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills(CompoSkill:由可通过扫描器的单个 Skill 组合成攻击链): https://arxiv.org/abs/2608.16246Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem(同名不同服务器:MCP 生态中静默漂移的安全普查): https://arxiv.org/abs/2609.14119Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations(扫描 Harness:AI 编码 Agent 配置中的供应链缺陷实证研究): https://arxiv.org/abs/2609.07360SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents(SkillJect:自动化 Skill-based Prompt Injection): https://arxiv.org/abs/2602.14211Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents(Poise:位置感知的一条指令 Skill 注入): https://arxiv.org/abs/2606.07943Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems(针对 LLM 编码 Agent Skill 生态的供应链投毒攻击): https://arxiv.org/abs/2604.03081SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents(SkillBloat:通过 Skill 注入放大 token 的攻击): https://arxiv.org/abs/2608.21929Chaining Skills to Hijack LLM Agents(串联 Skill 劫持 LLM Agent): https://arxiv.org/abs/2610.01564Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents(持续计费状态:工具调用 LLM Agent 的拒绝钱包攻击与防御): https://arxiv.org/abs/2609.28585A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem(A2M:MCP 生态中的轨迹优化型 Agent 劫持): https://arxiv.org/abs/2609.26761Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents(藏在显眼处:将表象与执行解耦,以掩盖 Skill 投毒): https://arxiv.org/abs/2609.39352Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents(Pretext:绕过 AI Agent 恶意 Skill 检测框架): https://arxiv.org/abs/2609.39607Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents(Agent 能信任它们的 Skill 吗?揭示基于 Skill 的 LLM Agent 中不安全信任链): https://arxiv.org/abs/2609.39065MCP Security Notification: Tool Poisoning Attacks(MCP 安全通知:工具投毒攻击): https://modelcontextprotocol.io/specification/2025-03-26/securityInvariant Labs: MCP Security Research(Invariant Labs:MCP 安全研究): https://invariantlabs.ai/blog/mcp-security权限提升与过度授权When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents(同意超出上下文:长生命周期 Agent 中的剩余权限重放): https://arxiv.org/abs/2609.33910Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents(小心你的 Agent!调查 LLM Agent 的后门威胁): https://arxiv.org/abs/2402.11208Pandora’s White-Box: Precise Training Data Detection and Extraction in Large Language Models(潘多拉的白盒:精确检测和提取大语言模型训练数据): https://arxiv.org/abs/2402.17012R-Judge: Benchmarking Safety Risk Awareness for LLM Agents(R-Judge:LLM Agent 安全风险感知基准): https://arxiv.org/abs/2401.10019TrustAgent: Towards Safe and Trustworthy LLM-based Agents(TrustAgent:安全可信的 LLM Agent): https://arxiv.org/abs/2402.01586FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion(FragFuse:通过基于记忆的查询碎片化与融合绕过访问控制): https://arxiv.org/abs/2606.15609Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards(后训练 Linux 提权 LLM Agent): https://arxiv.org/abs/2603.17673Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing(可信凭据,不可信行为:高性能计算中的 LLM-Agent 安全基准): https://arxiv.org/abs/2607.18485When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents(更低权限就够了:调查 LLM Agent 中的过度授权工具选择): https://arxiv.org/abs/2606.20023(A)I Sees What You Don’t: Exploiting New Attack Surfaces in Third-Party Mobile Agents((A)I 看到了你没看见的东西:利用第三方移动 Agent 的新攻击面): https://arxiv.org/abs/2607.00333″Allow” to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents(“允许”即达成,悄悄超权限:移动 GUI Agent 中任务完成驱动弹窗授权的代价): https://arxiv.org/abs/2608.04755The Missing Boundary: How Autonomous Agents Lose Control(缺失边界:自治 Agent 如何失控): https://arxiv.org/abs/2609.11024Loopjacking: Hijacking Human-in-the-Loop Approval(Loopjacking:劫持人类审批): https://arxiv.org/abs/2609.21081OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents(OverAct:测量与缓解 LLM 工具调用 Agent 的主动过度授权): https://arxiv.org/abs/2610.01508
###
### 数据外泄与隐私
* ```
API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Model(API 密钥不应成为 LLM 词表中的 token): https://arxiv.org/abs/2609.XXXXXAgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows(AgentSCOPE:评估 Agentic 工作流中的上下文隐私): https://arxiv.org/abs/2603.04902Silent Egress: LLM-Driven Data Exfiltration via Steganographic Channels(静默外泄:利用隐写通道进行 LLM 驱动的数据外流): https://arxiv.org/abs/2502.XXXXXPrivacy Risks of General-Purpose AI Systems: A Foundation for Investigating Practitioner Perspectives(通用 AI 系统的隐私风险): https://arxiv.org/abs/2407.02027IMMACULATE: A Framework for Analyzing Information Exposure in Agent-Based Systems(IMMACULATE:Agent 系统中的信息暴露分析框架): https://arxiv.org/abs/2502.XXXXXAgentRaft: Automated Detection of Data Over-Exposure in LLM Agents(AgentRaft:自动检测 LLM Agent 中的数据过度暴露): https://arxiv.org/abs/2603.07557You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents(你叫我做的:测量指令文本诱发的 LLM Agent 数据泄漏): https://arxiv.org/abs/2603.11862An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios(真实场景下工具使用 LLM Agent 的数据泄漏风险评估): https://arxiv.org/abs/2606.17114Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs(生成式 AI Agent 中的差分隐私): https://arxiv.org/abs/2603.17902Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?(能力强但粗心:计算机使用 Agent 是否遵循上下文完整性?): https://arxiv.org/abs/2606.23189ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents(ToolPrivacyBench:工具使用 LLM Agent 的目的绑定隐私基准): https://arxiv.org/abs/2606.28061What Happens Locally, Leaks Globally: Detecting Privacy Leakage Risks in MCP Servers(本地发生什么,全球泄露什么:检测 MCP Server 中的隐私泄漏风险): https://arxiv.org/abs/2606.21338Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies(Agentic 协商中的行为隐私泄漏): https://arxiv.org/abs/2607.06815PiSAs: Benchmarking Contextual Integrity in Multi-User Agentic Systems(PiSAs:多用户 Agentic 系统中的上下文完整性基准): https://arxiv.org/abs/2607.05318Data Leakage Prevention in Agentic Applications via Preemptive Hardening(通过预防性加固防止 Agentic 应用中的数据泄漏): https://arxiv.org/abs/2607.18847When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills(当 Agent 学会“像你”:Persona Skill 中的隐私泄漏和冒充风险基准): https://arxiv.org/abs/2608.03700Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills(行为技能重建:从 LLM Agent Skill 中还原隐藏功能): https://arxiv.org/abs/2608.04192SkillWatermark: An Embedded Skill Watermark of Progressive Privacy Inference via Benign Prompts(SkillWatermark:通过良性提示嵌入渐进式隐私推断水印): https://arxiv.org/abs/2608.16026Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control(混淆模型,控制流向:用信息流控制理解与缓解 LLM Agent 隐私泄漏): https://arxiv.org/abs/2609.14003AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing(AgentLeak:超越 Skill 窃取,将更强 LLM Agent 能力克隆到更弱 Agent 上): https://arxiv.org/abs/2609.07131Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines(测量并利用 LLM 工具调用流水线中的隐式信任): https://arxiv.org/abs/2609.18217CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents(CIPL:LLM Agent 中可恢复隐私泄漏的通道感知框架): https://arxiv.org/abs/2609.21686Codetta: High-Capacity, Keyless, and Undetectable Multi-Agent Collusion(Codetta:高容量、无密钥且难察觉的多 Agent 串通): https://arxiv.org/abs/2609.28900The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching(无辜信使:通过合法的 LLM Web Fetch 进行隐蔽外泄): https://arxiv.org/abs/2610.01768间接提示词注入CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents(CoDeL:面向 LLM Agent 的协同进化防御): https://arxiv.org/abs/2609.34463CodeSentinel: A Three-Layer Defense Against Indirect Prompt Injection in Code Contexts(CodeSentinel:代码上下文中的三层防御): https://arxiv.org/abs/2606.19235Origin Is All You Need: Provenance-Aware Transformers for Structural Trust-Boundary Separation(来源就是一切:面向结构化信任边界分离的来源感知 Transformer): https://arxiv.org/abs/2609.21088Agent Data Injection Attacks are Realistic Threats to AI Agents(Agent 数据注入攻击是现实威胁): https://arxiv.org/abs/2607.05120Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents(隐形墨水威胁:计算机使用 Agent 中隐藏的对抗目标): https://arxiv.org/abs/2608.02018Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure(你的 Agentic LLM 会隐秘编码间接 Prompt Injection 暴露信号): https://arxiv.org/abs/2608.02657DualView: Preventing Indirect Prompt Injection in Personal AI Agents(DualView:防止个人 AI Agent 中的间接 Prompt Injection): https://arxiv.org/abs/2607.03821Prismata: Confining Cross-Site Prompt Injection in Web Agents(Prismata:限制 Web Agent 中跨站 Prompt Injection): https://arxiv.org/abs/2607.08147AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations(AttriGuard:通过工具调用因果归因防御 LLM Agent 间接 Prompt Injection): https://arxiv.org/abs/2603.10749ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents(ActGuard:在执行前审计 LLM Agent 的动作): https://arxiv.org/abs/2609.14987IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection(IPI-proxy:用于 Web 浏览 Agent 间接 Prompt Injection 红队测试的拦截代理): https://arxiv.org/abs/2605.11868How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation(我们能多大程度信任 LLM 搜索 Agent?测量其对网页内容操纵的信任脆弱性): https://arxiv.org/abs/2606.16821ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents(ChatInject:利用聊天模板进行 LLM Agent Prompt Injection): https://arxiv.org/abs/2509.22830Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents(拆除爆炸式提示:理解并防止 LLM Agent 中基于 trigger 的 Prompt Injection): https://arxiv.org/abs/2609.22510Adaptive Attacks and Defenses Against Indirect Prompt Injection(针对间接 Prompt Injection 的自适应攻击与防御): https://arxiv.org/abs/2408.XXXXXHouYi: A Black-box Prompt Injection Attack on LLM-integrated Applications(HouYi:针对 LLM 集成应用的黑盒 Prompt Injection 攻击): https://arxiv.org/abs/2306.05499DMAST: Dual-Modality Multi-Stage Adversarial Safety Training(DMAST:双模态多阶段对抗性安全训练): https://arxiv.org/abs/2603.04364LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web Agents(LoginTrap:发现 LLM Web Agent 中与任务无关的钓鱼式间接 Prompt Injection): https://arxiv.org/abs/2608.04741Breadcrumbing Search Agents(面包屑式搜索 Agent): https://arxiv.org/abs/2608.04565Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots(用一张纸劫持机器人:VLM 控制机器人中的物理 Prompt Injection 系统研究): https://arxiv.org/abs/2608.05715StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection(StepJack:衡量计算机使用 Agent 对多步间接 Prompt Injection 的安全性): https://arxiv.org/abs/2608.06477Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection(不是 A11y:Android 无障碍功能如何让移动 AI Agent 面临间接 Prompt Injection): https://arxiv.org/abs/2608.08939Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection(迈向元认知一击式间接 Prompt Injection): https://arxiv.org/abs/2608.08795CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering(CounterSteer:用 activation steering 抑制间接 Prompt Injection): https://arxiv.org/abs/2609.36570Agent 欺骗与操控Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents(你的 Agent 站在哪一边?LLM Agent 的多方主体忠诚度问题): https://arxiv.org/abs/2606.30383Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction(劫持 Agent 记忆:通过对话交互进行隐蔽木马攻击): https://arxiv.org/abs/2605.29960When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents(当 Agent 记住却不说:持久化个人 Agent 中的隐蔽记忆注入): https://arxiv.org/abs/2607.05189Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses(你的 Agent 记忆不属于它:LLM Agent 记忆中的伪造推理攻击): https://arxiv.org/abs/2607.05029Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems(思想病毒:多 Agent 系统中的潜意识提示式传播性错误对齐): https://arxiv.org/abs/2603.00131FlowSteer: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems(FlowSteer:仅提示词驱动的工作流偏转暴露多 Agent LLM 系统的规划期漏洞): https://arxiv.org/abs/2605.11514Intentional Deception as Controllable Capability in LLM Agents(故意欺骗作为 LLM Agent 可控能力): https://arxiv.org/abs/2603.07848Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem(信任less Agent 能被信任吗?ERC-8004 去中心化 AI Agent 生态实证研究): https://arxiv.org/abs/2606.26028When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents(当 Agent 记得太多:LLM Agent 的记忆投毒攻击): https://arxiv.org/abs/2607.06595MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents(MemPoison:揭示 LLM Agent 中持续记忆威胁与结构盲区): https://arxiv.org/abs/2607.14651Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems(坏记忆:评估 Agentic 系统中记忆带来的 Prompt Injection 风险): https://arxiv.org/abs/2607.14611Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents(Agent 会梦到错误记忆吗?多模态 AI Agent 长期记忆的黑盒视觉攻击): https://arxiv.org/abs/2607.15657Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw(Salami Attack:针对 OpenClaw 的隐蔽串谋记忆投毒): https://arxiv.org/abs/2608.01637MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents(MAFIA:仅通过查询实现的记忆攻击): https://arxiv.org/abs/2608.03844When Coordination Becomes a Threat: Communication Attacks in LLM-Controlled Multi-Robot Systems(当协调变成威胁:LLM 控制的多机器人系统中的通信攻击): https://arxiv.org/abs/2608.06830Indirect tipping: a social attack surface in AI agent populations(间接诱导:AI Agent 群体中的社会攻击面): https://arxiv.org/abs/2609.25194
跨插件攻击
When LLM-based Code Generation Meets the Software Supply Chain(LLM 代码生成遇到软件供应链): https://arxiv.org/abs/2405.XXXXXShadow API: Covert Data Exfiltration via LLM-Mediated API Interactions(Shadow API:通过 LLM 中介 API 交互进行隐蔽数据外泄): https://arxiv.org/abs/2603.01919AgentSkillOS: Towards Secure and Composable Agent Skill Operating Systems(AgentSkillOS:面向安全且可组合的 Agent Skill 操作系统): https://arxiv.org/abs/2603.02176Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems(单独无害,组合有害:Agent Skill 生态中的安全风险): https://arxiv.org/abs/2606.15242MOSAIC: Knowledge-Guided CLI Command Composition Attack in LLM Coding Agents(MOSAIC:LLM 编码 Agent 中知识引导的 CLI 命令组合攻击): https://arxiv.org/abs/2607.02857
针对 Agent 的后门攻击
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents(SlowBA:针对 VLM-based GUI Agent 的效率型后门攻击): https://arxiv.org/abs/2603.08316SkillJack: Persistent Skill Backdoors in Self-Evolving Agents(SkillJack:自进化 Agent 中的持久性 Skill 后门): https://arxiv.org/abs/2608.03509
越狱与护栏绕过
- “`
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs(Claudini:自动研究发现 LLM 的 SOTA 对抗攻击算法): https://arxiv.org/abs/2603.24511
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails(从盾牌到目标:LLM Agent 护栏的 DoS 攻击): https://arxiv.org/abs/2606.14517Jailbreaking ChatGPT via Prompt Engineering(通过 Prompt Engineering 越狱 ChatGPT): https://arxiv.org/abs/2305.13860MasterKey: Automated Jailbreaking of Large Language Model Chatbots(MasterKey:自动化越狱大语言模型聊天机器人): https://arxiv.org/abs/2307.08715PentestGPT: An LLM-empowered Automatic Penetration Testing Tool(PentestGPT:LLM 驱动的自动渗透测试工具): https://arxiv.org/abs/2308.06782Self-Fulfilling Misalignment in AI Control(AI 控制中的自我实现误对齐): https://arxiv.org/abs/2603.XXXXXReasoning Models Struggle to Control Their Chains of Thought(推理模型难以控制思维链): https://arxiv.org/abs/2603.XXXXXCRAFT: Contrastive Reasoning Alignment — Reinforcement Learning from Hidden Representations(CRAFT:对比推理对齐—基于隐藏表示的强化学习): https://arxiv.org/abs/2603.17305It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents(它对医生撒谎去买毒药:量化真实世界中 Phone-use Agent 的误用): https://arxiv.org/abs/2606.27944Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring(拒答之后:通过行为监控确定护栏激活): https://arxiv.org/abs/2607.02121Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion(对齐是局部的:用户侧说服下 GUI Agent 的成对诊断): https://arxiv.org/abs/2607.29199
---
## 防御研究
权限与访问控制ActGov: Governing LLM Agent Actions via Policy-Constrained Validation(ActGov:通过策略约束校验治理 LLM Agent 行为): https://arxiv.org/abs/2609.24446Zero-Trust Authorization and Discovery for Enterprise MCP(企业 MCP 的零信任授权与发现): https://arxiv.org/abs/2609.22573Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems(超越单模型注入:多 Agent 系统中 Prompt Injection 的威胁模型与防御架构): https://arxiv.org/abs/2609.22949Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing(语言防火墙:多 Agent 系统路由中的几何防御): https://arxiv.org/abs/2606.30555PAuth: Precise Task-Scoped Authorization For Agents(PAuth:面向 Agent 的精确任务作用域授权): https://arxiv.org/abs/2603.17170Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies(使用动态实时组合策略保护多工具 AI Agent 链): https://arxiv.org/abs/2607.03423Rethinking Agent Security as a Networking Problem(把 Agent 安全重新视为网络问题): https://arxiv.org/abs/2608.12172Twin Agent: Context Residual Compression for Privilege Separated Agents(Twin Agent:权限隔离 Agent 的上下文残差压缩): https://arxiv.org/abs/2607.19595TrustAgent: Towards Safe and Trustworthy LLM-based Agents(TrustAgent:安全可信的 LLM Agent): https://arxiv.org/abs/2402.01586Autoformalization of Agent Instructions into Policy-as-Code(将 Agent 指令自动形式化为 Policy-as-Code): https://arxiv.org/abs/2606.26649ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents(ToolFence:安全工具使用 LLM Agent 的细粒度授权): https://arxiv.org/abs/2609.37196ActionGuard: Tool Call Authorization under Poisoned Skills(ActionGuard:在恶意 Skill 下对工具调用进行授权): https://arxiv.org/abs/2609.39450Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems(治理行为,而非治理 Agent:制度认证作为自治 AI 系统治理模型): https://arxiv.org/abs/2606.26298A Deterministic Control Plane for LLM Coding Agents(LLM 编码 Agent 的确定性控制平面): https://arxiv.org/abs/2606.26924A Dual-Helix Governance Approach for Reliable Agentic AI(可靠 Agentic AI 的双螺旋治理方法): https://arxiv.org/abs/2603.04390Talk Freely, Execute Strictly: Schema-Gated Agentic AI(自由表达,严格执行:Schema-Gated Agentic AI): https://arxiv.org/abs/2603.XXXXXESAA-Security: Event-Sourced Architecture for Agent-Assisted Security Audits(ESAA-Security:Agent 辅助安全审计的事件溯源架构): https://arxiv.org/abs/2603.XXXXXCaging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare(约束 Agent:医疗 AI 的零信任安全架构): https://arxiv.org/abs/2603.17419Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes(主权执行中间件:在 Agentic 控制平面中强制证书绑定权限): https://arxiv.org/abs/2606.20520The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems(不可熄灭的安全内核:AI Agent 等可执行系统的执行时对齐): https://arxiv.org/abs/2606.26057A First Measurement Study on Authentication Security in Real-World Remote MCP Servers(真实世界远程 MCP Server 的认证安全测量研究): https://arxiv.org/abs/2605.22333CmdNeedle: Measuring the Incompleteness of Command Denylists for AI Agents(CmdNeedle:测量 AI Agent 命令拒绝列表的不足): https://arxiv.org/abs/2606.15549SecureClaw: Clawing Back Control of LLM Agents(SecureClaw:夺回 LLM Agent 控制权): https://arxiv.org/abs/2606.09549SessionBound: Turning Enterprise Task Approval into Budgeted Database Sessions(SessionBound:将企业任务审批转化为预算化数据库 session): https://arxiv.org/abs/2607.00751Janus: A Playground for User-Involved Agentic Permission Management(Janus:用户参与式 Agentic 权限管理实验平台): https://arxiv.org/abs/2607.01510aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents(aiAuthZ:面向 AI Agent 的离主机身份绑定授权): https://arxiv.org/abs/2607.05518From Neural Intent to Cryptographic Authorization: Governing Agentic Workflows(从神经意图到密码学授权:治理 Agentic 工作流): https://arxiv.org/abs/2607.15596Steerability via Constraints: A Substrate for Scalable Oversight of Coding Agents(约束驱动的可控性:编码 Agent 可扩展监督基础): https://arxiv.org/abs/2607.02389Context-to-Execution Integrity for LLM Agents(LLM Agent 的上下文到执行完整性): https://arxiv.org/abs/2607.06000Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents(少想,多核验:确定性门禁修复工具调用 LLM Agent 的静默策略违规故障): https://arxiv.org/abs/2607.07405ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents(ScopeJudge:攻击性安全 Agent 的成本感知前执行门禁): https://arxiv.org/abs/2607.07774How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement(Agent 如何请求权限:从界面到执行的 AI Agent 用户权限研究): https://arxiv.org/abs/2607.13718Toward Cryptographically Verifiable Authorization for Autonomous AI Agents(面向自治 AI Agent 的密码学可验证授权): https://arxiv.org/abs/2607.21325ToolGuardian: Declarative Security for AI Agent-Tool Interactions(ToolGuardian:AI Agent 与工具交互的声明式安全): https://arxiv.org/abs/2607.21835Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture(企业 AI Agent 的动态能力作用域): https://arxiv.org/abs/2607.22445Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents(LLM Agent 污点约束的 Agentic 权限策略代数): https://arxiv.org/abs/2607.24625Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks(面向自治网络的跨厂商 Agent 工具信任管理标准化): https://arxiv.org/abs/2607.25914FAVA: Formal Authorization for Verified Agents with Evidence-Backed Permission Graphs(FAVA:带证据权限图的形式化已验证 Agent 授权): https://arxiv.org/abs/2607.27267Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents(你还是我授权的 Agent 吗?对演化 Agent 的固定上限已获授权问题): https://arxiv.org/abs/2607.23586Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory(LLM Agent 中的记忆来源洗钱:针对持久记忆的非扩增防火墙): https://arxiv.org/abs/2607.29167Beyond Single-Use Tokens: Durable Authorization State for Replay-Resistant LLM Agent Actions(超越单次 token:抗重放 LLM Agent 动作的持久授权状态): https://arxiv.org/abs/2608.01710MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication(MNC:面向私密 LLM-Agent 通信的作用域绑定语义解密): https://arxiv.org/abs/2608.01719Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture(AI Agent 签名工作流的硬件密钥库:零信任 MCP 执行架构): https://arxiv.org/abs/2608.06130NiyamAI: An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs(NiyamAI:使用零知识证明实现意图绑定的加密可验证护栏): https://arxiv.org/abs/2608.07167A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem(企业 MCP 身份认证网关架构): https://arxiv.org/abs/2608.10760InterSAGE: The Secure and Verifiable Interoperability Protocol for An Internet of Agents(InterSAGE:Agent Internet 的安全可验证互操作协议): https://arxiv.org/abs/2608.13030Bounded Agents: Delegation Security for Multi-Agent AI Systems(Bounded Agents:多 Agent AI 系统中的委托安全): https://arxiv.org/abs/2608.15888Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems(在上下文之前授权:跨受众记忆泄漏的模型中立界限): https://arxiv.org/abs/2608.17148PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance(PACE:去中心化金融中安全 AI Agent 的策略认证合约执行): https://arxiv.org/abs/2608.17220From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions(从意图到执行授权:高风险 AI 动作的执行边界合规概述): https://arxiv.org/abs/2609.11596Scan the Skill, Govern the Action: Composing Registry Verdicts with Runtime Consequence Control(扫描 Skill,治理行为:整合注册表判定与运行时后果控制): https://arxiv.org/abs/2609.12001The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents(随机代理:工具使用 LLM Agent 的结构化租户隔离): https://arxiv.org/abs/2609.14780AcquireBound: Runtime Authorization for Resources Acquired by AI Agents(AcquireBound:AI Agent 获取资源的运行时授权): https://arxiv.org/abs/2609.14744LLM Agent Capabilities Should Follow Task Intent and Context Source(LLM Agent 能力应遵循任务意图和上下文来源): https://arxiv.org/abs/2609.14631Trust Propagation and Structural Containment in Multi-Agent LLM Pipelines(多 Agent LLM 流水线中的信任传播与结构性约束): https://arxiv.org/abs/2609.17648Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution(长时间运行 AI Agent 的授权撤销): https://arxiv.org/abs/2609.21284Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery(渐进式 Skill 发现作为工具使用 LLM Agent 的访问控制): https://arxiv.org/abs/2609.28693MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes(MetaPermit:通过 LLM 推断元属性实现 AI Agent 的可扩展审计式访问控制): https://arxiv.org/abs/2609.31039Crypto-bound identity-verified capability tokens for coordinating distributed AI agents: A proposal(加密绑定的身份验证能力 token:分布式 AI Agent 协调提案): https://arxiv.org/abs/2609.30824A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents(医疗 AI Agent 的安全边界 SDC-to-MCP 网关): https://arxiv.org/abs/2609.31358PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents(PACE:面向工具使用 LLM Agent 的来源感知能力执行控制): https://arxiv.org/abs/2610.01349Sapien: A Stateful Policy Engine for Autonomous AI Agents(Sapien:自治 AI Agent 的状态化策略引擎): https://arxiv.org/abs/2610.00797
###
### 运行时监控与沙箱
* ```
Aletheia: Permission-Minimality Testing for Coding-Agent Rules(Aletheia:编码 Agent 规则的最小权限测试): https://arxiv.org/abs/2609.39678HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control(HARDE:优化 Agent harness 以进行运行时风险检测和执行控制): https://arxiv.org/abs/2609.38291Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks(推测式安全蜜罐:面向多轮 Agent 攻击的主动防御): https://arxiv.org/abs/2609.39549Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents(Tracekit:自治编码 Agent 的防篡改意图-推理-动作审计): https://arxiv.org/abs/2609.35659Planarian: Managing Agent State with Statepoints(Planarian:用 Statepoints 管理 Agent 状态): https://arxiv.org/abs/2609.35366The Price of Safety: Benign-Case Utility and Token Overhead of Memory-Poisoning Defenses in LLM Agents(安全的代价:LLM Agent 记忆投毒防御的 benign-case 效用与 token 开销): https://arxiv.org/abs/2609.22818Forensic Trajectory Signatures for Agent Memory Poisoning Detection(用于 Agent 记忆投毒检测的取证轨迹签名): https://arxiv.org/abs/2606.30566MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems(MESA:优先保障多 Agent 系统中易受攻击的通信通道): https://arxiv.org/abs/2606.30602AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework(AI 沙箱:威胁模型、分类与度量框架): https://arxiv.org/abs/2606.18532Cordon: Semantic Transactions for Tool-Using LLM Agents(Cordon:工具使用 LLM Agent 的语义事务): https://arxiv.org/abs/2606.17573VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills(VIGIL:AI Agent Skill 中行为规范的运行时强制执行): https://arxiv.org/abs/2606.26524AIRGuard: Guarding Agent Actions with Runtime Authority Control(AIRGuard:通过运行时权限控制保护 Agent 动作): https://arxiv.org/abs/2605.28914SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs(SKILLLITE:用紧凑 LLM 进行证据引导的恶意 Skill 审计): https://arxiv.org/abs/2609.36879Beyond Handcrafted Security: Self-Evolving Defense (SED) for LLM Agents(超越手工安全:LLM Agent 的自进化防御): https://arxiv.org/abs/2609.36603Safeguarding LLM Agents from Misalignment through Provenance Analysis(通过来源分析保护 LLM Agent 免受误对齐): https://arxiv.org/abs/2607.01236Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital(真实资本下链上语言模型 Agent 的运行层控制): https://arxiv.org/abs/2604.26091Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents(治理你无法观察的东西:自治 AI Agent 的自适应运行时治理): https://arxiv.org/abs/2604.24686AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents(AgentWard:自治 AI Agent 的生命周期安全架构): https://arxiv.org/abs/2604.24657Behavioral Integrity Verification for AI Agent Skills(AI Agent Skill 的行为完整性验证): https://arxiv.org/abs/2605.11770ADR: An Agentic Detection System for Enterprise Agentic AI Security(ADR:企业 Agentic AI 安全的 Agentic 检测系统): https://arxiv.org/abs/2605.17380Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Features, Architectures, and Evaluation Protocols(LLM Agent 工具调用流量中的内容感知攻击检测): https://arxiv.org/abs/2605.11053Arbiter: Detecting Interference in LLM Agent System Prompts(Arbiter:检测 LLM Agent 系统 prompt 中的干扰): https://arxiv.org/abs/2603.08993MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in MCP Agents(MCPShield:适配式信任校准的安全认知层): https://arxiv.org/abs/2602.14281OpenClaw PRISM: A Zero-Fork, Defense-in-Depth Runtime Security Layer for Tool-Augmented LLM Agents(OpenClaw PRISM:工具增强 LLM Agent 的零 fork 防御深度运行时安全层): https://arxiv.org/abs/2603.11853Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the SSGM Framework(治理 LLM Agent 中不断演化的记忆): https://arxiv.org/abs/2603.11768Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees(保护 LLM Agent 长期记忆免受投毒): https://arxiv.org/abs/2606.24322ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents(ZoneClaw:通过记忆分区缓解持久记忆攻击): https://arxiv.org/abs/2610.00450Detecting Malicious Agent Skills in the Wild using Attention(利用注意力检测真实世界中的恶意 Agent Skill): https://arxiv.org/abs/2606.23416AgentSentry: Real-time Monitoring for Agentic AI Systems(AgentSentry:Agentic AI 系统的实时监控): https://arxiv.org/abs/2502.XXXXXMonitoring Emergent Reward Hacking via Internal Activations(通过内部激活监控新兴的 reward hacking): https://arxiv.org/abs/2603.04069Self-Attribution Bias: When AI Monitors Go Easy on Themselves(自我归因偏差:AI 监控器对自己太宽松): https://arxiv.org/abs/2603.XXXXXSalient Directions in AI Control(AI 控制中的关键方向): https://arxiv.org/abs/2603.XXXXXGoverned Memory: A Production Architecture for Multi-Agent Workflows(受治理的记忆:多 Agent 工作流生产架构): https://arxiv.org/abs/2603.17787Agent-Native Immune System: Architecture, Taxonomy, and Engineering(Agent 原生免疫系统:架构、分类与工程): https://arxiv.org/abs/2606.28270Behavioral Attestation and Compaction Drift in Persistent AI Agents(持久化 AI Agent 中的行为证明与压缩漂移): https://morrow.run/after-the-safety-gate.htmlThe Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills(分解即指纹:Agent Skill 的组件级身份识别): https://arxiv.org/abs/2606.31272From Tool Connection to Execution Control: Benchmarking Security Invariants in MCP-Style Agent Runtimes(从工具连接到执行控制:MCP 风格 Agent 运行时安全不变量评测): https://arxiv.org/abs/2606.29073AgenticOS: An Intent-Oriented Secure Operating System Architecture for Autonomous AI Agents(AgenticOS:面向意图的自治 AI Agent 安全操作系统架构): https://arxiv.org/abs/2606.21129Proof of Execution: Runtime Verification for Governed AI Agent Actions(执行证明:受治理 AI Agent 动作的运行时验证): https://arxiv.org/abs/2607.05397NovaFabric: Tamper-Evident, Replayable Evidence for Autonomous AI Agent Runs(NovaFabric:自治 AI Agent 运行的防篡改可回放证据): https://arxiv.org/abs/2609.12582When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems(当 Agent 失控:基于激活的多 Agent 系统恶意行为检测): https://arxiv.org/abs/2607.06807Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents(Token-Flow Firewall:持久化 AI Agent 的语义运行时审计): https://arxiv.org/abs/2607.08395TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories(TRACE:面向 LLM-Agent 轨迹的双通道稳健归因水印): https://arxiv.org/abs/2607.08400FlowGuard: From Signals to Evidence for MCP Security Detection(FlowGuard:从信号到证据的 MCP 安全检测): https://arxiv.org/abs/2607.14754When Agents Look Like Beacons: NIDS Evasion by Model Context Protocol Traffic(当 Agent 像信标:MCP 流量对 NIDS 的规避): https://arxiv.org/abs/2609.19091Democratizing Agent Deployment Safety: A Structural Monitoring Approach(推动 Agent 部署安全民主化:结构化监控方法): https://arxiv.org/abs/2607.14570Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents(跨 Agent 攻击活动归因:链接 LLM Agent 之间的异步攻击): https://arxiv.org/abs/2607.18826ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems(ChainWatch:基于 kill chain 的多步骤 MCP Agent 攻击顺序检测框架): https://arxiv.org/abs/2607.19432Operational Hallucination and Safety Drift in AI Agents(AI Agent 的操作性幻觉与安全漂移): https://arxiv.org/abs/2607.18366JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety(JANUS:预见长时 Agent 安全中的潜在风险): https://arxiv.org/abs/2607.19913CARE: Pre-Execution Command Verification for Shell-Executing LLM Agents(CARE:对 shell 执行 LLM Agent 的执行前命令验证): https://arxiv.org/abs/2607.21642SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems(SafeFlow:语义信息流控制阻断多 Agent 系统中的恶意传播): https://arxiv.org/abs/2607.25255Hybrid Analysis for Secure MCP Tool Use in LLM Agents(混合分析:保护 LLM Agent 对 MCP 工具的安全使用): https://arxiv.org/abs/2607.25297Distributing Security Controls Through Harness Engineering(通过 harness engineering 分发安全控制): https://arxiv.org/abs/2607.25890SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents(SkillGate:编码 Agent 中恶意 Skill 文件的低成本运行时检测): https://arxiv.org/abs/2607.25619Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study(多 Agent LLM 系统的分布式后门早期检测): https://arxiv.org/abs/2607.24893$S^3$: Improving Agent Safety through Multi-Stage Defense($S^3$:通过多阶段防御提升 Agent 安全性): https://arxiv.org/abs/2608.02683AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection(AgentAntibody:自适应免疫系统,抵御 Prompt Injection): https://arxiv.org/abs/2608.04053DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model(DreamGuard:通过风险感知 World Model 构建高效运行时护栏): https://arxiv.org/abs/2608.05695Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents(超越手写安全:面向 LLM Agent 的自进化防御): https://arxiv.org/abs/2608.12977Proof-of-Execution Memory: Defending LLM Agents Against Forged-Reasoning Attacks by Verifying What Actually Happened(执行证明记忆:通过验证真实发生的事情来防御伪造推理攻击): https://arxiv.org/abs/2608.16032DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents(DriftNet:双头轨迹 Transformer 用于检测和定位 LLM Agent 中的 Prompt Injection): https://arxiv.org/abs/2609.10892Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures(工程可靠提交门禁:Agentic AI 下的成本感知验证组合): https://arxiv.org/abs/2609.10969Provisional Reachability: Containing Agents by Making Every Crossing Revocable(临时可达性:让每一次跨越都可撤销,从而限制 Agent): https://arxiv.org/abs/2609.21957On the Effectiveness of Kernel-Level Evidence for Agent Security(内核级证据对 Agent 安全有效性的研究): https://arxiv.org/abs/2609.28915输入/输出校验Certified Multi-Source Integrity for Structured Agent Actions(结构化 Agent 动作的认证多源完整性): https://arxiv.org/abs/2609.34245IHDec: Divergence-Steered Contrastive Decoding for Securing Multi-Turn Instruction Hierarchies(IHDec:用于多轮指令层级安全的发散引导对比解码): https://arxiv.org/abs/2606.29960AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization(AgentVisor:通过语义虚拟化防御 LLM Agent Prompt Injection): https://arxiv.org/abs/2604.24118Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems(分析针对 Agentic AI 系统的防御性误导): https://arxiv.org/abs/2606.20470Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment(通过推理增强任务对齐防御自适应 Prompt Injection 攻击): https://arxiv.org/abs/2606.15441Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated Planning(三思而后行:通过隔离规划防护 LLM Agent 免受工具描述投毒): https://arxiv.org/abs/2606.20922Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection(解码护栏:基于 XAI 的 Prompt Injection 检测扰动分析): https://arxiv.org/abs/2609.24801StruQ: Defending Against Prompt Injection with Structured Queries(StruQ:用结构化查询防御 Prompt Injection): https://arxiv.org/abs/2402.06363Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation(通过偏差界定评估实现可证明无偏的 LLM Judge): https://arxiv.org/abs/2603.05485Judge Reliability Harness: Stress Testing LLM Judges(Judge Reliability Harness:压力测试 LLM Judge): https://arxiv.org/abs/2603.05399GELO: Good-Enough LLM Obfuscation(GELO:足够好的 LLM 混淆): https://arxiv.org/abs/2603.05035Untrusted Content Masking for Web Agents with Security Guarantees(带安全保证的 Web Agent 不可信内容屏蔽): https://arxiv.org/abs/2607.05277Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions(通过安全感知工具描述缓解 MCP Server 中的污点式漏洞): https://arxiv.org/abs/2607.07461Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models(面向语言模型交互中敏感数据隐私保护的多 Agent 防火墙架构): https://arxiv.org/abs/2607.08282PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis(PVDetector:通过策略违规概念分析检测针对特定任务 LLM Agent 的 Prompt Injection): https://arxiv.org/abs/2607.12624SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification(SingGuard-NSFA:面向 Agentic AI 的可扩展护栏): https://arxiv.org/abs/2607.13081ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems(ChannelGuard:安全模型并不等于安全的多 Agent 系统): https://arxiv.org/abs/2607.19430MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck(MIND:通过意图感知信息瓶颈实现轻量且有效的 LLM Agent 记忆注入防御): https://arxiv.org/abs/2607.28103PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents(PromptShield Home:面向智能家居 Agent 的环境多模态 Prompt Injection 防御): https://arxiv.org/abs/2608.05495Robust Context-Aware Detection of Malicious Instructions in Text(鲁棒的上下文感知文本恶意指令检测): https://arxiv.org/abs/2608.05430Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents(工具规范很重要:揭示并缓解 AI Agent 的安全风险): https://arxiv.org/abs/2607.29254PIPES: Securing Agent Perception with Provenance and Priors(PIPES:通过来源和先验保护 Agent 感知): https://arxiv.org/abs/2608.12789Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach(Agent Skill 的跨层错位检测): https://arxiv.org/abs/2607.10534Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks(工具集成 LLM Agent 的通用防御): https://arxiv.org/abs/2609.16098The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents(可验证动作卡:安全自治 Agent 的可信人类在环控制): https://arxiv.org/abs/2609.18411Detecting and Localizing Segment-Level Poisoning in Multi-Source LLM-Agent Inputs(检测和定位多源 LLM-Agent 输入中的分段投毒): https://arxiv.org/abs/2609.14723Prompt Injection Detection for Email Agents Through Attack Chain Modeling(通过攻击链建模检测邮件 Agent 的 Prompt Injection): https://arxiv.org/abs/2609.30657形式化验证与分析Efficient and Sound Probabilistic Verification for AI Agents(AI Agent 的高效且可靠概率验证): https://arxiv.org/abs/2606.20510Agentics 2.0: Logical Transduction Algebra for Agentic Data Workflows(Agentics 2.0:Agentic 数据工作流的逻辑转导代数): https://arxiv.org/abs/2603.04241Knowledge Divergence and the Value of Debate for Scalable Oversight(知识分歧与辩论在可扩展监督中的价值): https://arxiv.org/abs/2603.XXXXXAutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming(AutoSpec:通过归纳逻辑编程演化 LLM Agent 安全规则): https://arxiv.org/abs/2606.24245Local LLM Agents as Vulnerable Runtimes: A Source-Code Audit of the Agent Runtime Layer(本地 LLM Agent 作为脆弱运行时:Agent 运行时层源码审计): https://arxiv.org/abs/2606.21071AgentFlow: Building Agent Dependency Graphs for Static Analysis of Agent Programs(AgentFlow:构建 Agent 依赖图用于静态分析): https://arxiv.org/abs/2607.01640SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills(SkillsMetric:绘制静态分析对恶意 Agent Skill 检测边界): https://arxiv.org/abs/2608.08468SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment(SkillConsist:通过双向图对齐检测 Agent Skill 不一致): https://arxiv.org/abs/2608.07639Correct Is Not Governed: Provenance Integrity in Agentic Workflows(正确并不等于受治理:Agentic 工作流中的来源完整性): https://arxiv.org/abs/2608.12761CaMeLoT: CaMeL orchestrated with Temporal logic for static verification and liveness(CaMeLoT:将 CaMeL 与时序逻辑结合进行静态验证和活性分析): https://arxiv.org/abs/2609.18674Formal Analysis and Supply Chain Security for Agentic AI Skills(Agentic AI Skill 的形式化分析与供应链安全): https://arxiv.org/abs/2603.00195
评估与红队测试
- “`
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation(评估系统 1 模型用于 Agent 安全决策的可靠性、校准与选择性自动化): https://arxiv.org/abs/2609.33401Ajar: Measuring Open Privilege in Agent Defenses(Ajar:测量 Agent 防御中的开放式权限): https://arxiv.org/abs/2609.26900MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG(MIRROR:针对 Agentic RAG 的新颖性约束、记忆引导 MCTS 红队测试): https://arxiv.org/abs/2606.26793CONTRA: Red-Teaming Configurations of Personalizable Agents(CONTRA:可定制 Agent 配置红队测试): https://arxiv.org/abs/2607.03220Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification(大规模安全测试 LLM Agent:从风险发现到证据化验证): https://arxiv.org/abs/2607.01793The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems(精心设计的 SCHEMEs:多 Agent 系统中的协同破坏与监控): https://arxiv.org/abs/2605.29178ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D(ResearchArena:评估自动化 AI R&D 中的破坏与监控): https://arxiv.org/abs/2607.19321Real-Time Trust Verification for Safe Agentic Actions using TrustBench(TrustBench:面向安全 Agent 动作的实时信任验证): https://arxiv.org/abs/2603.09157Agent Security Bench (ASB)(Agent 安全 benchmark): https://arxiv.org/abs/2410.02644SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills(SkillVetBench:LLM-as-Judge 对开源 LLM Agent Skill 进行多维度安全风险评估): https://arxiv.org/abs/2606.15899Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents(对 LLM Agent 中旁路防御的适配性评估): https://arxiv.org/abs/2606.26479AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents(AutoDojo:自适应攻击揭示表面型防御与用户约束不足): https://arxiv.org/abs/2606.15057SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces(SkillSafetyBench:评估 Skill 攻击面下 Agent 的安全性): https://arxiv.org/abs/2605.12015No More, No Less: Task Alignment in Terminal Agents(恰到好处:终端 Agent 的任务对齐): https://arxiv.org/abs/2605.12233Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems(Proteus:Agent Skill 生态的自进化红队): https://arxiv.org/abs/2605.11891SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills(SkillSecurer:检测和修补 AI Agent Skill 中的 Prompt Injection 漏洞): https://arxiv.org/abs/2609.14079
* ```
SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills(SkillSentry:用于 Agent Skill 动态安全测试的自适应蜜罐世界): https://arxiv.org/abs/2608.03485AgentDyn: A Dynamic Open-Ended Benchmark for Prompt Injection Attacks(AgentDyn:动态开放式 Prompt Injection 攻击 benchmark): https://arxiv.org/abs/2602.03117NAAMSE: Framework for Evolutionary Security Evaluation of Agents(NAAMSE:Agent 进化式安全评估框架): https://arxiv.org/abs/2602.07391R-Judge: Benchmarking Safety Risk Awareness(R-Judge:安全风险感知 benchmark): https://arxiv.org/abs/2401.10019InjecAgent: Benchmarking Indirect Prompt Injections(InjecAgent:间接 Prompt Injection benchmark): https://arxiv.org/abs/2403.02691SIABENCH: Evaluating LLMs for Security Incident Analysis(SIABENCH:评估 LLM 在安全事件分析中的表现): https://arxiv.org/abs/2603.XXXXXEVMbench: Evaluating AI Agents on Smart Contract Security(EVMbench:评估 AI Agent 在智能合约安全上的表现): https://arxiv.org/abs/2603.04915τ-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge(τ-Knowledge:评估会话 Agent 对非结构化知识的处理): https://arxiv.org/abs/2603.04370Interactive Benchmarks(交互式 benchmark): https://arxiv.org/abs/2603.04737VeriGrey: Greybox Agent Validation(VeriGrey:灰盒 Agent 验证): https://arxiv.org/abs/2603.17639LAAF: Logic-layer Automated Attack Framework for Agentic LLM Systems(LAAF:Agentic LLM 系统的逻辑层自动化攻击框架): https://arxiv.org/abs/2603.17239VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers(VIPER-MCP:检测和利用 MCP Server 中的污点式漏洞): https://arxiv.org/abs/2605.21392SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents(SafeClawBench:区分工具使用 LLM Agent 中的语义、审计证据和沙箱危害): https://arxiv.org/abs/2606.18356Red-Teaming the Agentic Red-Team(Agentic 红队的红队测试): https://arxiv.org/abs/2606.24496Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming(保护 AI Agent:多层 Agent 红队测试的统一框架): https://arxiv.org/abs/2606.31227Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens(通过计算机系统视角理解和评估 Claw-like Agent 安全): https://arxiv.org/abs/2606.30755Security–Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense(安全—保真度权衡:Prompt Injection 防御的隐藏成本): https://arxiv.org/abs/2606.30783When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents(当 AUC 0.998 还不够时:多模态计算机使用 Agent 中间接 Prompt Injection 隐状态探针评估方案): https://arxiv.org/abs/2606.22864Distributed Attacks in Persistent-State AI Control(持续状态 AI Control 中的分布式攻击): https://arxiv.org/abs/2607.02514Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents(超越攻击成功率:工具使用 AI Agent 的动作分级严重程度量表): https://arxiv.org/abs/2607.07474Rethinking MCP Security: A Large-Scale Study of Runtime MCP Servers and Security Scanner Reliability(重新思考 MCP 安全:运行时 MCP Server 和安全扫描器可靠性的实证研究): https://arxiv.org/abs/2607.11086Exposed by Design: A Dynamic Security Assessment of Internet-Facing MCP Servers at Scale(暴露即设计:互联网暴露式 MCP Server 的大规模动态安全评估): https://arxiv.org/abs/2608.00150Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation(重构 AI 系统渗透测试:从资源妥协到行为目标违规): https://arxiv.org/abs/2607.14006Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security(适应性对手:LLM Agent 安全的多轮、多模型 benchmark): https://arxiv.org/abs/2607.18063RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery(RECEIPT:白盒 Agentic XSS 发现的确定性、抗 reward hacking 验证): https://arxiv.org/abs/2607.18575Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming(Agent 对抗 Agent:用于自动 Prompt Injection 红队测试的 Agentic 系统): https://arxiv.org/abs/2608.05108Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents(了解你的 Agent:基于侦察驱动的 AI Agent 渗透测试): https://arxiv.org/abs/2607.19837Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents(把护栏当替罪羊:审计工具增强 LLM Agent 中不忠实的安全拒答): https://arxiv.org/abs/2607.19449IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests(IssueTrojanBench:评估 AI 编码 Agent 对恶意 issue 请求的抗性): https://arxiv.org/abs/2607.20759OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills(OpenSkillRisk:评估真实世界风险第三方 Skill 使用时 Agent 安全性): https://arxiv.org/abs/2607.20121Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation(标签不是终点:MCP Agent 安全评估中的处理泄漏与构念效度): https://arxiv.org/abs/2608.12880Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming(Agent 攻击 Agent:生产环境 Agent 的自动研究红队测试): https://arxiv.org/abs/2607.11698Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection(DeepSeek Harness 安全评估:评估其对间接 Prompt Injection 的抵抗能力): https://arxiv.org/abs/2608.16393BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure(BenchShield:LLM-Agent 评测基础设施中奖励完整性的形式化模型支撑仪表化): https://arxiv.org/abs/2609.11028
- “`
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems(Emergence World:长时多 Agent 系统的对抗性压力测试): https://arxiv.org/abs/2609.17320Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery(黑盒 Agentic AI 红队测试:基于分类法的自动化风险发现框架): https://arxiv.org/abs/2609.09647Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents(Auto Mode 红队测试:增强阻断分类器对恶意编码 Agent 的识别能力): https://arxiv.org/abs/2609.19587AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination(AgentLSD:在对抗性任务污染下评估 AI 安全 Agent): https://arxiv.org/abs/2609.19140Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery(LLM Agent 系统中的集体失控:突变、传播与恢复的流行病学解释): https://arxiv.org/abs/2609.18460Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI(谁来审计谁?在什么基础设施上?用什么证据?Agentic AI 的独立性审计协议): https://arxiv.org/abs/2609.18272A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems(可信大语言模型、Agentic AI 和多模态系统的统一评估框架): https://arxiv.org/abs/2609.19524Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure(在正常任务压力下出现工具性监控规避): https://arxiv.org/abs/2609.30217
---
## 基准测试与数据集
| 基准测试 | 关注点 | 规模 | 论文 |
| --- | --- | --- | --- |
| ASB(Agent Security Bench) | 全面 Agent 安全评测 | 10 个 agent,398 个环境 | Zhang et al. |
| InjecAgent | 间接 Prompt Injection | 1,054 个测试用例 | Zhan et al. |
| R-Judge | 安全风险感知 | 162 条记录,27 个场景 | Yuan et al. |
| ToolSword | 工具学习安全 | 6 个场景,3 个阶段 | Ye et al. |
| AgentDyn | 动态 Prompt Injection | 开放式、可扩展 | Li et al. |
| SkillSafetyBench | Skill 中介的 Agent 安全 | 155 个案例,47 个任务 | Jin et al. |
| SkillVetBench | 开源 Agent Skill 风险评估 | 实时 leaderboard | Hossain et al. |
| SCR-Bench | Skill 组合风险 | 多个技能链 | Xie et al. |
| SafeClawBench | 工具使用 Agent 的阶段性危害 | 600 个对抗任务 | Tian et al. |
| ToolPrivacyBench | 工具使用 Agent 中的目的绑定隐私 | 2,150 个案例 | Hu et al. |
| ToxicBench | 工具输出投毒与盲目遵循 | 118 个任务 | Tao & Yin |
| TAB | 终端 Agent 的选择性跟随 | 89 个终端任务 | Mavali et al. |
| Skill-Inject | Skill 文件攻击 | 多种场景 | Schmotz et al. |
| NAAMSE | Agent 安全的进化性评估 | 自适应红队 | Pai et al. |
| AgentHarm | Agent 误用 | 110 个行为,440 个变体 | Andriushchenko et al. |
| SkillGuard Dataset | 恶意 Skill 检测 | 157 个恶意 Skill | Liu et al. |
| WIPI | Web 间接注入 | 多种场景 | Liu et al. |
| DUMA-Bench | 双控制 Agent 安全(8 类漏洞) | 8 个领域,14 个模型 | Aleksandrov et al. |
| SkillAtlas | Agent Skill 攻击轨迹库 | 3,014 个案例,6,589 条轨迹 | Tian et al. |
| IssueTrojanBench | 恶意 issue 请求对编程 Agent 的影响 | 4 类攻击,6 种向量 | Singh et al. |
| OpenSkillRisk | 对真实风险第三方 Skill 的安全评估 | 263 个 Skill,7 类别 | Liu et al. |
| AudioAgentSecurity | 音频 Prompt Injection vs 多模态 Agent | 8 个场景,10 种模式 | Liu et al. |
| AIP-Bench | Agentic commerce 协议级安全 | 3 个平台,33 个漏洞 | Louck |
| ContainmentBench | 注入后的遏制能力 | 轨迹驱动,多指标 | Lan et al. |
| MemSecBench | Agent 记忆投毒生命周期安全 | 310 个案例,24 配置矩阵 | Chen et al. |
| IH-Benchmark | 指令层级鲁棒性(含工具介导冲突) | 44 个约束族,37 个模型 | McCauley et al. |
| AgentS4D | 工作区 Agent 的生命周期运行时风险 | 沙箱化、轨迹感知 | Zhou et al. |
| StepJack | 多步间接 Prompt Injection | 480 个测试用例 | Zhan et al. |
| HarnessSafe | 持续载体下的安全性 | 328 个案例,7 类载体 | Zhang et al. |
| ToolHazard | 对抗性工具环境的可扩展合成 | 可通过 seed domain + compute 扩展 | Mou et al. |
| ATOBench | 电子渗透测试 Agent 在欺骗性证据下的漏洞验证 | 运行时响应转换 | Chen et al. |
| HarnessRisk | Agent harness 生命周期安全 | 6 个操作阶段 | Bai et al. |
| ChemMat-AgentSafetyBench | 化学与材料 Agents 中的长时攻击/防御 | 432 个案例,8 种危害类型 | Zhang et al. |
| PIDS-Bench | Prompt-Injection 检测器在过度防御、混淆和分布漂移下的表现 | 7 个检测器,多指标 | Shir... |
| HTB (Hallucinated-Tools Benchmark) | Tool 幻觉问题(伪造工具/参数) | 322 个单注册表 + 154 个多注册表案例 | Sun et al. |
| ClashBench | 特权 Agent 对资源的破坏性抢占 | 268 个冲突案例,55 个环境 | Liu et al. |
| MemoryStackBench | Agent 记忆安全与可审计性 | 5 个场景,33 个检查项 | Repository |
| mcp-defense-bench | MCP 防御代理攻击面覆盖 | 24 个向量,35 个案例 | Arumugam |
| VLoc Bench | 安全 Agent 的仓库级漏洞定位 | 500 个漏洞,290 个 repos,147 个 CWEs | Priyanshu et al. |
| APort Vault | 工具使用 Agent 的支付授权 | 4,371 个人工 CTF 攻击 | AporHQ |
| EvasionBench | 常规任务压力下的运行时监控规避 | 50 个任务-策略对 | Schmotz et al. |
| ACE | 内核系统调用 + 应用层证据 | 4,047 个配对会话,17 个威胁模型 | King et al. |
---
## 工具与框架
| 工具 | 说明 | 链接 |
| --- | --- | --- |
| Agent Memory Guard | OWASP 对 ASI06(Memory Poisoning)的参考实现 | ``` https://github.com/OWASP/www-project-agent-memory-guard ``` |
| Bounty Sieve | 针对 coding agent 的离线默认 bounty 入口保护 | ``` https://github.com/junbuilds96/bounty-sieve ``` |
| SkillGuard | 基于 LLM 的 Agent Skill 安全审计器 | ``` https://github.com/LLMSecurity/skillguard ``` |
| SkillCI | 回归测试 + OWASP Agentic Skills Top 10 映射静态安全 lint | ``` https://github.com/kabirnarang39/skillci ``` |
| Pipelock | 开源 AI Agent 防火墙与 MCP-aware egress proxy | ``` https://github.com/luckyPipewrench/pipelock ``` |
| NemoClaw | NVIDIA 参考实现,用于安全运行常驻型 AI Agent | ``` https://github.com/NVIDIA/NemoClaw ``` |
| CubeSandbox | 硬件隔离(per-kernel)沙箱 | ``` https://github.com/TencentCloud/CubeSandbox ``` |
| Invariant Guardrails | 基于策略的 Agent 安全护栏 | ``` https://github.com/invariantlabs-ai/invariant ``` |
| Armorer Guard | 本地 Rust 扫描器 | ``` https://github.com/ArmorerLabs/Armorer-Guard ``` |
| Sunglasses | Agent skills / tool use 的运行时信任扫描器 | ``` https://github.com/sunglasses-dev/sunglasses ``` |
| LLM Guard | 对 LLM 应用进行输入/输出扫描 | ``` https://github.com/protectai/llm-guard ``` |
| Rebuff | 自我强化 Prompt Injection 检测器 | ``` https://github.com/protectai/rebuff ``` |
| NeMo Guardrails | NVIDIA 提供的 LLM 应用护栏工具包 | ``` https://github.com/NVIDIA/NeMo-Guardrails ``` |
| Lakera Guard | 企业级 Prompt Injection 防御 API | ``` https://www.lakera.ai/ ``` |
| Promptfoo | LLM 红队测试和评估框架 | ``` https://github.com/promptfoo/promptfoo ``` |
| Garak | LLM 漏洞扫描器 | ``` https://github.com/leondz/garak ``` |
| IPI-Proxy | Web-browsing Agent 的间接 Prompt Injection 红队代理 | ``` https://github.com/VulcanLab/IPI-Proxy ``` |
| Tuning Engines CLI | MCP server + CLI,用于治理 Agent/Skill/Tool 权限 | ``` https://github.com/cerebrixos-org/tuning-engines-cli ``` |
| AgentSkillsScanner | Agent Skill 定义的静态分析扫描器 | ``` https://github.com/sumleo/AgentSkillsScanner ``` |
| repo-agent-scan | 针对 Agent Skill 与仓库指令文件的本地确定性扫描 | ``` https://github.com/sunxiayi/repo-agent-instruction-security-scan ``` |
| SkillTotal | 对 AI 组件进行静态离线扫描 | ``` https://github.com/pezhik/skilltotal ``` |
| SkilLock | Behavior-pinning lockfile + capability-delta PR review | ``` https://github.com/skills-lock/skil-lock ``` |
| agent-diff-guard | 预推送 guardrail | ``` https://github.com/cubxxw/agent-diff-guard ``` |
| Skillid | 策略驱动的 Claude Code 插件 | ``` https://github.com/dkoly/skillid ``` |
| Agent Audit | LLM Agent 应用安全分析系统 | ``` https://arxiv.org/abs/2603.22853 ``` |
| mcp-sec-audit | MCP Server 安全工具包 | ``` https://arxiv.org/abs/2603.21641 ``` |
| SkillGate | 本地 CLI,静态检查 Agent Skill 包 | ``` https://github.com/selfradiance/skillgate ``` |
| Assay Harness | CI gate,校验 Agent 声称的工具副作用与实际运行时证据 | ``` https://github.com/Rul1an/Assay-Harness ``` |
| Agent Scan | Snyk 的本地 Agent 供应链扫描器 | ``` https://github.com/snyk/agent-scan ``` |
| DScan | 开源 Agent 安全套件 | ``` https://github.com/DeepScan-Security/dscan ``` |
| Sayfos SDK | AI Agent 运行时护栏 SDK | ``` https://github.com/sayfos-labs/sayfos-sdk ``` |
| PIC Standard | 本地优先标准及参考校验器 | ``` https://github.com/madeinplutofabio/pic-standard ``` |
| TraceFold | Rust 引擎 | ``` https://github.com/TraceFold/tracefold ``` |
| agent-sentinel | eBPF/BPF-LSM 原型 | ``` https://github.com/junlinwk/agent-sentinel ``` |
| mcp-sploit | Metasploit 风格的授权安全测试框架 | ``` https://github.com/Prasanna-27eng/mcp-sploit ``` |
| Clawvisor | Agent 网关 | ``` https://github.com/clawvisor/clawvisor ``` |
| trentclaw | OpenClaw 环境安全评估 Skill | ``` https://github.com/trnt-ai/trent-openclaw-security-assessment ``` |
| Nobulex | Agent 信任资本评分层 | ``` https://github.com/arian-gogani/nobulex ``` |
| Clay Seal | Agent 的 attested runtime identity 与 Biscuit capability token | ``` https://github.com/clayseal/clayseal-identity ``` |
| SkillWatch | 定期重新检查外部 URL | ``` https://github.com/kuzivaai/SkillWatch ``` |
| Humanbound | 开源对抗性测试引擎与 SDK | ``` https://github.com/humanbound/humanbound ``` |
| whatileaked | 扫描 coding agent 已写入磁盘的凭据 | ``` https://github.com/selan-ai/whatileaked ``` |
| Skill-audit | Claude/agent skills 的静态扫描器 | ``` https://github.com/AgentPostmortem/Skill-audit ``` |
| AgentWarden | MCP 配置/AI Skill 的静态安全门禁 | ``` https://github.com/juangh123/AgentWarden ``` |
| AVE | Agentic AI 组件的开放标准与行为漏洞分类 | ``` https://github.com/aveproject/ave ``` |
| pikit | 组合式间接 Prompt Injection 研究工具包 | ``` https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/pikit ``` |
---
## Agent Skill 规范
| 规范 | 组织 | 关注点 |
| --- | --- | --- |
| AgentSkills.io | Open Standard | Agent Skill 定义与安全要求 |
| Model Context Protocol (MCP) | Anthropic | LLM 的工具/资源集成协议 |
| OpenAI Function Calling | OpenAI | GPT 模型的工具调用规范 |
| Tool Use (Claude) | Anthropic | Claude 原生工具使用接口 |
| LangChain Tools | LangChain | Agent 框架中的工具抽象 |
| AutoGPT Plugins | AutoGPT | 自主 Agent 插件系统 |
| OpenAPI/Swagger | Linux Foundation | API 规范,常被用作工具定义 |
---
## 行业报告与博客文章
* ```
Snowflake Cortex AI Escapes Sandbox and Executes Malware(Snowflake Cortex AI 逃逸沙箱并执行恶意代码): https://www.promptarmor.com/resources/snowflake-ai-escapes-sandbox-and-executes-malwareConfused Deputy Attacks on Autonomous AI Agents(自治 AI Agent 的混淆代理攻击): https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-agent-confused-deputy-prompt-injection/How AI Assistants are Moving the Security Goalposts(AI 助手如何改变安全目标线): https://krebsonsecurity.com/2026/03/how-ai-assistants-are-moving-the-security-goalposts/Hackers Used Meta’s AI Support Bot to Seize Instagram Accounts(黑客利用 Meta AI 支持机器人控制 Instagram 账号): https://krebsonsecurity.com/2026/06/hackers-used-metas-ai-support-bot-to-seize-instagram-accounts/Anthropic: Challenges in Red Teaming AI Systems(Anthropic:AI 系统红队测试的挑战): https://www.anthropic.com/index/challenges-in-red-teaming-ai-systemsOpenAI: Safety of Advanced AI Agents(OpenAI:高级 AI Agent 的安全性): https://openai.com/research/practices-for-governing-agentic-ai-systemsCompromising Agents via MCP(通过 MCP 攻破 Agent): https://invariantlabs.ai/blog/mcp-securitySimon Willison: Prompt Injection Explained(Simon Willison:Prompt Injection 解析): https://simonwillison.net/2023/Apr/14/worst-that-can-happen/The sorry state of skill distribution(Skill 分发的糟糕状态): https://blog.trailofbits.com/2026/06/03/the-sorry-state-of-skill-distribution/TRAIL: Trusted Reasoning and AI Logging(TRAIL:可信推理与 AI 日志): https://arxiv.org/abs/2502.XXXXXCyber Threat Intelligence for AI Systems(AI 系统网络威胁情报): https://arxiv.org/abs/2603.05068AI Safety Has 12 Months Left(AI 安全只剩 12 个月): https://mhdempsey.substack.com/p/ai-safety-has-12-months-leftLiteLLM Hack: Were You One of the 47,000?(LiteLLM 被攻破:你是 47,000 人中的一员吗?): https://futuresearch.ai/blog/litellm-hack-were-you-one-of-the-47000/Lack of Isolation in Agentic Browsers Resurfaces Old Vulnerabilities(Agentic 浏览器隔离缺失重现旧漏洞): https://blog.trailofbits.com/2026/01/13/lack-of-isolation-in-agentic-browsers-resurfaces-old-vulnerabilities/OpenAI Help: Lockdown Mode(OpenAI 帮助:锁定模式): https://simonwillison.net/2026/Jun/5/openai-help-lockdown-mode/ClawHub by the Numbers: Metadata on All 52,652 Packages(ClawHub 数据概览:全量 52,652 个包的元数据): https://trent.ai/blog/clawhub-by-the-numbers/The Memory Heist(记忆窃取事件): https://www.ayush.digital/blog/the-memory-heistOpenAI's Accidental Cyberattack Against Hugging Face is Science Fiction That Happened(OpenAI 对 Hugging Face 的意外网络攻击): https://simonwillison.net/2026/Jul/22/openai-cyberattack/Auto Mode Is Now the Default in Claude Code(Auto Mode 现在是 Claude Code 默认模式): https://simonwillison.net/2026/Aug/8/auto-mode/OpenAI Agents Carried Out an Undisclosed Attack on RubyGems(OpenAI Agent 执行了未披露的 RubyGems 攻击): https://simonwillison.net/2026/Sep/12/openai-agents-rubygems/Is sandboxing sufficient to contain rogue agents?(沙箱是否足以遏制恶意 Agent?): https://blog.cryptographyengineering.com/2026/09/30/is-sandboxing-sufficient-to-contain-rogue-agents/Caught in the Hook: RCE and API Token Exfiltration Through Claude Code Project Files(被钩住:通过 Claude Code 项目文件实现 RCE 和 API Token 外泄): https://research.checkpoint.com/2026/rce-and-api-token-exfiltration-through-claude-code-project-files/相关 Awesome 列表awesome-llm-security(通用 LLM 安全资源): https://github.com/corca-ai/awesome-llm-securityawesome-ai-safety(AI 安全研究与资源): https://github.com/hari-sikchi/awesome-ai-safetyawesome-chatgpt-prompts(Prompt engineering,含对抗性样例): https://github.com/f/awesome-chatgpt-promptsawesome-ml-for-cybersecurity(应用于网络安全的 ML 资源): https://github.com/jivoi/awesome-ml-for-cybersecurityawesome-mcp-servers(MCP Server 生态,攻击面参考): https://github.com/punkpeye/awesome-mcp-serversawesome-ai-agents(AI Agent 框架与项目): https://github.com/e2b-dev/awesome-ai-agents
免责声明:
本文所载程序、技术方法仅面向合法合规的安全研究与教学场景,旨在提升网络安全防护能力,具有明确的技术研究属性。
任何单位或个人未经授权,将本文内容用于攻击、破坏等非法用途的,由此引发的全部法律责任、民事赔偿及连带责任,均由行为人独立承担,本站不承担任何连带责任。
本站内容均为技术交流与知识分享目的发布,若存在版权侵权或其他异议,请通过邮件联系处理,具体联系方式可点击页面上方的联系我。
本文转载自:小猫信安 小猫信安
小猫信安《AI安全、agent安全所有外刊文章汇总(全网最全)》