Posted in last 7 days

Agentic AI Researchers - Harness Safety & Forensics

๐——๐—ฎ๐˜† ๐Ÿญ/๐Ÿญ๐Ÿฌ๐Ÿฌ โ€” ๐Ÿญ๐Ÿฌ๐Ÿฌ ๐——๐—ฎ๐˜†๐˜€ ๐—ผ๐—ณ ๐—”๐—ด๐—ฒ๐—ป๐˜๐—ถ๐—ฐ ๐—”๐—œ ๐—™๐—ผ๐—ฟ๐—ฒ๐—ป๐˜€๐—ถ๐—ฐ๐˜€ This is Day 1 of my โ€œ100 Days of Agentic AI Forensicsโ€ series, where Iโ€™ll explore agent security, attacks, forensics, harness safety, and trustworthy agentic systems. Iโ€™d really like to hear from researchers working on AI agents, agent security, multi-agent systems, or trustworthy AI. ๐—”๐—ฟ๐—ฒ ๐—ช๐—ฒ ๐—ง๐—ฟ๐˜‚๐—น๐˜† ๐— ๐—ฒ๐—ฎ๐˜€๐˜‚๐—ฟ๐—ถ๐—ป๐—ด ๐—›๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€ ๐—ฆ๐—ฎ๐—ณ๐—ฒ๐˜๐˜† ๐—ถ๐—ป ๐—”๐—ด๐—ฒ๐—ป๐˜๐—ถ๐—ฐ ๐—”๐—œ ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ๐˜€? Imagine this: ๐—ง๐—ฎ๐˜€๐—ธ: Prepare a financial report using the companyโ€™s authorized financial files. The agent produces the correct report. But during execution, it: โœ… Accesses financial.csv โŒ Accesses unauthorized employee_salary.csv โŒ Shares salary information with another agent โœ… Still completes the task successfully Soโ€ฆ ๐—ช๐—ฎ๐˜€ ๐˜๐—ต๐—ฒ ๐—ฎ๐—ด๐—ฒ๐—ป๐˜ ๐˜€๐—ฎ๐—ณ๐—ฒ? This is the key question behind: ๐Ÿ“„ HarnessAudit: Auditing Agent Harness Safety https://lnkd.in/g9rDGB5X The paper argues that evaluating an agent only by its final answer or task success is not enough. ๐—›๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€ ๐—ฆ๐—ฎ๐—ณ๐—ฒ๐˜๐˜† โ‰ˆ Task Completion + Step-Level Permission Checks + Cross-Agent Information Flow + Workflow Fidelity + Robustness One striking observation is that safety violations can accumulate as agents execute more actions. Even strong evaluated configurations achieve only around ๐Ÿฌ.๐Ÿฏ๐Ÿฎ overall safety score, with many violations related to unauthorized resource access and inter-agent information sharing. But, ๐—”๐—ฟ๐—ฒ ๐˜„๐—ฒ ๐—บ๐—ฒ๐—ฎ๐˜€๐˜‚๐—ฟ๐—ถ๐—ป๐—ด ๐—ฒ๐—ป๐—ผ๐˜‚๐—ด๐—ต? ๐—ช๐—ต๐—ฎ๐˜ ๐—ถ๐—บ๐—ฝ๐—ผ๐—ฟ๐˜๐—ฎ๐—ป๐˜ ๐—ฑ๐—ถ๐—บ๐—ฒ๐—ป๐˜€๐—ถ๐—ผ๐—ป ๐—ผ๐—ณ ๐—ต๐—ฎ๐—ฟ๐—ป๐—ฒ๐˜€๐˜€ ๐˜€๐—ฎ๐—ณ๐—ฒ๐˜๐˜† ๐—ฎ๐—ฟ๐—ฒ ๐˜„๐—ฒ ๐˜€๐˜๐—ถ๐—น๐—น ๐—บ๐—ถ๐˜€๐˜€๐—ถ๐—ป๐—ด? Some questions I kept thinking about while reading the paper: โ€ข Did the agent violate the policy naturally, or was it intentionally manipulated? โ€ข If an attack exists, is it direct (active attack) where the attacker agent itself breaks the rules or indirect (passive attack), where a malicious agent manipulates a benign agent into violating them? โ€ข What about private information that is never explicitly sent through a tool, but is implicitly stored or transferred through the agentโ€™s internal state, memory, or communication? โ€ข Could an agent contain a backdoor that remains dormant until a specific trigger appears? โ€ข Could sensitive information leak indirectly through knowledge transfer or subliminal learning, even when the victim agent does not realize it is revealing anything? Subliminal Learning: https://lnkd.in/g4zUWP4y โ€ข Can an agent achieve the goal while silently bypassing a required security layer? โ€ข What about attacks that do not steal information at all, but push other agents into deadlocks, infinite loops, or resource-exhaustion states? Comments, corrections, paper suggestions, and different viewpoints are very welcome. Iโ€™m also open to discussions and collaborations around these questions. #AgenticAI #AISecurity #AIForensics #MultiAgentSystems #TrustworthyAI #LLMSecurity #AgentSecurity #ArtificialIntelligence #Research
Sourcee Logo

Brought to you by Sourcee

We find journo requests from across the web and deliver them directly to your inbox.

We Monitor the Web for Journo Requests