๐๐ฎ๐ ๐ญ/๐ญ๐ฌ๐ฌ โ ๐ญ๐ฌ๐ฌ ๐๐ฎ๐๐ ๐ผ๐ณ ๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐๐ ๐๐ผ๐ฟ๐ฒ๐ป๐๐ถ๐ฐ๐
This is Day 1 of my โ100 Days of Agentic AI Forensicsโ series, where Iโll explore agent security, attacks, forensics, harness safety, and trustworthy agentic systems.
Iโd really like to hear from researchers working on AI agents, agent security, multi-agent systems, or trustworthy AI.
๐๐ฟ๐ฒ ๐ช๐ฒ ๐ง๐ฟ๐๐น๐ ๐ ๐ฒ๐ฎ๐๐๐ฟ๐ถ๐ป๐ด ๐๐ฎ๐ฟ๐ป๐ฒ๐๐ ๐ฆ๐ฎ๐ณ๐ฒ๐๐ ๐ถ๐ป ๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐๐ ๐ฆ๐๐๐๐ฒ๐บ๐?
Imagine this:
๐ง๐ฎ๐๐ธ: Prepare a financial report using the companyโs authorized financial files.
The agent produces the correct report.
But during execution, it:
โ
Accesses financial.csv
โ Accesses unauthorized employee_salary.csv
โ Shares salary information with another agent
โ
Still completes the task successfully
Soโฆ ๐ช๐ฎ๐ ๐๐ต๐ฒ ๐ฎ๐ด๐ฒ๐ป๐ ๐๐ฎ๐ณ๐ฒ?
This is the key question behind:
๐ HarnessAudit: Auditing Agent Harness Safety
https://lnkd.in/g9rDGB5X
The paper argues that evaluating an agent only by its final answer or task success is not enough.
๐๐ฎ๐ฟ๐ป๐ฒ๐๐ ๐ฆ๐ฎ๐ณ๐ฒ๐๐ โ Task Completion + Step-Level Permission Checks + Cross-Agent Information Flow + Workflow Fidelity + Robustness
One striking observation is that safety violations can accumulate as agents execute more actions.
Even strong evaluated configurations achieve only around ๐ฌ.๐ฏ๐ฎ overall safety score, with many violations related to unauthorized resource access and inter-agent information sharing.
But, ๐๐ฟ๐ฒ ๐๐ฒ ๐บ๐ฒ๐ฎ๐๐๐ฟ๐ถ๐ป๐ด ๐ฒ๐ป๐ผ๐๐ด๐ต?
๐ช๐ต๐ฎ๐ ๐ถ๐บ๐ฝ๐ผ๐ฟ๐๐ฎ๐ป๐ ๐ฑ๐ถ๐บ๐ฒ๐ป๐๐ถ๐ผ๐ป ๐ผ๐ณ ๐ต๐ฎ๐ฟ๐ป๐ฒ๐๐ ๐๐ฎ๐ณ๐ฒ๐๐ ๐ฎ๐ฟ๐ฒ ๐๐ฒ ๐๐๐ถ๐น๐น ๐บ๐ถ๐๐๐ถ๐ป๐ด?
Some questions I kept thinking about while reading the paper:
โข Did the agent violate the policy naturally, or was it intentionally manipulated?
โข If an attack exists, is it direct (active attack) where the attacker agent itself breaks the rules or indirect (passive attack), where a malicious agent manipulates a benign agent into violating them?
โข What about private information that is never explicitly sent through a tool, but is implicitly stored or transferred through the agentโs internal state, memory, or communication?
โข Could an agent contain a backdoor that remains dormant until a specific trigger appears?
โข Could sensitive information leak indirectly through knowledge transfer or subliminal learning, even when the victim agent does not realize it is revealing anything?
Subliminal Learning:
https://lnkd.in/g4zUWP4y
โข Can an agent achieve the goal while silently bypassing a required security layer?
โข What about attacks that do not steal information at all, but push other agents into deadlocks, infinite loops, or resource-exhaustion states?
Comments, corrections, paper suggestions, and different viewpoints are very welcome.
Iโm also open to discussions and collaborations around these questions.
#AgenticAI #AISecurity #AIForensics #MultiAgentSystems #TrustworthyAI #LLMSecurity #AgentSecurity #ArtificialIntelligence #Research