AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.…
Opening of the original on arXiv cs.AI
Summary
Researchers developed AgentProv, a framework to audit Anthropic's Claude API for tool-use policies. AgentProv probes Claude's behavior when interacting with external tools, revealing potential policy violations. This work aims to ensure AI agents adhere to defined rules, enhancing safety and trustworthiness in their operations. The system tests how Claude handles specific tool requests and data exchanges, providing a method for verifying compliance.
Why it matters
Why it matters: AgentProv introduces a novel auditing method for AI agents, specifically targeting tool-use policies. This is crucial as AI models like Anthropic's Claude increasingly integrate with external services. Current auditing methods may not fully capture the nuances of agentic behavior. AgentProv offers a concrete way to test compliance, potentially influencing how API providers design and verify their safety mechanisms. Future work will likely focus on expanding these probes to cover more complex interactions and a wider range of policies across different AI providers.
Related: OpenAI: Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study · Google: AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes · Microsoft: Responsible AI in 2026: How we are adapting for what’s ahead · Meta: Trump Administration Sides With OpenAI in New York Times Copyright Lawsuit · xAI: Ajeya Cotra – "This might be the clearest warning shot we ever get"
Rated low: routine. Worth knowing, not worth rearranging your day for.
- Sam Altman :‘AGI in 2026’, just as Models Start to [Mis]Train ThemselvesAI Explained · Web · Aug 27, 2026
- Sam Altman "AGI by December"Wes Roth · Web · Aug 29, 2026
- Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model StudyarXiv cs.AI · Web · Sep 1, 2026
- Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented AgentsarXiv cs.AI · Web · Sep 1, 2026
- SIR: Self-improving Red-teaming for Compute Use AgentsarXiv cs.AI · Web · Aug 30, 2026
- Claude's new system prompt really doesn't want to reproduce song lyricsSimon Willison · Web · Sep 2, 2026
- Automated researchers can reliably mitigate alignment failuresAnthropic Research
- Model Hardware Standard: AI operating physical equipmentAnthropic on YouTube
- Sam Altman :‘AGI in 2026’, just as Models Start to [Mis]Train ThemselvesAI Explained
- Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being PostsarXiv cs.AI
- SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in AutoformalizationarXiv cs.AI
- Sam Altman "AGI by December"Wes Roth
Questions people ask
- What is AgentProv?
- AgentProv is a framework designed to audit AI agents, specifically probing their adherence to tool-use policies when interacting with APIs like Anthropic's Claude.
- What is the goal of AgentProv?
- The goal is to ensure AI agents comply with defined rules and policies during tool use, thereby enhancing safety and trustworthiness.
More from arXiv cs.AI 23 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to arXiv cs.AI.

