AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.…
Opening of the original on arXiv cs.AI
Summary
Researchers developed AgentProv, a framework to audit how large language model (LLM) API providers handle tool use policies. The system probes LLM agents to understand their adherence to specified rules, aiming to improve transparency and safety in agentic AI systems. This work addresses the need for verifiable compliance in the growing field of AI agents that interact with external tools and services.
Why it matters
Why it matters: AgentProv introduces a method to test LLM agent adherence to tool-use policies, a critical aspect of AI safety and reliability. This affects developers building AI agents and API providers like Google Gemini, Anthropic, and OpenAI. Current auditing methods are often manual or limited. AgentProv offers a more systematic approach. Future work should focus on scaling these probes and integrating them into continuous monitoring systems for deployed agents.
Related: OpenAI: Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study · Anthropic: AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes · Microsoft: Responsible AI in 2026: How we are adapting for what’s ahead · Meta: Trump Administration Sides With OpenAI in New York Times Copyright Lawsuit · xAI: Ajeya Cotra – "This might be the clearest warning shot we ever get"
Rated low: routine. Worth knowing, not worth rearranging your day for.
- SIR: Self-improving Red-teaming for Compute Use AgentsarXiv cs.AI · Web · Aug 30, 2026
- Claude's new system prompt really doesn't want to reproduce song lyricsSimon Willison · Web · Sep 2, 2026
- Black Box: The Chatbots | Spirals | Ep 1 – podcastThe Guardian AI · Web · Sep 3, 2026
- The Most Overhyped and Underhyped New AI ModelsMatt Wolfe · Web · Sep 2, 2026
- GPT-6 Astra Just Went CRITICAL...Wes Roth · Web · Sep 2, 2026
- LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, DronesLast Week in AI · Web · Aug 31, 2026
- When millions of AI agents meetGoogle DeepMind on YouTube
- Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being PostsarXiv cs.AI
- Piloting the world's first double-blind AI evaluationsGoogle DeepMind Blog
- AI is getting a little out of controlAI Explained
- Understanding the inner thoughts of AIGoogle DeepMind on YouTube
- Investing in multi-agent AI safety researchGoogle DeepMind Blog
Questions people ask
- What is AgentProv?
- AgentProv is a framework designed to audit how large language model (LLM) API providers adhere to their tool-use policies. It probes LLM agents to check their compliance with specified rules.
- What problem does AgentProv solve?
- It addresses the need for transparency and verifiable compliance in AI agent systems that use external tools. This helps ensure agents operate safely and according to defined policies.
More from arXiv cs.AI 6 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to arXiv cs.AI.



