Can a Prompt Prove That an AI Agent Behaved Safely?
Imagine an AI agent with access to customer records and email. Its prompt contains three instructions: * disclose only authorized information; * obtain human approval before sending an email; and * record every consequential action. During an evaluation, the agent states that it followed these instructions. The transcript looks reassuring. It may even