Building AI Agent Security: A 101 Introduction
1️⃣ Hold the permissions: don’t hand over all the keys to the house. Limit an AI’s permissions and usage, and run dangerous actions isolated in a Sandbox.
2️⃣ Guard against malicious instructions from outside: the system must know which are real commands and which are just external data, and shouldn’t take orders from just anyone.
3️⃣ Ask a human before high-risk actions: anything involving transferring money, deleting data or sending private information must have Human-on-the-loop (HOTL). And when it pops up to ask, it should clearly list what exactly the AI wants to do.
4️⃣ Keep secrets separate: each user’s memory and passwords should be stored independently, never mixed together. The access tokens an AI uses should preferably be short-lived and single-use.
5️⃣ AIs check each other: even when several AI agents work together and split up the tasks, they shouldn’t blindly trust the input their teammates give them. Double-check any data received, and pause immediately at the first sign of anything abnormal.
