AI application security: control API resource use as well as attacks
Use OWASP resource-consumption risks to review account quotas, concurrency, timeouts and cost monitoring, and separate edge protection from application authorization.
Contents of this article
Valid requests can still exhaust resources
OWASP API Security Top 10 identifies unrestricted resource consumption as API4:2023; its LLM application risks also include unbounded consumption. An authenticated AI request may run inference for a long time, invoke paid tools or hold a connection open. Blocked-request counts alone do not show whether the service can operate sustainably.
Apply limits to the account that owns the workload
Track tenant, account, API key and entry point together. IP rate limits provide one layer, but shared networks group legitimate users and distributed traffic spans many addresses. At the application layer, enforce identity-based request rates, concurrent jobs and model permissions, with server-side quota checks rather than client-reported balances.
Give long-running requests a defined end
For streaming APIs, define first-response timeout, maximum generation time, output size and cancellation after client disconnect. Closing a page does not necessarily stop the backend task or upstream charges. Bound queue length and job age, and return a clear overload response instead of hiding failures behind an endless queue.
Divide responsibility between the edge and application
The edge handles abusive traffic, common web attacks and ingress load. The application owns object authorization, model quotas, tool permissions and sensitive-data access. A WAF cannot replace those rules. After enabling protection, verify authentication headers, streaming, timeouts and retries with real clients; avoid browser challenges that API clients cannot complete.
Run a controlled check before release
In a test environment, cover quota exhaustion, excess concurrency, user cancellation and upstream timeout. Check that tasks stop promptly, costs can be attributed to accounts and rejected requests receive consistent responses. Monitor success rate, generation latency, active connections and cost per business operation together.
