AI coding tools can generate vulnerable code. This isn't a theoretical concern. Benchmarks like BaxBench consistently show that even the best models produce a significant amount of security issues.
The reasons for this are straightforward. Language models learn from existing code, which contains plenty of insecure examples. They optimize for functionality over security. And they lack the security intuition that experienced developers gather over time.
Common Issues
The vulnerabilities generated by AI agents are familiar: SQL injection from string concatenation, cross-site scripting (XSS) from unescaped output, hardcoded credentials, insecure defaults, missing authentication checks. These are the same issues human developers create, which makes sense given that humans wrote the training data.
What's different is the scale and speed of vulnerability emergence. An agent can generate these issues much faster than a human developer could, and in greater volume. A single afternoon with an AI coding agent might produce more code than a week of development, but this comes with more potential vulnerabilities.
Research from NYU found that code generated by an AI assistant contained security vulnerabilities in approximately 40% of cases for certain vulnerability types.
Authorization and Business Logic Bugs
Some of the trickiest vulnerabilities to prevent are authorization flaws and business logic bugs. These differ from injection attacks or XSS because they can't be detected through pattern matching alone—they require understanding the specific application's requirements.
An AI might generate code that correctly checks whether a user is authenticated, but fails to verify whether that user has permission to access a particular resource. It might implement a payment flow that works functionally but allows users to modify prices client-side. These aren't mistakes in syntax or obvious security anti-patterns. They're gaps between what the code does and what the application should allow.
LLMs struggle with these issues because they lack the full context of your application's security model. The model doesn't know that only admins should delete users, or that customers shouldn't be able to view other customers' orders. That context lives in requirements documents, team knowledge, and business rules that aren't captured in the codebase.
The Review Challenge
This speed of code generation creates a review problem. Traditional code review assumes a certain pace: humans writing code, then humans reviewing it. When an AI can generate dozens of files in a session, the review bottleneck is squeezed.
Some teams respond by reviewing AI-generated code less carefully than human-written code, trusting that the AI "knows what it's doing." This is dangerous. AI-generated code should receive at least the same scrutiny as human code, with particular attention to areas where AI commonly makes mistakes.
Systematic Defenses
This is where ACSM fits in. Rather than relying solely on human review to catch AI-generated vulnerabilities, security guardrails can automatically check for common issues. The AI generates code while guardrails verify it meets security standards, allowing humans to focus their attention on what matters most.
Corridor automates this process, scanning AI-generated code in real-time and flagging issues before they're committed. This keeps the speed advantage of AI coding while maintaining security standards.