AWS’ Deception Benchmark tests whether AI models can tell real security vulnerabilities from safe code that looks risky.