Researchers achieved 61-67% accuracy detecting code flaws by analysing AI model activations, surpassing prompted responses.