Ask HN: Which frontier model can do code security reviews

I can't complain about Opus 5.5, more than extracted my money's worth out of it but the final stages require pen-testing and Claude's masters stomp on it every time it tries to help do security reviews [cyber]. Ever so often it sneaks in a security fix. Even found a race condition in haproxy by mistake and Claude took it upon itself to find the crash string. That went horribly bad. Each time they try to up-sell Mythos and say I have to go through a verification program that I am not permitted to go through.

Aside from the uncensored Qwen forks, which frontier models can do extensive code security reviews, security fixes? Ideally something close to the quality of the NCC Group. This is for my own hobby craft. Maybe this does not exist and that is fine too.

4 points | by Bender 50 minutes ago

1 comments

  • bigyabai 35 minutes ago
    GLM 5.3. Anthropic even made the mistake of comparing it to Mythos (lol): https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

    > Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.

    > We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.

    • Bender 20 minutes ago
      I may end up going that direction. I would ideally like to find something that is purpose built to do code pen-testing so I do not have to bypass anything. There are forks of other models built for this, maybe there is a fork of GLM too.