AI News Feed
Market watch
Cybersecurity

Cantina Releases Open-Weights apex-flash-1 for Vulnerability Research

Cantina Security and Yeta Labs released apex-flash-1, a 321.3B open-weights vulnerability research model under MIT. It solved 40 of 60 held-out bug tasks in company-reported tests.

apex-flash-1 has 321.3B total parameters, according to its Hugging Face safetensors metadata. The GLM-5.3-Flash base is a Mixture-of-Experts model with 18B active parameters. Cantina trained it with GRPO using a rank-256 LoRA plus selective full-parameter training. The MIT weights can be served on vLLM, SGLang or Transformers, but BF16 deployment requires roughly 640 GB of GPU memory. The comparison table lists about 642 GB for BF16 weights, and community 4-bit ports exist.

The training data covers 150 tasks built from 50 real vulnerability cases. Each case appears in three variants: guided whitebox, focused whitebox and focused blackbox. Authorization, identity and scope flaws make up 72% of the cases. Accounting and numerical precision bugs add 18%, while time validation, business rules and SSRF cover the rest. According to the model card, RL rollouts ran inside the Codex agent harness on production-like software and protocol environments.

Cantina evaluated 60 tasks from 20 held-out vulnerability cases. Each model ran the set once, with costs estimated from provider pricing. apex-flash-1 solved 40 of 60 tasks, a 66.7% pass@1 rate, at about $2.38. GLM-5.3-Flash solved 36 of 60, or 60.0%, at about $4.56. Claude Opus 5 High solved 43 of 60, or 71.7%, at about $74.68. Opus solved three more tasks but cost about 31 times more per run. That is roughly $0.06 per solved task for apex-flash-1 versus $1.74 for Opus. These are company-reported numbers on an internal benchmark.

Cantina positions apex-flash-1 as a worker model orchestrated by a larger model. The model card lists code reading, tool use, exploit development and verification as target skills. An experimental apex-flash-1-abliterated variant ships with modified refusal behavior; it was not separately evaluated. Cantina's rationale is that defenders need capable models they can run and control locally.

The model card compares apex-flash-1 with Aikido Altar-1, Cisco Foundation-Sec-8B-Reasoning and GLM-5.3-Flash. Aikido Altar-1 is described as using expert pruning and quantization for air-gapped autonomous pentesting. Cisco Foundation-Sec-8B-Reasoning is described as instruction-tuned and RLHF-trained on security QA for SOC triage and threat defense. GLM-5.3-Flash is listed as a general-purpose base for general coding and agents.