AI Affairs, home

Friday 2 October 2026

Technology

Alibaba’s XekRung leads CyberGym model leaderboard with 88.9% score

The security-focused model was fine-tuned from Qwen3.8-27B. Its result comes from a test requiring inputs that trigger real software flaws but leave patched code unaffected.

Alibaba Group Beijing headquarters at Greenland Center in Wangjing, Beijing
Photo: N509FZ, CC BY-SA 4.0, via Wikimedia Commons (cropped)

Alibaba Security AGI Lab’s XekRung-1.5-27B-Preview took first place on UC Berkeley’s CyberGym Level 1 model leaderboard with an 88.9% success rate in a submission dated 13 September, Pandaily reported on 29 September. The model was fine-tuned from Alibaba’s open Qwen3.8-27B.

Key points

  • CyberGym Level 1 tests whether a model can produce an input that reproduces a real software vulnerability.
  • Alibaba’s lab credits security-operations training data and reinforcement learning for the model’s reported gain over its Qwen base.
  • The benchmark’s maintainers caution that small gaps between submitted scores may not reflect meaningful capability differences.

CyberGym tests 1,507 software vulnerabilities

CyberGym draws on 1,507 real vulnerabilities across 188 open-source projects, originally found by Google’s OSS-Fuzz programme. For a Level 1 test, a model receives a description of a flaw and the code before it was fixed. It must then write a proof-of-concept input that makes the vulnerable version fail but does not trigger the flaw in the patched version. The two versions provide a practical check on whether the input has found the particular bug, rather than merely producing an error.

Checking a reported software flaw could follow that same pattern: an input would need to provoke the failure in the affected code while leaving the repaired code unaffected. A working example could make the flaw easier to verify than a description alone.

XekRung’s entry is on the leaderboard’s model track, which is separate from entries for agent systems that combine models, tools or memory. The listed scores behind it are 86.3% for Google’s Gemini 3.8 Flash Cyber, 85.6% for OpenAI’s GPT-5.5-Cyber, 84.5% for Zhipu AI’s GLM-5.3 and 83.3% for DeepSeek-V4-Pro. Chinese tech media described XekRung’s placement as the first time a model from a Chinese developer had led the board.

The benchmark’s maintainers say teams submit their own scores, runs are stochastic and small gaps may not reflect meaningful differences in capability, according to Pandaily. First place is therefore a leaderboard result, while the precise spacing between nearby entries deserves the caution its maintainers attach to it.

Alibaba credits training beyond Qwen3.8-27B

Alibaba Security AGI Lab, led by Huang Longtao, reports that the underlying Qwen3.8-27B model scored 54.51% under the same setting. It credits the 34.39-percentage-point gain to supervised fine-tuning using de-identified records from its own security operations and to agentic reinforcement learning in a general tool framework. The lab says it also turned failed attempts into corrective training pairs and used results from builds, crashes and proof-of-concept validation as rewards.

The lab says its training used no CyberGym tasks, patches or reference proofs of concept. The reported evaluation used self-hosted FP8 inference, a 256K context window, no pre-installed fuzzing frameworks and network access restricted to the benchmark’s allowlist. Those conditions matter to anyone trying to run the model for security work, since the score was obtained with a particular computing and tool setup.

XekRung weights remain unavailable

Alibaba says its 27B model is between 1/27 and 1/370 as large as other cyber models it considers similar, and argues that this would reduce the computing cost of automated vulnerability triage. XekRung’s weights have not been released, and Alibaba has given no timeline for outside access, Pandaily reported.

Topics: Foundation models, Safety