Bilibili Launches 'AI Infinite Arena' as GPT-6 Astra Tops First Model Leaderboard
Bilibili launched the AI Infinite Arena on Sept. 16, publishing a first leaderboard drawn from creator-run tests of more than 100 models. GPT-6 Astra finished first, while three Chinese models placed in the top five.
GPT-6 Astra won first place in 10 evaluations by Bilibili creators, known on the platform as UP hosts, beating GLM-5.3 for the most first-place finishes, according to QbitAI, which republished the announcement from Bilibili. Three of the top five models were Chinese.
The arena collects model reviews produced by UP hosts across a range of fields. It covers more than 100 large models tested in real scenarios, spanning themes such as coding, reasoning, collaboration and knowledge.
Unlike conventional benchmark scoring, the arena imposes no restriction on subject matter or evaluation dimensions. Creators from different sections of the site devise their own tasks, including real work flows, professional applications and playful ideas, and the models compete on the same prompts so that differences in practical performance become visible. The first leaderboard includes comparative tests of DeepSeek, Kimi, ChatGPT, Claude, Gemini, Doubao, Qwen, Hy, MiniMax and other mainstream models.
Bilibili said its AI ecosystem grew sharply over the past year, with time spent watching AI knowledge content rising 72 percent year on year and more than 190 million users viewing AI content in an average month. Live streams, videos and bullet-screen comments suit the way AI enthusiasts seek information and exchange views, and model reviews have accumulated in the community across many fields, dimensions and topics, the company said, adding that the arena was created as a gathering place for that output.
The rankings will be updated in real time and remain open for UP hosts to sign up. As more evaluation cases are added, Bilibili said, the real-scenario samples will test model capabilities more thoroughly and give users a broader view of where the limits of AI performance lie.