AI News Feed
Market watch
Large Language Models

Moonshot's Kimi K2.8 Preview Brings 1M-Token Context to Every Paid Tier as Hong Kong IPO Push Advances

Moonshot AI released Kimi K2.8 Preview on Sept 11, opening a 1M-token context window to all membership tiers, including the free one, and claiming performance close to flagship K3 with improved coding and agent abilities.

Kimi's release notes describe K2.8 Preview as close to K3 in overall performance, with coding and agent capabilities raised across the board and reasoning efficiency improved over K2.7 Code. It supports low, high and max reasoning effort levels aligned with K3's thinking tiers, accepts image and video input, and carries a 1M context window. No benchmark figures accompanied the preview, so the published material does not show how near the model comes to K3 in measurable terms.

Access is where the two models diverge. K3 requires at least the Moderato tier at 99 yuan a month, and its 1M context window needs Allegretto at 199 yuan or above, while K2.8 Preview is available on all five tiers, including the free Adagio plan, and supplies the full 1M context directly. The tier names — Adagio, Andante, Moderato, Allegretto and Allegro, priced at free, 49, 99, 199 and 699 yuan a month — were introduced after the K3 launch and follow musical tempo markings.

Existing traffic is being redirected rather than replaced by user choice. According to the documentation, kimi-for-coding has been upgraded to K2.8 Preview without users noticing, with the Model ID left unchanged. Inside Kimi Code, requests to K3 with thinking switched off are routed to the non-thinking version of K2.8 Preview. Kimi Work received the same preview model alongside the coding product.

Kimi positions K2.8 Preview for code completion and routine development work rather than headline benchmark runs. K3 was built the other way: a mixture-of-experts model with 2.8 trillion total parameters and roughly 104 billion active parameters, released in July with open weights, a long context window and flagship coding ability. It was priced at $3 per million input tokens and $15 per million output tokens, about a third of comparable models, and the company said it beat Claude Opus 4.8 and GPT-5.5 on coding benchmarks. On Arena.ai's front-end code arena, K3 topped the leaderboard with 1,679 points, ahead of Claude Fable 5 and GPT-5.6 Sol, after the predecessor K2.6 placed 18th. K3 series models now generate about 300 billion tokens a day, according to the company.

Even a flagship has limits. Moonshot's technical blog noted that K3 is sensitive to historical thinking content, so quality can become unstable if an agent framework fails to return previous reasoning traces or if a session switches to K3 mid-conversation. The company also said K3 can act too aggressively on ambiguous tasks and make decisions beyond what users expect. Those constraints leave room for a cheaper, more controllable model aimed at mass deployment.

Hiring K2.8 to carry volume fits Moonshot's revenue curve. Reported annualized recurring revenue moved from about $100 million in March to roughly $200 million in April, above $300 million in June and past $1 billion in August. Bloomberg reported the company aims internally for $2 billion by the end of the year.

The same report said Moonshot confidentially filed an A1 application with the Hong Kong stock exchange this month, formally starting an IPO process. Valuations have climbed steeply: $4.3 billion post-money in the Series C round at the end of 2025, $20 billion in May 2026, $35 billion after the Series F round in July, and $50 billion pre-money when a Pre-IPO round opened in late July — nearly an eightfold rise in about eight months. Against that $50 billion pre-money figure and the $300 million ARR then public, the price-to-sales multiple works out to about 167 times. For comparison, Anthropic trades at roughly 20 times, OpenAI near 40 times, and Zhipu at about 94 times when it reached a trillion-yuan valuation.