- إعلان رعاية رئيسي -إعلان

Kimi K3: How Moonshot AI Built the World’s Largest Open-Weight AI Model

    Moonshot AI has unveiled Kimi K3, the world's largest open-weight AI model with 2.8 trillion parameters. Rather than relying solely on massive computing power, the model uses a new architecture that reduces computational demands while requiring significant memory and infrastructure, raising questions about how practical it will be for enterprise deployment.

Moonshot AI has introduced Kimi K3, a new open-weight large language model that has quickly attracted global attention. With 2.8 trillion parameters, it is the largest open-weight AI model announced to date, surpassing previous Chinese open models in scale.

However, Kimi K3 is more than just a bigger model. Its architecture reflects a different approach to building advanced AI under hardware constraints. Instead of relying entirely on more computing power, Moonshot AI redesigned the model to reduce computation while making better use of available memory.

A Major Leap in Model Size

Kimi K3 represents a significant jump from Moonshot AI’s previous generation, which featured just over one trillion parameters.

The new model also exceeds DeepSeek V4 Pro, which contains around 1.6 trillion parameters, making Kimi K3 the largest open-weight model currently available.

Despite its enormous size, Moonshot says the goal was not simply to build a larger model but to improve efficiency during inference.

How Kimi K3 Works

Instead of activating the entire model every time it generates text, Kimi K3 uses a Mixture of Experts (MoE) architecture.

The model is divided into 896 specialized experts, but only 16 experts are activated for each token.

This approach dramatically reduces the amount of computation required while maintaining strong performance.

The trade-off is that all of the model’s parameters must remain available in memory, creating much higher memory requirements than smaller AI models.

Reducing Memory Requirements

To make deployment more practical, Moonshot AI trained Kimi K3 using four-bit precision instead of the traditional sixteen-bit format.

This compression reduces the model’s storage requirements from an estimated 5.6TB to roughly 1.4TB.

Although this is a substantial reduction, the model still requires far more memory than a typical enterprise server can provide.

Built for Large-Scale Infrastructure

Moonshot recommends running Kimi K3 across 64 or more AI accelerators connected together as a single memory pool.

Rather than depending on one extremely powerful processor, the model distributes its workload across many connected accelerators.

This design allows organizations with large AI infrastructure to deploy the model more efficiently, but it also increases deployment costs.

Can Businesses Run Kimi K3?

Many enterprises are interested in open-weight AI models because they offer greater control over sensitive data and allow organizations to customize models for their own needs.

However, Kimi K3’s hardware requirements make local deployment challenging.

The model alone requires approximately 1.4TB of memory, before accounting for additional resources needed for long documents, multiple users, and production workloads.

As a result, most organizations are expected to rely on dedicated cloud infrastructure instead of hosting the model entirely on their own servers.

Software Support Is Still Developing

Hardware is not the only challenge.

Because Kimi K3 introduces new architectural techniques, many open-source inference tools do not yet fully support the model.

Moonshot AI says it is working with infrastructure partners and the open-source community to improve compatibility before wider deployment.

This means organizations should expect some delay between the public release of the model weights and full production readiness.

Pricing

Moonshot AI prices Kimi K3 at $3 per million input tokens, dropping to $0.30 when cached inputs are reused.

Output tokens cost $15 per million.

While these prices remain competitive compared with several premium AI models, they are higher than some Chinese open-weight alternatives.

Organizations should therefore evaluate the total cost of completing AI tasks rather than comparing token prices alone.

Performance and Current Limitations

Kimi K3 achieved first place in Arena’s Frontend Code benchmark, outperforming several competing models in coding evaluations.

However, Moonshot acknowledges that Kimi K3 still trails leading proprietary models such as Claude Fable 5 and GPT 5.6 Sol in overall performance.

The company also notes that the model may occasionally generate inconsistent results or make unexpected decisions when user instructions are ambiguous.

Most published benchmark results remain based on Moonshot’s own testing and will require independent verification after the model weights become publicly available.

What Kimi K3 Means for AI

Kimi K3 demonstrates that innovation in AI architecture can continue even under hardware restrictions.

Rather than eliminating infrastructure challenges, Moonshot AI has shifted them from raw computing power to memory capacity and large-scale hardware deployment.

For enterprises, the model offers greater flexibility through open weights but also demands significant investment in infrastructure.

Whether Kimi K3 becomes widely adopted will depend not only on its performance but also on how easily organizations can deploy and operate it in real-world environments.

Read the article in Arabic

- إعلان -إعلان داخلي
Share
Copy link

Report an issue

نورهان فؤاد

كاتبة محتوى متخصصة، تجمع بين السلاسة والأسلوب الصحفي، تساهم في صياغة مقالات ريادة الأعمال والشركات الناشئة بأسلوب جذّاب وسهل الفهم

Report An Error

You are now reporting an error in the article: Kimi K3: How Moonshot AI Built the World’s Largest Open-Weight AI Model

For Media Partnership