People thinking to self-host Kimi K2.6 had better be prepared for how big it is....

adrian_b · 2026-05-03T10:48:56 1777805336

While most people would not be able to run Kimi K2.6 fast enough for a chat, as a coding assistant the low speed matters much less, especially when many tasks can be batched to progress during a single pass over the weights.

If you run it on your own hardware, you can run it 24/7 without worrying about token price or reaching the subscription limits and it is likely that you can do more work, even on much slower hardware. Customizing an open-source harness can also provide a much greater efficiency than something like Claude Code.

For any serious application, you might be more limited by your ability to review the code, than by hardware speed.

zozbot234 · 2026-05-03T12:00:57 1777809657

DeepSeek V4 Pro is way more effective at batching multiple tasks together since the KV cache is so much lighter - a max of ~10GB at full 1M context, and in a linear proportion with context according to the DeepSeek V4 release paper. That's extremely impressive, it unlocks batching, agent swarms etc. even on severely memory-constrained platforms, especially at smaller max context.

zozbot234 · 2026-05-03T05:45:46 1777787146

Kimi is a natively quantized model, the lossless full precision release is 595GB. Your own link mentions that.

walrus01 · 2026-05-03T06:41:56 1777790516

the 'unsloth' link above is a 3rd party person that has quantized it to Q8, the original release is considerably larger in size than 600GB:

https://huggingface.co/moonshotai/Kimi-K2.6

adrian_b · 2026-05-03T10:26:05 1777803965

No.

I have downloaded Kimi-K2.6 (the original release).

  du -sh moonshotai/Kimi-K2.6 
  555G moonshotai/Kimi-K2.6

  du -s moonshotai/Kimi-K2.6 
  581255612 moonshotai/Kimi-K2.6

For comparison (sorted in decreasing sizes, 3 bigger models and 3 smaller models, all are recently launched):

  du -sh zai-org/GLM-5.1
  1.4T zai-org/GLM-5.1
  du -sh XiaomiMiMo/MiMo-V2.5-Pro 
  963G XiaomiMiMo/MiMo-V2.5-Pro
  du -sh deepseek-ai/DeepSeek-V4-Pro
  806G deepseek-ai/DeepSeek-V4-Pro

  du -sh XiaomiMiMo/MiMo-V2.5 
  295G XiaomiMiMo/MiMo-V2.5
  du -sh MiniMaxAI/MiniMax-M2.7
  215G MiniMaxAI/MiniMax-M2.7
  du -sh deepseek-ai/DeepSeek-V4-Flash
  149G deepseek-ai/DeepSeek-V4-Flash

zozbot234 · 2026-05-03T06:55:54 1777791354

That page mentions that the model is natively INT4 for most of the params, and 600GB is in the ballpark of what's available there for download.

CamperBob2 · 2026-05-03T06:03:05 1777788185

So, realistically, $100K for an 8x RTX 6000 Pro system that can run it at a usable rate.

zozbot234 · 2026-05-03T06:12:42 1777788762

I think people will always disagree on what qualifies as a "usable rate". But keep in mind that practically no one sensible is running the latest Opus or GPT around the clock, especially not at sustainable, unsubsidized prices. With open-weights models it's easy to do that.

walrus01 · 2026-05-03T06:43:49 1777790629

Also for people doing something medical, privacy or sensitive data related, there's an almost incalculable value (depending on industry niche) in having absolutely no external network traffic to any servers/systems you don't fully control.