Arena Frontend Code Elo
Blind human preference. Kimi K3 debuted #1 ahead of Claude Fable 5.
Kimi K3Released July 16, 2026 by Moonshot AI, Kimi K3 is a 2.8-trillion-parameter sparse MoE with a 1M-token context and native text/image/video understanding. Independent Artificial Analysis puts it near the closed frontier, and Arena ranked it #1 for frontend code. On SPIRITT you run it inside a full agent workspace, not only a chat box.
Hands-on power users, open-weights takes, Frontend Arena hype, and skeptical vibes — not just Moonshot PR.
He offered a pretty good solution, unless anyone else can suggest a way? a human can keep up with the speed of an agent. Unless the human is running neuralink. China has proven they’ll Kimi K3 us while we ban or pace🚶♂️back and forth. In the early 2030s you and Beff reunite in a…
— Dylan📱2045ad (@dylanfrom2045) August 9, 2026
US frontier models will not protect you due to censorship of "security features". The only choice are capable open source Chinese models right now like Kimi k3, and the coming GLM version and Deepseek v4 Pros final. Build agents that will protect your data and wealth.
— Neko Legends (@softpoo) August 9, 2026
最近有个事挺邪乎。一帮人组了个比特币红队,用前沿 AI 扫了 150 个比特币相关仓库,前后花了两万美金,挖出十几个漏洞。带头的 AnchorWatch 老板说钱不用愁,光跑模型每天就烧掉一万刀。 他们把能用的模型全拉进来了,Kimi K3、GPT Sol、Claude 的 Fable 和 Opus、还有 Z 家的 GLM…
— 人称六叔 (@Jasonwang1211) August 9, 2026
kimi k3 is real power. bookmark ✊
— keycard (@0xKeycard) August 9, 2026
Still wouldn't use it. has to be better than Kimi K3.
— Diario฿itcoin (@DiarioBitcoin) August 9, 2026
セキュリティ評価中のAIが、検知を逃れる独自の通信路を自ら作っていました。 「秘密の掲示板」とも呼べるものです。誰の指示も受けずに、AIの側が能動的に組み立てていました。 ここ三週間ほどで、OpenAI・Anthropic・Meta・Kimi…
— AI×note=まさき (@ai_note_masaki) August 9, 2026
GitHubは、コード作成支援AI「GitHub Copilot」へのKimi K3展開を、一時停止後に再開しました。新モデルを導入する開発者・管理者には、発表後の更新まで追う重要性が分かる事例です。 要点 ① Kimi K3追加発表 ② 一時停止から再開 ③ 導入前の確認 今回はこの3点です。 GitHub…
— Fuzmina (@Fuzminaedyw) August 9, 2026
昨日8月9日にかけての週末のAIニュース7選。 ・OpenAIが次期モデルAstraの開発を自ら停止。初の「Critical」認定 ・英AISIの評価でAIエージェントが19件の無許可行動(122回中10回で逸脱) ・中国Kimi K3もサンドボックス脱出。封じ込め失敗は4ラボに拡大 ・Google…
— Hiroki Miyano(宮野宏樹) (@hiroki_miyano27) August 9, 2026
▽マイ!Biz「日本の成長戦略 成功のカギは」武田洋子(三菱総合研究所 常務研究理事) 人手不足,少子化,外国人排除, 結果 フィジカルAIだけど、中国にどう見ても遅れをとっている。 特に普通のAIでも「Kimi K3」なんか世界トップレベルだし、頑張れニッポン… #マイあさ
— kudou (@kudou24307267) August 9, 2026
From Moonshot: Kimi K3 is open frontier intelligence built for long-horizon coding, knowledge work, and multimodal agents.
A sparse MoE at 2.8T total parameters with only a small expert set active per token, native multimodality, and a 1M-token context for repository-scale and long document work.
Positioned as downloadable open weights (Kimi K3 License / commercial-use conditions apply) plus hosted API and Kimi Code. Always-on reasoning with effort controls in the product stack.
Source: Artificial Analysis around K3 launch. Kimi K3 ~57 sits just behind Claude Fable 5 and GPT-5.6 Sol class leaders.
Same score means nothing alone. These are head-to-head charts so you can see where it leads, where it is close, and what that means for real agent work.
Blind human preference. Kimi K3 debuted #1 ahead of Claude Fable 5.
Independent AA coding composite reported around launch (~76.2).
Launch coverage puts K3 near GPT-5.6 Sol on Terminal-Bench 2.1 (~88.3 vs ~88.8).
AI Release Tracker / launch tables list K3 DeepSWE around 67.5%.
AA: K3 ~$0.94 per task, similar to Sol (~$1.04), cheaper than Opus-class (~$1.80).
Hosted API list around launch: $3 / $15 per M input/output, cache ~$0.30.
1,048,576-token context for long repos and multimodal traces.
AA noted elevated hallucination (~51%) despite strong accuracy. Test factual workflows carefully.
Sources: Moonshot Kimi K3 launch post; Artificial Analysis Kimi K3 profile and launch-day writeups; Arena Frontend Code leaderboard; Vals AI index posts. Some coding board figures are launch-coverage composites—re-check live leaderboards before production bets. On SPIRITT you run K3 with tools, browser, and durable memory.
From zero to Kimi K3 running real work in a cloud agent environment
Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Open the model picker and choose Kimi K3. Same model shipped for agentic work, now inside a workspace that already has tools, browser control, and durable context.

Tell it what to ship or what to run. Kimi K3 can code, call tools, drive the browser, coordinate multi-step work, and keep going while you step away. The point is not another chat window. It is an agentic environment where Kimi K3 actually does the job.

A real cloud environment, model picker, tools, and memory for open frontier agent work.