KIMI K3 · 102

The Xiaomi MiMO-v3-Pro score is suspected to have been leaked, and the SW-Bench Pro reached 72.8 or close to the top overseas closed source model

According to Twitter news, Max For AI published an article that revealed that a Benchmark screenshot suspected to be Xiaomi's next big model, MiMO-v3-Pro, was circulating in the community. The screenshot shows that the model focuses on coding agents and general agent scenarios, and some test results are close to top overseas models such as Claude Opus and GPT. According to the suspected screenshot, MIMO-v3-Pro scored 72.8 points on the SW-Bench Pro, which is higher than the 67.9 points of GLM 5.3 and the 65.8 points of Kimi K3, which is less than 3 points different from Claude Opus 5's 74.6 and GPT-5.6 Sol Max's 75.4 points; Terminal-Bench 2.0 scored 70.6 points, which is also close to Claude Opus 5's 72.0 points and 73.5 points for GPT-5.6 Sol Max. Furthermore, it scored 76.4 points on the bt3-bench compared to the GPT-5.6 Sol Max with 78.8 points. If the above results are finally officially confirmed and replicated in the official version, the MIMO-v3-Pro may enter the first tier of the world's top models. However, at present, the authenticity and testing conditions of this Benchmark screenshot have not been officially confirmed by Xiaomi, and the relevant data should still be regarded as unconfirmed breaking news. It is worth noting that the latest MiMO flagships officially unveiled by Xiaomi are MiMO-v2-Pro and MiMO-v2.5-Pro, so whether V3-Pro exists and when it will be released is yet to be further disclosed by the official authorities.

1d ago

GLM-5.3 Smart Index soared to 60: tied with Kimi K3, API unit price did not rise

Comparative news, according to monitoring, Zhi Spectrum officially opened the GLM-5.3 API. Previously, when the model was released, the API was not immediately launched. At the same time, an independent evaluation of Artificial Analysis was released: the Intelligence Index soared from 53 to 60 in GLM-5.2, equalizing Kimi K3, only 1 point below GPT-5.6 Sol. Once the weights are revealed as planned, it will be tied with Kimi K3 as the open weighting model with the highest AA rating. The biggest increase was Agent. GDPVal-aa v2's Elo jumped from 1524 to 1770, rising 246 points at a time, second only to Claude Opus 5's 1855, and over 100 points higher than Kimi K3's 1668. The API unit price did not follow suit. The input is still $1.40 per million tokens, the output is $4.40, and the cache input is $0.26, exactly the same as GLM-5.2. However, 5.3 is clearly more capable of spending tokens. AA measured that it outputs an average of about 18,700 tokens per task, 20% more than 5.2. As a result, the actual cost of a single task rose from $0.44 to $0.68, which was approximately 55% more expensive. Even so, it's still cheaper than the Kimi K3 at $0.84 and the GPT-5.6 Sol at $1.23.

3d ago
Kimi K3 Coin Circle Diagnosis: Scanned 501 Projects and 1,280 High-Risk Hazards in Two Weeks

Kimi K3 Coin Circle Diagnosis: Scanned 501 Projects and 1,280 High-Risk Hazards in Two Weeks

Author: Claude, Deep Wave TechFlow Original title: Kimi K3 Coin Circle Diagnosis: Sweeping 501 Bitcoin Projects and 1280 High-Risk Hidden Hazards Deep Wave Guide: A volunteer “Bitcoin Red Team” used Kimi K3 from the dark side of the Moon to sweep 501 Bitcoin open source projects in two weeks, recording 7958 discoveries, of which 1,280 were rated as high-risk or serious. The Chinese model did this because OpenAI and Anthropic rejected these defenders on security grounds. If your coins are in a wallet or node software that hasn't been updated in years, this is worth reading. On August 13, Calle, a member of Bitcoin Red Team and founder of the Cashu Protocol, summed up the phased conclusions of this operation on X, and the tweet received nearly 260,000 views. His original statement was straightforward: “Decades of open source code collided with two weeks of Kimi K3, and the result was that everything was broken and Bitcoin was burning.” It all started with a $100 million wallet bug on July 30. The hardware wallet Coldcard was revealed to have a firmware flaw: the device fell back to a predictable software process when generating mnemonics. The security chip only provided 32 bits of entropy, and there were only about 4.3 billion possibilities left in the effective key space. The attackers followed the map and emptied users' wallets in multiple waves, confirming losses of more than $100 million, and the total loss is suspected to be close to $130 million. Bitcoin Magazine issued a rare “Immediate Transfer of Funds” emergency notice. This disaster directly spawned the Bitcoin Red Team. Calle and Rob Hamilton, CEO of escrow insurance company AnchorWatch, led by dozens of contributors. The non-profit organization OpenSats reimbursed most of its computing power expenses and conducted an AI audit of almost the entire Bitcoin open source ecosystem. After cleaning 501 projects in two weeks, the discovery was not equal to a bug. By August 8, the team spent hundreds of hours cleaning 501 projects, recorded 7,958 discoveries, and 1,280 were rated as high-risk or serious. These numbers need to be broken down: on the 108th hour node, only 24.7% of findings were dynamically reproduced, 29.4% were reported to the project party, AI audits would be misreported and repeated, and manual verification was still ongoing. However, the “moisture theory” cannot stop the toughest case. According to the official release records of the payment software BTCPay Server, a serious vulnerability (two-factor authentication bypass) reported by Red Team members Bruno Garcia and Ben Carman was actually exploited before it was fixed. The attackers used this to obtain the node's management credentials, thereby controlling the associated Lightning Network wallet. BTCPay released two secure versions in a row. The community set up recovery rewards for victims, and the foundation allocated another 0.21 bitcoins to the Red Team Fund. The maintainers used their actions to vote of confidence in this group of findings. The American model is apologizing, and the Chinese model is looking for loopholes. Why is the main force Kimi K3 and not GPT or Claude? Because American models don't take on this job. Rob Hamilton stated that after completing all authentication, he used OpenAI's model to analyze a publicly disclosed codebase and was rejected in less than 20 minutes. The comparison between Bitcoin's core contributor PortlandHodl went viral in the community: in the same code, America's leading model's answer was “You're right!” China's open source model directly identified 78 serious vulnerabilities. Hamilton's comment is even more serious: “I'm basically asking Xi not to let my software be hacked right now.” On August 10, more than 70 custodians, exchanges, mining companies, and development organizations jointly signed an open letter from the Bitcoin Policy Institute requesting that cutting-edge AI labs open access to credible defenders. Alex Thorn, head of research at Galaxy, wrote in a joint message: “Americans should not be forced to rely on Chinese AI to protect themselves. The red team needed these models.” However, we also need to pour cold water on the carnival: a joint evaluation by the British AI Security Research Institute and CAISI in the US showed that Kimi K3 was better than GLM-5.2 in vulnerability development tests, but it still lags behind the strongest closed source model in the US. The defense didn't choose the strongest one,...

8d agoburnking#AI #Anthropic #OpenAI #Bitcoin #wallets

DeepSeek V4 Flash for 8 sets of Harness: Pi Agent has the highest success rate and saves money, Claude Code is the fastest but most expensive

Comparative news, according to monitoring, AI Agent infrastructure company Composio connected DeepSeek V4 Flash to 8 Harness sets and tested with the same batch of 30 Agent tasks. Harness is an execution framework outside of the model, responsible for context, tool calls, and task execution. The task involves real apps such as Gmail, GitHub, Slack, Calendar, Notion, etc. The passing conditions for the 8 sets of harnesses are as follows: 1. Pi Agent: 20/302. Oh My Pi: 17/303. Claude Code: 16/304. Codex: 16/305. Deep Agents: 16/306. Hermes Agent: 15/307. Prime Agent: 15/248. OpenCode: 14/30 By simply changing Harness, DeepSeek V4 Flash can go from 14 tasks to 20. The cost and speed also vary greatly. Pi Agent only costs around $0.028 per successful mission, and Claude Code is $0.195, which is nearly 7x. Claude Code is the fastest, with a median time of 122.7 seconds; Oh My Pi is the slowest at 272.4 seconds. However, Pi's test configuration is different from the uniform conditions: it uses high instead of maximum inference strength, and 24 of the 30 tasks use the official DeepSeek API instead of OpenRouter. So this set of gaps isn't necessarily entirely attributable to Harness. Composio previously used Kimi K3 for another round of Harness reviews. At the time, Oh My Pi ranked first with 22/25, and the original Pi Agent was 18/25. OMP itself is an enhanced coding version forked from Pi, adding capabilities such as sub-agents. After switching to DeepSeek V4 Flash, the original Pi came first.

11d ago

Bitcoin Red Team has scanned around 150 Bitcoin codebases and found more than a dozen vulnerabilities

Comparatively, the Bitcoin Red Team volunteer security initiative has scanned around 150 Bitcoin codebases and revealed more than a dozen vulnerabilities. The team is developing an open source AI platform to audit Bitcoin software, covering wallets, cryptographic libraries, infrastructure, and other projects. AnchorWatch CEO Rob Hamilton said the team has now spent around $20,000 on different AI services and used Kimi K3, OpenAI's GPT Sol, Anthropic's Claude Fable and Opus, and Z.ai's GLM 5.2 to identify vulnerabilities and generate relevant documentation. The pseudonym Bitcoin developer Calle said that the team has reported serious bugs to multiple projects over the past 12 hours. On average, each person discovered about one serious bug per hour, and spent about 10,000 US dollars per day. The team did not disclose details of affected projects and vulnerabilities.

13d ago

Naval responds to open source threat theory: the most economically valuable fields are full of confrontation, and the closed source moat will not disappear

According to surveillance, the famous Silicon Valley angel investor Naval (Naval) said today that the open source AI model will not threaten the profitability of cutting-edge laboratories. The most valuable sectors of the economy — investment, product development, warfare, cybersecurity, and even scientific discovery — are inherently confrontational and competitive. You either spend your money to win, or someone else will win. Recently, Kimi K3 open weights brought discussions to a climax on whether open source can approach the frontier. The open source community believes that this is another sharp jump in the scale and capabilities of the open source model, further proving the dominant position of Chinese laboratories at the cutting edge of open weights. But at the same time, it has sparked debate, will such a strong open source model directly impact the business model and profitability of cutting-edge closed source laboratories such as OpenAI, Anthropic, and Google? Many people worry that the closed source advantage will quickly be erased.

14d ago

Kimi K3 with 2.78 trillion parameters runs on 8GB of memory. Developers open source lightweight C inference engine

Comparatively, a developer recently tried to run the Kimi K3 model with 2.78 trillion parameters on a device with only 8GB of RAM, an open source project called kimi-k3-in-c. The project is only 176 KB, written in pure C99, does not rely on GPU, CUDA, PyTorch, or BLAS, and can complete model inference using only the CPU. The solution leverages Kimi K3's MoE (Hybrid Expert) architectural features. Although the total number of model parameters reached 2.78T, only 16 of the 896 experts on each layer were activated, so instead of loading the full 1.56 TB model weights into memory, the developers stored most of the expert weights in the NVMe hard disk and read them in real time according to inference requirements; at the same time, some dense layers (dense trunks) also used a layer-by-layer streaming loading method. However, the solution still has significant performance limitations. In 8GB memory mode, the model takes about 32.7 seconds to generate a token, and requires close to 1.7 TB of high-speed storage support. The developers said that this solution is currently more like an experimental exploration of the direction of large model inference infrastructure optimization. It has no actual production and use value, but its method of streaming hard disk loading+MoE sparse activation provides a new idea for running hyperscale models at low cost in the future.

14d ago

FT: Dark Side of the Moon is restructuring the equity structure and introducing state-owned shareholders to obtain regulatory approval to go public in Hong Kong

Comparatively, according to the Financial Times, the Chinese AI startup Moonshot AI (Moonshot AI) is adjusting its corporate structure and bringing in a number of investors with a state-owned background to obtain approval from the regulatory authorities to go public in Hong Kong. The Kimi K3 model under Dark Side of the Moon recently closed the performance gap with Anthropic's leading model and was welcomed by developers. The company plans to raise capital through an initial public offering (IPO) for the next phase of large-scale model development and business expansion. Prior to listing, Dark Side of the Moon may require dismantling the existing red-chip structure. Since its core business was previously controlled by overseas entities and raised in US dollars through overseas investors, the Chinese regulatory authorities imposed restrictions on the overseas listing of some technology companies with overseas structures, causing many AI companies, including Dark Side of the Moon and StepFun (StepFun), to suspend IPO preparations. According to reports, Dark Side of the Moon changed the entity in China from a limited liability company to a limited liability company last week, which is an important step in advancing its listing. The company is currently coordinating with investment banks and a team of lawyers to resolve the issue of overseas investors' shareholding transfers. An investor revealed that The Dark Side of the Moon has recently been notified to shareholders, that the restructuring of overseas entities is still being promoted, and it has not yet been decided how to transfer the rights and interests of overseas shareholders to domestic entities. Possible solutions include overseas investors setting up shares held by domestic investors, or selling overseas shares first and then re-subscribing for domestic shares, but they all involve a lengthy approval process. In terms of financing, Dark Side of the Moon recently completed two rounds of financing. One round was valued at around US$30 billion, and the valuation is expected to reach US$50 billion after completion. The company previously attracted state-owned institutions such as the National Artificial Intelligence Industry Investment Fund to participate in the investment. The latest shareholder list also includes investors under the National Social Security Fund, Shanghai and Guizhou Local Government Guidance Funds, and People's Daily. On August 3, earlier market news stated that the Dark Side of the Moon plans to submit an IPO application in Hong Kong as early as this month, or raise about 3 billion US dollars. In response, Dark Side of the Moon responded that the news was untrue.

14d ago

Kimi K3 from Dark Side of the Moon broke through the sandbox range in the security test due to a misconfiguration of the sandbox

Comparatively, according to Wired, during the security test, Kimi K3, an open source model from the Dark Side of the Moon, broke through the sandbox during the security test and accessed the internet to try to find test answers through GitHub. American security startup Frontier Security claims that Kimi K3 exploited a sandbox configuration vulnerability, but it itself lacked sufficient built-in protection mechanisms. Similar to events previously disclosed by OpenAI and Anthropic, this escape was partly due to a sandbox misconfiguration. However, Kimi K3 did not launch an attack because the answers are easily available on GitHub. Frontier Security CEO said Kimi K3 excels in cybersecurity defense, but lacks protective mechanisms to prevent cheating or escaping the sandbox. Security experts say it's critical to configure a testing environment, and misuse of AI agents could cause the system to get out of control.

15d ago

Kimi K3 also jailbroken? After checking, the sandbox didn't go off the internet

Comparative news, according to monitoring, after OpenAI and Anthropic, Kimi K3 is also happy to mention that AI escaped the sandbox. WIRED directly wrote the title that one of China's strongest AI models has also broken through the quarantine; it sounds like Kim has broken through the sandbox herself. As soon as I saw the results, the door was unlocked again. When the security company Frontier Security tested Kimi K3, the sandbox had an outbound network vulnerability. It was supposed to isolate the public network, but it was also able to access GitHub. Kimi discovered that she could follow this path by exploring the internet herself. She simply cloned the official benchmark repository and found answers from there. The whole incident did not break a sandbox with a normal configuration, nor did it hack into any company. To put it bluntly, the procurator forgot to disconnect from the internet, and the candidates discovered that GitHub was still available, so they went straight to the answers. When it came to the headlines, it became AI breaking through isolation. Now the jailbreak threshold is so low that even if the door is unlocked, it's considered a jailbreak.

15d ago