DeepSeek · 475

DeepSeek launches multi-modal model V4-Flash-Vision-Exp and opens an API service

According to the DeepSeek website, DeepSeek has launched the multi-modal visual understanding model V4-Flash-Vision-Exp and opened an API service. The model adds image understanding and visual generation capabilities to V4-Flash, supports multi-modal input and output such as text and images, and provides developers with calling interfaces for inference and generation capabilities.

1d ago

The DeepSeek visual model's multi-modal Agent ability is close to Opus 4.8, 4 items are 2:2

In comparison, DeepSeek announced the first batch of V4-Flash-Vision-Exp agent scores. There were 2 wins and 2 losses against Opus 4.8 in 4 multi-modal Agent reviews: Aggregation Last Exam 27.3 versus 25.7, ZeroBench 35.0 versus 34.0; ApexBench 36.5 versus 39.4; and Chartography 64.3 versus 65.0. Compared to the plain text version of V4-Flash-0731, the visual version increased from 26.2 to 36.5 in ApexBench and from 25.2 to 27.3 for Aggregated Last Exam. The official statement indicates that the plain text version will ignore multi-modal content, so it mainly reflects the ability to supplement visual input. Coupled with the fact that the ability of the post-visual text agent was not significantly reduced, 6 of the 7 text reviews were superior to V4-Flash-0731. For example, DeepSWE upgraded from 54.4 to 59.3 (58.0 over Opus 4.8), and Toolathlon 75.9 almost tied with 76.2 of Opus 4.8. The results are from DeepSeek's official self-test, not a third-party independent list; the public Code Agent text task uses the DeepSeek Harness minimal mode, and the inference level is max.

1d ago

DeepSeek visual model officially launched API: the price is exactly the same as V4 Flash

In comparison, DeepSeek officially launched a new visual model, deepseek-v4-flash-vision-exp. The official API documentation has juxtaposed it with V4 Flash and V4 Pro, and developers can directly import images through the DeepSeek API. The model supports 1 million token contexts, a maximum output of 384,000 tokens, and also supports JSON Output, Tool Calls, Responses API, and Anthropic API. The price is directly aligned with the V4 Flash. The peak period for each million token input is 3 yuan and the idle period is 1.5 yuan; the cache hit is only 0.1 yuan and 0.05 yuan, respectively. Each million tokens are worth 9 yuan during peak output periods and 4.5 yuan during idle periods. Compared to the V4 Pro, the price of the same class is only one-third. There is no separate charge per image. DeepSeek will convert the image into a token based on the image size, and then bill it together with the text token. Peak periods are 9:00 to 12:00 and 14:00 to 18:00 Beijing time; prices are halved the rest of the day.

1d ago

DeepSeek V4 Flash vision model suspected to be released soon

Comparative news, according to X user MaxForAI, signs related to the DeepSeek visual model have appeared. Today at 3 p.m., the name deepseek-v4-flash-vision-exp has appeared in DeepSeek Harness related code, and the latest version of Harness has begun to add native image request adaptation. However, until the model endpoint is ready, the model will not be placed in the official model directory. Currently, DeepSeek should be testing the visual version of V4 Flash, and the public API has not been officially opened yet. Existing users directly use the model ID to request the DeepSeek API. Currently, an error will still be returned, and the model has not been listed in the official documentation. The adaptation path between the client and the Harness side has basically been paved, and DeepSeek V4 Flash is expected to soon have visual capabilities.

1d ago
From 4 models to more than 500, OpenRouter was acquired after growing 30,000 times in three years

From 4 models to more than 500, OpenRouter was acquired after growing 30,000 times in three years

Author: Menlo Ventures Compiled by: Jia Huan, ChainCatcher Original title: Early Investors Behind OpenRouter Revisited Investments Today, OpenRouter announced that it has reached an acquisition agreement with Stripe. OpenRouter was launched in 2023, just over three years ago. OpenRouter was initially launched as a “unified interface for LLM” and only supported 4 models at the time: GPT-3.5, GPT-4, GPT NeoXt and Cohere xlarge by Together. When the company was founded, it was based on two core judgments: first, AI will eventually be used on a large scale and penetrate various fields; second, there will be many different models on the market, each with trade-offs, and users will choose different models according to different needs. As it turned out, both judgments far exceeded expectations at the time. Since its launch, the number of tokens processed by the OpenRouter platform has increased by about 30,000 times. Currently, it has exceeded 4,500 trillion tokens on an annualized basis, and the scale of expenditure on the platform has reached a very impressive level. Meanwhile, the number of models supported by OpenRouter has grown from the original 4 to over 500. Figure: OpenRouter Token usage growth from inception to acquisition Menlo Ventures is fortunate to be part of this journey. In March 2025, we participated in OpenRouter's seed funding round through the Anthology Fund set up in partnership with Anthropic. OpenRouter founder and CEO Alex Atallah previously founded OpenSea, which was once valued at $13.3 billion. His co-founders include tech guru Louis Vichy, whom he met on Discord, and highly executive COO Chris Clark. In May 2025, we led OpenRouter's Series A funding round, with Matt joining the company's board of directors, and Deedy as a board observer. Earlier this year, after seeing OpenRouter's rapid growth in customer numbers and revenue, and the company built a product route with stronger “model intelligence” capabilities around model selection and evaluation, we continued to step up Series B financing. In the tech industry, it often takes years for an idea to change from the judgment of a few people to industry consensus. And just a few weeks ago, this happened: from Ramp to Cursor, more than 10 companies launched their own model routing products almost simultaneously. In just a few years, OpenRouter has become one of the most important companies in the AI era. Picture: Group photo when deciding to lead OpenRouter Round A At first glance, Stripe doesn't seem like the most natural buyer of OpenRouter, but the two companies are actually strikingly similar. Both use an API that can be directly accessed to simplify the otherwise complicated transaction process and charge a certain percentage of the fee. It's just that OpenRouter deals with AI models. As Stripe has always said, the two companies combined and are still doing the same thing: increasing “internet GDP.” In fact, over a year ago, OpenRouter called itself the “Stripe of LLM.” OpenRouter's core value OpenRouter was one of the first companies Deedy came into contact with after joining Menlo in 2024. This company is almost right at the heart of our AI infrastructure investment logic. Menlo presented two judgments necessary to invest in OpenRouter in the 2024 Enterprise AI Report: AI spending will increase dramatically, and developers will not only use one model, but multiple models at the same time. Figure: Menlo's initial contact email to OpenRouter As someone who can also write code and actually use these models, we realized long ago that there is a very clear difference in cost, latency, and performance between the different models...

2d agoburnking#OpenRouter

Rumor has it that Tencent Hy4 has appeared on the Yuanbao App model list, and gray testing has started

Comparative news, according to X user MaxForAI, Tencent has begun gray testing of the new flagship model Hybrid Hy4. Some users discovered that Hy4 has already appeared in the model selection list of the Tencent Yuanbao App and is labeled as an expert model, ranking above Hy3 and DeepSeek. Official account @TencentHunyuan This model option is simultaneously visible in the relevant portal. Tencent confirmed in its Q2 earnings report last week that Hy4 with larger parameters will be launched in the near future, further improving model performance and multi-modal capabilities. Currently, Tencent has not officially released Hy4. Outsiders are unable to confirm that this is a small-scale gray test or early opening of the entrance, but judging from the progress, the model is nearing launch.

2d ago

DeepSeek V4 Pro runs 5 sets of Harnesses: Pi has the highest success rate, DSH saves the most money

Comparative news, AI News, Agent Tool Company Composio put the same DeepSeek V4 Pro 0813 into Pi Agent, DeepSeek Harness 0.1, Claude Code, OpenCode, and Hermes Agent respectively, and unify them to the max level, running 30 multi-step Agent tasks. Pass rate: 1. Pi Agent: 21/302. DeepSeek Harness: 20/303 Claude Code: 19/304 OpenCode: 19/305. Hermes Agent: In the 18/3030 question, all 5 out of 15 were passed, and all 7 were broken. Only the Harness was replaced with the remaining 8 tracks, and the result was reversed. Speed ranking: 1. Claude Code: 181.8 seconds 2. DeepSeek Harness: 252.1 s 3. Hermes Agent: 273.6 seconds 4. OpenCode: 280.6 seconds 5. Pi Agent: 362.9 seconds Cost per success: 1. DeepSeek Harness: $0.028 2. Pi Agent: $0.031 3. OpenCode: $0.032 4. Hermes Agent: $0.037 5. Claude Code: $0.074 Pi has the highest pass rate, DeepSeek Official Harness saves the most money, and Claude Code is the fastest but most expensive. Composio previously used DeepSeek V4 Flash for similar tests, and Pi also had the highest pass rate at the time.

2d ago
Yao Shunyu reorganizes Tencent's multi-modal route: closer to Liang Wenfeng and away from Li Feifei

Yao Shunyu reorganizes Tencent's multi-modal route: closer to Liang Wenfeng and away from Li Feifei

Text | Miao Zheng Editor | Wang Jing Source | Letter AI Tencent's mixed element multi-modal team has undergone another personnel change. According to media reports, Lin Xudong, who was responsible for xAI's multi-modal understanding, has left xAI and joined Tencent's mixed element as the head of the multi-modal content generation algorithm. The reason this personnel news is worth paying attention to is that it takes place in the context of continuous adjustments of mixed and multi-modal teams. Over the past period of time, news of the departure of the person in charge, the transfer of researchers, and the addition of new members came out one after another within the mixed yuan. Hu Han, the former head of multimodal understanding, left his career to start a business, and Tian Yonglong and others joined Tencent. The reporting relationship between the original multi-modal team also changed with the integration of the big language model department and the multimodal model department. However, does this mean that Tencent's multi-modal team is “changing the dynasty” is currently unable to draw a direct conclusion. What can be confirmed by public information is that mixed forces have indeed experienced personnel movements and organizational restructuring. The rumor of Lin Xudong's addition is more like a new signal in this adjustment: Tencent is recombining the two routes of multimodal understanding and content generation. So the question is, what exactly did Lin Xudong come from, and what abilities can he add to Tencent? And is Tencent's multi-modal approach shifting from “generating content” to Yao Shunyu's more biased “understanding context and acting in the world”? What is Lin Xudong's origin and what can he do after joining Tencent? According to public information, Lin Xudong graduated from Tsinghua University in 2018 and then went to Columbia University to study for his doctorate. While studying at the blog, his research interests included embedded learning, video analysis, and generative models. He also participated in the Vx2Text project in collaboration with Columbia University and Facebook AI. V indicates video, x indicates unknown, can be sound, voice, or even ambient sound. 2 represents TO, and Text represents subtitles. Its logic is to first convert different modes such as video and sound into vectors similar to “language tokens”, then uniformly feed the language model for fusion, and finally generate open text by an autoregressive decoder. Transformer can only understand tokens, so AI essentially doesn't understand video and audio file formats, making it even less likely to convert them into text. For example, if a dog jumps into the water next to a swimming pool, Vx2Text's video recognizer (V) will output keywords: dog, jump, pool; sound reader (x) will output: sound of water, fluttering. Although the product function of Vx2Text is “generation,” the core difficulty of the product is “understanding.” Of course, Vx2Text doesn't simply “translate” a screen into a few sentences. Models need to recognize people, objects, movements, and events from videos, understand how these things change over time, and finally organize visual information into language. After graduating from his PhD, Lin Xudong joined DeepMind and participated in Gemini-related multi-modal pre-training and post-training work. In 2025, he also joined xAI. According to public information, it is responsible for the direction of multimodal understanding and participating in the training of multimodal content understanding and generation models. Now that he has joined Tencent Hybrid, he will be responsible for the hybrid multi-modal content generation algorithm. Lin Xudong was added not so much to improve the performance of mixed-element multi-modal generation, but rather to solve a problem that plagues all multimodals — understanding. The previous generation model was more like a picture maker. Give it a hint, and it can generate an image or a video. But as long as users make more complex requests, the model just can't keep up. For example, the characters change in the long video, the shape of the object is not consistent before and after, the camera movement does not match the spatial relationship, etc. It's not because the model doesn't generate, but because it doesn't remember and understand the world steadily. Therefore, putting Lin Xudong in the position of multi-modal content generation is probably because he “translated” multi-modality into something AI can understand. Lin Xudong's addition can only be clearly seen in a larger context. That is, now Tencent's mixed element is reorganizing its multi-modal route. In January 2025, Tencent Outstanding Scientist (Tencent Distinguished Scientist) Hu Han succeeded Liu Wei, who had previously left his job, and was fully responsible for the research and development of mixed-element multi-modal models, and also served as Tencent's mixed-element big model Tech Lead. Tencent's internal organization was adjusted in the second half of 2025. He transferred from the Multimodal Model Department to the “Frontier” Frontier Technology Research Group under the Big Language Model Department. The title was changed to Head of the Multimodal Understanding Direction, and the reporting line was also changed to report to Yao Shunyu. The actual position changed from “the head of an independent department” to a “big language model...

2d ago字母AI#AI #Li Feifei #Liang Wenfeng #Tencent

Lyon: Smart Spectrum's training ability improved significantly after GLM-5.3, maintaining an outperforming market rating

Comparative news, according to a report by Jin Shi, Lyon published a research report stating that the Intelligent Spectrum (02513.HK) GLM-5.3 API will now be open for use, showing leading performance in the domestic industry in terms of complex coding and long-term proxy tasks. Although the parameter scale is small, it has achieved open-weighted SOTA performance in many benchmark tests, and scored 60 points in the Artificial Analysis Intelligence Index, which is on par with Kimi 3, surpassing Tongyi Qianwen 3.8 Max. The bank estimates that the GLM model's OpenRouter revenue share has risen from 1% in January to 7% in July, leading DeepSeek's 6% and Dark Side of the Moon's 3%. It is believed that DeepSeek's recent price increase reflects healthy competition in the industry, and GLM 5.3 has regained Pareto's leading position. The recent weak stock price may reflect the market's overreaction to Anthropic's annualized recurring revenue deceleration, maintaining a smart score that outperforms the market rating. The target price is HK$2,061.

3d ago
Why is capital chasing AI Native and ignoring the old Internet

Why is capital chasing AI Native and ignoring the old Internet

Capital doesn't reward being old-fashioned, not because old-fashioned people are at fault. The old part is clearly priced. There is no bad information, so there is no excess profit. Global venture capital was $510 billion in the first half of 2026, surpassing $44 billion for the full year of 2025 in one and a half months. More than 70% have entered AI; OpenAI and Anthropic took 217 billion dollars, accounting for 43%. With that much money, you'd think everyone could share a little bit. The truth is that distribution is more extreme than total volume, and the first sieve doesn't screen the industry, it screens people. The category that has been screened out now has an unkind name: the internet is old. Let's just say one thing: the “old man” in this article has nothing to do with age. It refers to a set of methodologies that have been formed in the mobile internet cycle, have been tested over and over, and have brought huge returns to holders. The person holding it may be 45 years old or 32 years old. It was this methodology that was being repriced, not the year of birth. Confusing these two things is Lao Deng's most common mistake and one of the most comfortable mistakes — because if the problem is someone else's age discrimination, you don't need to change a single word. 01 What is AI Native The term has been misused. They can use ChatGPT not called AI native, nor AI in the company name, let alone in their twenties. There are three things that really separate people. First, the starting point is a model, not a requirement. The order in which Lao Deng makes a product is: look at what the user wants, write down the requirements, and find technology to implement it. The order of AI natives is reversed: first figure out what level the model is capable of today and what step it is likely to reach tomorrow, and then move from this capability boundary to the external product. The former uses the model as a tool, and the latter uses the model as the foundation. There was no difference between these two kinds of things made by humans in the first edition; by the third edition, there was a difference of one species. Article 2. The default unit of an organization is not a person. The division of labor in the Internet age is the division of one thing into ten people. AI Native's division of labor is to take ten things from one person and add a bunch of agents. The CEO of a domestic application company said that the team consists of less than ten people, but a large number of AI work at night, and the first thing employees do every morning is check the work the AI handed in the night before. Cursor's side is even more extreme. Public reports mention that the company doesn't have a product manager; engineers write their own code, talk to users themselves, and participate in recruiting people themselves. Article 3. Information is first-hand. AI Native's input sources are papers, model cards, GitHub issues, original discussions on X, and self-run evals. Lao Deng's input sources are industry summits, closed-door meetings, brokerage reports, interpretation of public accounts, and finding someone to drink coffee with. This one is the least obscure and most lethal; I'll talk about that separately later. I'm satisfied with all three. The 25-year-old is an AI native, and so is the 45-year-old. I'm not satisfied with the three rules; I'm still an old man at the age of 25. AI natives are a state, not an age group. The trouble is that tickets in this state are works, not resumes. 02 The two lists spread the results of this round on the table. These are two lists. The first one is an all-AI native company. Their valuations are not rising; they are exchanging orders of magnitude. List 1 · Upstream OpenAI raised $122 billion in a single round of financing in Q1 2026, followed by $852 billion, the largest private equity financing in history. Anthropic Q2 had a single round of $65 billion, after investing $965 billion, accounting for about half of the total global venture capital for the quarter; the revenue operating rate in May reached about $47 billion. DeepSeek raised about 70 billion yuan in its first round of financing in May 2026. In April of the same year, Liang Wenfeng raised his direct shareholding from 1% to 34%, and controlled a total of about 84.29% of the shares through related entities. The Dark Side of the Moon (Kimi) was estimated at $4.3 billion in December 2025; it went for three consecutive rounds from January to February 2026 to reach 18 billion; the D round in May was about $2 billion, breaking 20 billion dollars after the investment; the July round surpassed $3.5 billion, after investing 35 billion dollars; the pre-IPO target was 50 billion dollars. ARR broke 100 million in March, 200 million in May, and held steady at 300 million US dollars in June, with APIs accounting for more than 70%. Smart Spectrum · MiniMax successively landed in Hong Kong stocks in early 2026, with a market capitalization exceeding 100 billion yuan. It was one of the first major model companies listed in China. The second one...

4d agoWendy#AI #DeepSeek