Yao Shunyu reorganizes Tencent's multi-modal route: closer to Liang Wenfeng and away from Li Feifei

source字母AI·字母AI·05:12 编辑
Yao Shunyu reorganizes Tencent's multi-modal route: closer to Liang Wenfeng and away from Li Feifei
Miao Zheng
editsWang Jing
Source|Letter AI

Tencent's mixed multi-modal team underwent another personnel change. According to media reports, Lin Xudong, who was responsible for xAI's multi-modal understanding, has left xAI and joined Tencent's mixed element as the head of the multi-modal content generation algorithm.

The reason this personnel news is worth paying attention to is that it takes place in the context of continuous adjustments of mixed and multi-modal teams.

Over the past period of time, news of the departure of the person in charge, the transfer of researchers, and the addition of new members came out one after another within the mixed yuan.

Hu Han, the former head of multimodal understanding, left his career to start a business, and Tian Yonglong and others joined Tencent. The reporting relationship between the original multi-modal team also changed with the integration of the big language model department and the multimodal model department.

However, does this mean that Tencent's multi-modal team is “changing the dynasty” is currently unable to draw a direct conclusion.

What can be confirmed by public information is that mixed forces have indeed experienced personnel movements and organizational restructuring.The rumor of Lin Xudong's addition is more like a new signal in this adjustment: Tencent is recombining the two routes of multimodal understanding and content generation.

So the question is, what exactly did Lin Xudong come from, and what abilities can he add to Tencent? And is Tencent's multi-modal approach shifting from “generating content” to Yao Shunyu's more biased “understanding context and acting in the world”?

Image

What is the origin of Lin Xudong

What can he do after joining Tencent?

According to public information, Lin Xudong graduated from Tsinghua University in 2018 and then went to Columbia University to study for his doctorate.

Image

While studying at the blog, his research interests included embedded learning, video analysis, and generative models. He also participated in the Vx2Text project in collaboration with Columbia University and Facebook AI.

V indicates video, x indicates unknown, can be sound, voice, or even ambient sound. 2 represents TO, and Text represents subtitles. Its logic is to first convert different modes such as video and sound into vectors similar to “language tokens”, then uniformly feed the language model for fusion, and finally generate open text by an autoregressive decoder.

Transformer can only understand tokens, so AI essentially doesn't understand video and audio file formats, making it even less likely to convert them into text.

For example, if a dog jumps into the water next to a swimming pool, Vx2Text's video recognizer (V) will output keywords: dog, jump, pool; sound reader (x) will output: sound of water, fluttering.

Although the product function of Vx2Text is “generation,” the core difficulty of the product is “understanding.”

Of course, Vx2Text doesn't simply “translate” a screen into a few sentences. Models need to recognize people, objects, movements, and events from videos, understand how these things change over time, and finally organize visual information into language.

After graduating from his PhD, Lin Xudong joined DeepMind and participated in Gemini-related multi-modal pre-training and post-training work. In 2025, he also joined xAI. According to public information, it is responsible for the direction of multimodal understanding and participating in the training of multimodal content understanding and generation models.

Now that he has joined Tencent Hybrid, he will be responsible for the hybrid multi-modal content generation algorithm.

Lin Xudong was added not so much to improve the performance of mixed-element multi-modal generation, but rather to solve a problem that plagues all multimodals — understanding.

The previous generation model was more like a picture maker. Give it a hint, and it can generate an image or a video. But as long as users make more complex requests, the model just can't keep up. For example, the characters change in the long video, the shape of the object is not consistent before and after, the camera movement does not match the spatial relationship, etc.

It's not because the model doesn't generate, but because it doesn't remember and understand the world steadily.

Therefore, putting Lin Xudong in the position of multi-modal content generation is probably because he “translated” multi-modality into something AI can understand.

Lin Xudong's addition can only be clearly seen in a larger context. That is, now Tencent's mixed element is reorganizing its multi-modal route.

In January 2025, Tencent Outstanding Scientist (Tencent Distinguished Scientist) Hu Han succeeded Liu Wei, who had previously left his job, and was fully responsible for the research and development of mixed-element multi-modal models, and also served as Tencent's mixed-element big model Tech Lead.

Tencent's internal organization was adjusted in the second half of 2025. He transferred from the Multimodal Model Department to the “Frontier” Frontier Technology Research Group under the Big Language Model Department. The title was changed to Head of the Multimodal Understanding Direction, and the reporting line was also changed to report to Yao Shunyu.

The actual position changed from “the leader of an independent department” to “the head of a certain direction in a research group under the Big Language Model Department,” and the scope of management was reduced from the overall multi-modal model to only multi-modal understanding, and directions such as multimodal generation were delineated.

In March 2026, Tencent withdrew its AI Lab, which had been established for nearly ten years, and some personnel were merged into the Mixed-Yuan Language Model Department and transferred to Yao Shunyu to report. After this adjustment, Jiang Jie, who was the actual leader of mixed yuan and vice president of Tencent, officially completed the handover of work with Yao Shunyu, and no longer managed both AI Lab and mixed yuan.

At the beginning of July 2026, former OpenAI researcher Yao Shunyu and Tsinghua undergraduate alumnus Tian Yonglong joined Tencent to take charge of the visual language model direction.

Shortly after that, Tencent TEG officially published an article, abolished the Big Language Model Department and the Multi-modal Model Department, and merged to establish the “Basic Model Department”. Yao Shunyu became the person in charge and reported to TEG President Lu Shan. According to Tencent, the move is aimed at improving the efficiency of model development and collaboration, and exploring the upper limit of intelligence for full-modal models.

Afterwards, it was revealed that Hu Han left his career and started a business.

Linus, a former chief scientist at Amazon, chief scientist at JD Mathematics, head of the application vision team at Ali Tongyi Laboratory, and former head of Tencent's multi-modal model department, changed from being in a parallel position with Yao Shunyu to report to Yao Shunyu after this adjustment.

It's not over yet. Zhong Zhao, head of Tencent's mixed-element multi-modal basic model, is responsible for the two generation lines of HunyuanVideo and HunyuAnimage, and is the actual person at the helm of the hybrid multi-modal generation direction.

He posted a circle of friends on August 14, saying “the upcoming version may be my final version in the HunyuanVideo and HunyuAnimage projects,” and the lines are full of farewell.

Lin Xudong is probably just a member of the mixed-element multi-modal exchange, but his addition sends a clear signal that Tencent's entire direction of multi-modal research and development has changed.

Image

Prior to that,

How does Tencent make multi-modal?

Previously, Tencent's language model and multimodality were two fundamental lines, and the two core products of hybrid multi-modality were HunyuanVideo and HunYuanImage mentioned earlier.

HunyuanVideo is a Wensheng video model. First released in December 2024, it became the world's largest open source video generation model at the time with 13 billion parameters, supporting 5 seconds of 720p output. Intensive expansion capabilities in 2025. Tucson Video was launched in March, and customized generation of HunyuanCustom and digital human-driven Avatar was launched in May. Version 1.5 was released in November, with parameters reduced to 8.3 billion. It supports native generation of 5 to 10 seconds of 480p or 720p video, and 1080p output can be obtained through super resolution.

HunyuAnimage started even earlier. It was the first native Chinese DIT architecture image model launched in May 2024. It iterated the 2.0 version of the industry's first industrial-grade real-time raw map in 2025. In September of the same year, version 2.1 (17 billion parameters, native 2K resolution) and version 3.0 (80 billion parameters, the first industrial-grade open source native multi-modal biograph model) were released, and an Instruct version with inference capabilities was further launched in January 2026, which supports self-rewriting of prompts and graph editing.

Another line is Hy 3D, which is a model dedicated to generating single 3D models. It was open sourced with Hunyuan-Large in November 2024, upgraded to 2.0 in January 2025, split into a two-stage assembly line for DIT shape generation and paint texture synthesis. The complete open source training code was released in June 2.1, and 3.0 was released in September to increase modeling accuracy by 3 times and the geometric resolution reached 1536³. It has now been iterated to 3.1. It is widely used in e-commerce modeling, product design, and 3D printing.

After these three mountains, perhaps Li Feifei's spatial intelligence theory influenced Tencent, or many years ago, the relevant knowledge acquired by Tencent in digital twins was reused with the support of AI.

Just as Liu Guanzhang was followed by Zhao Yun, the mixed-yuan multi-modal model also welcomed its fourth brother, Hy World 1.0, in late July 2025.

Image

The so-called spatial intelligence, according to Li Feifei's World Labs, refers to allowing AI to sense, generate, reason, and interact with the world in three-dimensional space.

Language models learn the statistical structure of text, while world models need to learn the structure of space and time, such as how light falls on the surface of an object, what a garden would look like from an unphotographed perspective, and how objects move when exposed to external forces.

According to Tencent, Hy World 1.0 is the world's first open source, immersive 3D world generation model that supports simulation. The technical route is to first generate a 360-degree panorama as a “world agent”, then break it down into semantic layers for hierarchical 3D reconstruction, and finally output a hierarchical mesh that can be imported into a game engine.

Since Tencent itself is very good at developing game engines and game content, although Hy World's expressiveness is amazing, it was also ridiculed by netizens as “Tencent doesn't want to leave its comfort zone.”

More than a month after Hy World 1.0 was released, Tencent also released the Hy World-Voyager. The positioning is an ultra-long world roaming model based on camera trajectory control.

Hy World-Voyager addresses a pain point left by 1.0. Although the 3D world generated by 1.0 is of good quality, the roaming range is quite limited, and geometric drift and molding will occur if you don't go far.

Voyager's idea is to change the technical route. Instead of directly generating 3D assets, it first generates an RGB-D bimodal video. Each frame simultaneously outputs a depth map while outputting color, and then instantly reconstructs a 3D point cloud using the depth map.

All in all, Hy World can almost be said to have perfectly recreated Li Feifei's spatial intelligence vision through Tencent's advantage in engine and art.

You know, the full technical report for Hy World 2.0 is “HY-World 2.0: A Multi-Modal World Model for Reconstructing, Rebuilding, and Simulating 3D Worlds,” which includes “Reconstructing, Generating, and Simulating 3D Worlds” (Reconstructing, Generating, and Simulating 3D Worlds) Simulate a 3D world), which is World Labs' explanation of spatial intelligence, as mentioned earlier.

However, not everyone agreed with Li Feifei, including Yao Shunyu. On the multi-modality issue, Yao Shunyu seems to agree more with Liang Wenfeng's views.

Image

Yao Shunyu agreed with Liang Wenfeng

Li Feifei's opinion is that the world model, as part of spatial intelligence, is an important line of AGI, while Liang Wenfeng opposes it.

An investor once asked Liang Wenfeng about the multi-modal question. Liang Wenfeng replied, “We only do the main AGI line — GPT, CoT, Agent. The field of AI is very broad, but 3D, video generation, and world models don't have much to do with the main line of intelligence.”

DeepSeek has a product called DeepSeek OCR, which is almost a technical figurative of Liang Wenfeng's multi-modal view. Think of visual modals as tools that serve language models rather than independent intelligent directions.

The core of DeepSeek OCR is called “Contexts Optical Compression” (Contexts Optical Compression). The core idea is to render long text into an image, use a visual encoder to compress the text into a compact visual token, and then hand it over to the language model for decoding.

Image

Isn't that a bit familiar? Looking back at Lin Xudong's Vx2Text, Yao Shunyu probably just wanted to learn from Liang Wenfeng, transform the existing multi-modal foundation of mixed elements into such a way of understanding, and then better integrate it into the main line of the Hy model.

Including OpenAI, many big AI companies actually do similar things. Yao Shunyu worked for OpenAI and probably established this direction at that point.

Yao Shunyu's first flagship product after joining Tencent was Hy3. He focused on reasoning, intelligence, long context, and productivity tasks.

According to officials, Hy3 has optimized tasks such as software development, office production, financial modeling, front-end design, and game production, and continues to improve tool call stability, complex context handling, and multi-round intent maintenance.

The goal of Hy3 is not to “break the list,” but to make the model less error-prone, able to understand the context, and continue to perform tasks in the actual product.

This product itself is very “Yao Shunyu”.

On June 5, 2026, during a conversation with Tang Daosheng at the Tencent Cloud AI Industry Application Conference, Yao Shunyu made this judgment. The competitive barrier for AI in the second half was not the amount of model parameters, but the context (context).

“We now feel like we have an all-purpose hammer that can smash any nail. The methodology has become very mature; on the contrary, it is becoming more and more difficult to find really worthwhile questions.” Yao Shunyu said this.

Also at that conference, the three-tier structure of Tencent's long-term AGI construction was announced: basic foundation (pre-training+post-training+infrastructure), product implementation, and cutting-edge exploration. Reasoning ability is a basic problem to be solved, and contextual ability is at the intersection of product implementation and cutting-edge exploration.

Previously, in order to break the test model away from the original knowledge and only seek answers from the context, Yao Shunyu also specially created the test context benchmark CL-Bench and CL-Bench Life.

As can be seen, Yao Shunyu is actually like Liang Wenfeng. He feels that the main line is agent-related content rather than multi-modal generation.

As for Hy3, it depends on WorkBuddy. WorkBuddy is one of the launch platforms of HY3. The two are in a “product+model” double spiral relationship and achieve mutual success.

WorkBuddy's performance improved significantly after connecting to HY3. Based on internal evaluation of the office scenario, the task resolution rate jumped from 72% to 90%, the average time required was reduced by 34%, and complex requirements can be processed about one-fifth more.

After Hy3 went live, the number of users actively choosing HY3 on WorkBuddy increased 6 times, and the average daily text processing volume increased 20 times compared to the preview version. On July 8, computing power consumption peaked at one point, and the queuing rate exceeded 50%. Tencent urgently expanded its capacity for this reason.

One conjecture is that after Hy improves its multi-modal understanding, its experience in the GUI Agent scenario will take it to the next level. The core link of GUI Agent is “Perceiving the screen — understanding the interface — making decisions — performing verification”. Among them, multi-modal understanding is a key bottleneck connecting visual perception with language decision making.

But that's not enough proof that Li Feifei was wrong; it can only be said that Yao Shunyu agreed with Liang Wenfeng.


Twitter:https://twitter.com/BitpushNewsCN

Compare the TG exchange group:https://t.me/BitPushCommunity

Compare TG subscriptions:https://t.me/bitpush

Original Link
#AI#李飞飞#梁文峰#腾讯
说明: All Bitpush articles reflect the author's views only and do not constitute investment advice.

Related

Loading...