李飞飞 · 14
Yao Shunyu reorganizes Tencent's multi-modal route: closer to Liang Wenfeng and away from Li Feifei

Yao Shunyu reorganizes Tencent's multi-modal route: closer to Liang Wenfeng and away from Li Feifei

Text | Miao Zheng Editor | Wang Jing Source | Letter AI Tencent's mixed element multi-modal team has undergone another personnel change. According to media reports, Lin Xudong, who was responsible for xAI's multi-modal understanding, has left xAI and joined Tencent's mixed element as the head of the multi-modal content generation algorithm. The reason this personnel news is worth paying attention to is that it takes place in the context of continuous adjustments of mixed and multi-modal teams. Over the past period of time, news of the departure of the person in charge, the transfer of researchers, and the addition of new members came out one after another within the mixed yuan. Hu Han, the former head of multimodal understanding, left his career to start a business, and Tian Yonglong and others joined Tencent. The reporting relationship between the original multi-modal team also changed with the integration of the big language model department and the multimodal model department. However, does this mean that Tencent's multi-modal team is “changing the dynasty” is currently unable to draw a direct conclusion. What can be confirmed by public information is that mixed forces have indeed experienced personnel movements and organizational restructuring. The rumor of Lin Xudong's addition is more like a new signal in this adjustment: Tencent is recombining the two routes of multimodal understanding and content generation. So the question is, what exactly did Lin Xudong come from, and what abilities can he add to Tencent? And is Tencent's multi-modal approach shifting from “generating content” to Yao Shunyu's more biased “understanding context and acting in the world”? What is Lin Xudong's origin and what can he do after joining Tencent? According to public information, Lin Xudong graduated from Tsinghua University in 2018 and then went to Columbia University to study for his doctorate. While studying at the blog, his research interests included embedded learning, video analysis, and generative models. He also participated in the Vx2Text project in collaboration with Columbia University and Facebook AI. V indicates video, x indicates unknown, can be sound, voice, or even ambient sound. 2 represents TO, and Text represents subtitles. Its logic is to first convert different modes such as video and sound into vectors similar to “language tokens”, then uniformly feed the language model for fusion, and finally generate open text by an autoregressive decoder. Transformer can only understand tokens, so AI essentially doesn't understand video and audio file formats, making it even less likely to convert them into text. For example, if a dog jumps into the water next to a swimming pool, Vx2Text's video recognizer (V) will output keywords: dog, jump, pool; sound reader (x) will output: sound of water, fluttering. Although the product function of Vx2Text is “generation,” the core difficulty of the product is “understanding.” Of course, Vx2Text doesn't simply “translate” a screen into a few sentences. Models need to recognize people, objects, movements, and events from videos, understand how these things change over time, and finally organize visual information into language. After graduating from his PhD, Lin Xudong joined DeepMind and participated in Gemini-related multi-modal pre-training and post-training work. In 2025, he also joined xAI. According to public information, it is responsible for the direction of multimodal understanding and participating in the training of multimodal content understanding and generation models. Now that he has joined Tencent Hybrid, he will be responsible for the hybrid multi-modal content generation algorithm. Lin Xudong was added not so much to improve the performance of mixed-element multi-modal generation, but rather to solve a problem that plagues all multimodals — understanding. The previous generation model was more like a picture maker. Give it a hint, and it can generate an image or a video. But as long as users make more complex requests, the model just can't keep up. For example, the characters change in the long video, the shape of the object is not consistent before and after, the camera movement does not match the spatial relationship, etc. It's not because the model doesn't generate, but because it doesn't remember and understand the world steadily. Therefore, putting Lin Xudong in the position of multi-modal content generation is probably because he “translated” multi-modality into something AI can understand. Lin Xudong's addition can only be clearly seen in a larger context. That is, now Tencent's mixed element is reorganizing its multi-modal route. In January 2025, Tencent Outstanding Scientist (Tencent Distinguished Scientist) Hu Han succeeded Liu Wei, who had previously left his job, and was fully responsible for the research and development of mixed-element multi-modal models, and also served as Tencent's mixed-element big model Tech Lead. Tencent's internal organization was adjusted in the second half of 2025. He transferred from the Multimodal Model Department to the “Frontier” Frontier Technology Research Group under the Big Language Model Department. The title was changed to Head of the Multimodal Understanding Direction, and the reporting line was also changed to report to Yao Shunyu. The actual position changed from “the head of an independent department” to a “big language model...

2d ago字母AI#AI #Li Feifei #Liang Wenfeng #Tencent

Li Feifei warns that anti-AI sentiment in the US is heating up: if there is no positive development path, the world will be affected

Comparing the news, AI pioneer Li Feifei said that the tech industry needs to better explain the value brought by artificial intelligence to the public and warned that growing anti-AI sentiment in the US could pose a risk to global AI development. In an interview with Bloomberg, Li Feifei said that technology practitioners need to strengthen communication with the public to show the positive impact AI technology can bring. She pointed out that “if we don't show a positive attitude and a positive direction of development, everyone will be affected”. The US has an important influence in the field of global technology, and if the US fails to show a “positive AI attitude and development path,” it will eventually affect the entire world. As an important researcher in the field of computer vision, Li Feifei has continued to promote AI technology from laboratories to real-world applications in recent years. World Labs, the AI startup she founded, is dedicated to developing spatial intelligence (Spatial Intelligence) technology to explore AI's ability to understand and interact with the real world. Li Feifei believes that the AI industry is currently facing not only technical challenges, but also issues such as social acceptance, public trust, and the regulatory environment. With the rapid spread of generative AI, concerns about employment impacts, data privacy, and security risks continue to increase, and American society's backlash against AI development is expanding. She emphasized that the AI industry needs to be more proactive in explaining how technology can improve fields such as healthcare, scientific research, and productivity, rather than just focusing on technology competition itself. In recent years, the world's major technology companies have continued to increase investment in AI, promoting the development of technologies such as big models, AI infrastructure, and intelligent agents. But at the same time, issues of AI regulation, employment substitution, and social impact have also become the focus of attention of policy makers and the public.

3d ago
Hinton, Li Feifei, and Wu Enda are on the same stage for the first time: targeting AI companies together

Hinton, Li Feifei, and Wu Enda are on the same stage for the first time: targeting AI companies together

Source: First Electric Network Author: First Electric Editorial Office On the morning of August 6, Beijing time (morning of August 5, local time in Las Vegas), Jeffrey Sinton, Li Feifei, and Wu Enda were on the same stage for the first time. The round table called “Smart Architects: A Historic Gathering” was the most watched event at this year's Ai4 conference. Since Hinton and Wu Enda are almost completely opposed to each other on “will AI destroy humanity”, everyone is looking forward to a live confrontation before the meeting. In fact, all three have their own opinions on regulation, unemployment, and China's open weight model. On stage, Hinton directly said, “We don't always agree,” and Li Feifei immediately added “But we are still friends.” But these three people, whose positions are far apart, are pointing the finger in the same direction: big AI companies. A person seen as a representative of apocalyptic theory, a standard-bearer of individualistic AI, and an open source faction that has long opposed existential risk narratives. They all agree on this matter. As of press release, the organizers have not released the video or full transcript of this round table. First Electric compiled the core views of the three people based on the participants' immediate records, post-conference summaries, and live media reports. ▍ Li Feifei: Increased productivity does not equal prosperity. Li Feifei is the current CEO of the space intelligence company World Labs. The ImageNet data set she led and its 2012 competition are widely regarded as the starting point of this wave of deep learning. She is also the co-founder of Stanford's Human-Centered AI Research Institute. At the round table, she criticized current “irrational and unscientific” remarks surrounding AI. Among the three sources she listed, in addition to critics and many journalists, there are also big AI companies that have financial motives to keep competitors out of their doors. “Let's bring science, not sci-fi, back to the AI debate.” At the same time, she criticized that the utopian statement is also unhelpful; “if used improperly, it can also cause harm.” One of Li Feifei's most quoted words at this roundtable was that increased productivity does not automatically translate into shared prosperity. Improving efficiency is one thing; who the benefits go to is another. On the employment issue, she believes that AI will improve the efficiency of certain work processes rather than the disappearance or retention of an entire job. After all, “no job is a single task.” But for the jobs that are actually being replaced, she thinks a “soft landing” is needed. Many participants also described Li Feifei as the most cautious of the three about the rapid acceleration of AI and its impact on society, and was applauded by the audience. Li Feifei also said that describing AI development as a dispute between open source and closed source routes is a false debate. She used two analogies. The first is nuclear physics. Basic research is carried out publicly, but the most sensitive downstream applications, such as uranium enrichment, are always strictly controlled, and AI is likely to fall on a similar spectrum rather than converge into a single model. The second is the human genome project in the 90s of the last century: at the time, a private company and a consortium of publicly funded universities were the first to complete sequencing. If the private sector completely wins and applies for a patent, this technology may have been blocked from the wider scientific research and pharmaceutical industry. However, in reality, the results of the two parties were jointly announced by the Clinton administration, and the resulting public data integrated the foundation for many years of drug development. With this, she emphasized the importance of the government supporting AI research and making basic breakthroughs widely applied. “This debate, especially on the general level of “we can only tolerate one kind,” is a false debate.” Li Feifei named the reporter's responsibility to break out of this framework and discuss when, where, and how to use varying degrees of openness or closure. Asked what kind of AI headlines the world might wake up to in the next five or six years, her answer was: Using AI as a tool, the world announced the elimination of illiteracy. ▍ Hinton: Regulation is needed, but not this kind of regulation. Hinton won the 2024 Nobel Prize in Physics for his neural network research. After leaving Google in 2023, he continued to speak out about AI risks. He has publicly estimated that the probability of human extinction due to AI development is between 10% and 20%. Although tech leaders generally believe regulation will stifle innovation, Hinton believes regulation is essential to guide the safe development of AI. “You can't hand over the control of artificial intelligence to people like Elon Musk and Mark Zuckerberg” received the most enthusiastic applause that morning. What he wants is not fewer rules, but rules not to be written by people with the most motive to miswrite them. When it comes to employment, Hinton's judgment is pessimistic. He pointed out that AI has surpassed humans in some ways, and predicted that jobs such as call center operators and paralegals will increasingly face the risk of unemployment. “Once artificial intelligence is capable of routine mental work, anything involving routine mental work...

14d agoWendy#AI #Wu Enda #big model #Li Feifei #Hinton

Li Feifei's World Labs buys robotics company SceniX

Comparing news, World Labs, an artificial intelligence company owned by Li Feifei, announced the acquisition of robotics company SceniX. The exact amount of the acquisition has not yet been disclosed. World Labs said that robots are an important carrier for spatial intelligence to land in the real world, and future robot systems need to have environmental perception, spatial understanding, behavioral reasoning, and reliable execution capabilities. Following this acquisition, SceniX will be integrated into the World Labs team, and the two sides plan to combine their respective strengths in spatial modeling and robotics to drive the implementation of more application scenarios.

30d ago
Messi voted for Li Feifei; Liang Wenfeng's four-hour investor conference quotes flashed the screen; the second wave of US stock market signals...

Messi voted for Li Feifei; Liang Wenfeng's four-hour investor conference quotes flashed the screen; the second wave of US stock market signals...

Dear readers, what have the KOLs on X been talking about in the past 24 hours? Note: The following content is compiled from the X platform. They are all personal opinions. They do not represent the platform's position, let alone constitute investment advice. Macy invests in World Labs founded by Li Feifei Read more: https://36kr.com/p/3906511292994950梁文锋四小时投资人会议语录刷屏原文链接:https://www.bitpush.news/articles/7662849美股第二波行情信号? Twitter: https://twitter.com/BitpushNewsCN比推 TG Community: https://t.me/BitPushCommunity比推 TG Subscriptions: https://t.me/bitpush

30d agoWendy#KOL

Messi voted for Li Feifei, and the King of Soccer began betting on world models

Comparative news, according to monitoring, Macy's investment platform Play Time has included World Labs founded by Li Feifei in its investment portfolio. World Labs studies spatial intelligence and world models, with the goal of enabling AI to understand, generate, and manipulate 3D worlds. Play Time is Macy's main investment platform founded in 2022, initially focusing on sports, media, and technology. Today, the investment list also includes robotics company FieldAI, voice AI company Fish Audio, and AI data platform SuperAnnotate. World Labs raised $1 billion in February this year. Major investors include Nvidia, AMD, and Autodesk.

31d ago
Li Feifei: Functional Classification of World Models and Prospects for Spatial Intelligence

Li Feifei: Functional Classification of World Models and Prospects for Spatial Intelligence

Author: Ga Yang Original title: Li Feifei Latest Long Article: When video generation, robots, and NVIDIA all call themselves world models, we need a taxonomy “world model” which is probably the hottest and most confusing concept in the AI field since 2025. When Sora came out, OpenAI called it a world simulator; Genie lets you walk around in the generated images, also called a world model; the robotics company said it was making a world model; NVIDIA said Omniverse was the infrastructure for the world model, and even the game engine was pulled into this story. Everyone is using the same words, but they aren't saying the same thing at all. Today, Li Feifei published a new article on his personal Substack clarifying this concept. She first went back to the most classic diagram in the reinforcement learning textbook (POMDP closed loop: intelligence → action → state → observation → smart body), then pointed out that what is now called a “world model” is actually three different projections of this closed loop. The output pixel (observation) is the renderer, the output state is the emulator, and the output action is the planner. The classification criteria are very simple, depending on which part of the closed loop you are outputting. (Source: MIT Technology Review) Of the three, she determined that of the three, the renderer is the most commercialized but has a ceiling (good looking doesn't mean physically correct), the planner is the most exciting but furthest from actual deployment (the gap between lab demonstration and actual use is still huge), and the emulator is a critical hub that is seriously underestimated. Because the simulator works at the level of geometry, physics, and dynamics, it can not only project pixels upward for human consumption, but also derive action consequences downward for robots to use. Once you master simulation, you have the foundation for rendering and planning at the same time; not the other way around. This post is, of course, a World Labs product declaration. Their Marble is already outputting both Gaussian spatter and collision meshes in an attempt to unify the renderer and simulator into a single model. The end story depicted at the end of the article is a unified world basic model that can freely switch between rendering, simulation, and planning according to downstream requirements. Whether this vision can be realized is another story, but as an analytical framework, the renderer/simulator/planner's rule of three may indeed help penetrate some of the noise of the current “world model” concept. The full text is translated below. “The world is the sum of everything that happened.” ——Wittgenstein, “Philosophy of Logic,” 1921 The world is not composed of words. In an earlier article, we proposed that spatial intelligence is the next frontier of AI, and the world model is the path to it. Now, the World Labs team and I wanted to go one step further: Of the many things that are now called “world models,” which functional modules actually make up this capability? What are their respective uses? Language models give machines strong control over concepts, vocabulary, and reasoning, but the physical world, whether virtual or real, operates on a completely different basis. The language model learns the statistical structure of text, and the world model learns the statistical structure of space and time: how light falls on a surface, what a garden looks like from an angle never captured by a camera, and how responsive objects are and follow the laws of physics. This makes “world model” one of the most important and most misused terms in AI today. Computer vision, robotics, reinforcement learning, and generative AI all claim to be modeling the world, but they each refer to very different things. A video model that can generate gorgeous but physically impossible flames, a language model that improvises playable games, and a physical engine that faithfully simulates the combustion process are all called by the same name. The ancient Greeks were never able to agree on what constituted the world, whether it was fire, water, or inseparable atoms, because the “world” was never a single thing. It's always an alternative word used by a certain thinker to reason about a certain generality. AI inherits the same problem, and it just happened at a time when accuracy was most needed in this field. To clear this confusion behind the closed loop of taxonomy, we can start with a map that is older than all of the techniques described above. All reinforcement learning materials, including the classic Sutton and Barto, have used variants of the same picture to describe how agents interact with the world for decades. The official name of this map is Partial Observable Markov Decision Process (POMDP...

46d ago谢伟伦#AI #Omniverse #robots #Li Feifei
The “Everything Is Possible World Model”: How can a vague concept support the $10 billion financing narrative?

The “Everything Is Possible World Model”: How can a vague concept support the $10 billion financing narrative?

Author: Motion Detective BeatingOriginal title: Everything Is Possible World Model The term World Model is almost being sold out by investors. Li Feifei's World Labs completed financing of 1 billion US dollars in February this year, with a valuation of 5 billion US dollars. A year ago, it was only valued at 1 billion dollars. Yang Likun's new company, AMI Labs, raised around a billion dollars in seed round, setting the record for the largest seed round in the history of European AI startups. There were 25 cases of financing related to the World Model in the first quarter in China. Some companies took two consecutive rounds of 2.5 billion dollars in a month, and the valuation jumped from 5 billion to 10 billion dollars. What's even more surprising is that don't look at investor FOMO like this. Currently, the entire AI industry has yet to agree on what the four words “world model” actually mean. It's a word that hasn't even been defined, and it's already worth tens of billions. Every day they shouted slogans to find an anti-consensus, and in the end, they invested their money in the absence of consensus, and then called it an outlet. However, I recently heard from a few big VC investors that they could clearly see the bubble in this direction themselves, and it wasn't that no one had taken a picture of the table during the internal rehearsal session and discussed it back and forth. The conclusion was that they still had to vote. If you don't vote, next year's LP will ask you why you missed the world model; if you vote, even if you make a mistake in the end, it will be the whole industry's fault. However, when the money is hot enough, it is time to ask an impolite question. Since no one can say exactly what it is, what exactly is everyone voting for? First of all, let's be fair about the process. There was something really about this concept. In 2018, two researchers published a paper titled “World Models.” They let AI create a dream for themselves in a racing game. First, master the car in the dream, and then run back in the game. At the time, ChatGPT didn't exist, and this paper was only circulating in a small circle of researchers. Yang Likun has actually been adhering to this research direction for a long time. In the years when all of Silicon Valley bet money on the big language model, he repeatedly reiterated his view that by predicting the next word, machines would never be able to touch human intelligence, so it must be made to understand the physical world. I've been saying this for almost ten years, but the wind hasn't blown this way for him. People like him didn't turn away after hearing the wind; in their perception, there really is something about the world model. However, when “something really does” enter the venture capital industry, it usually has to go through a process first. This process will process a research direction into a term that can be wholesale. The venture capital community's favorite has never been a technical concept; it's the franchisability of a technical concept. The world model is a child of choice in this regard. It is more technological than the “metaverse,” sexier than “spatial intelligence,” broader than “embodied intelligence,” and fresher than “multi-modal.” Most importantly, it's very difficult to falsify. In BP, the harder it is to prove falsification, the more valuable it is, because not being able to falsify means that new investors can be found to take over in the next round. What the story earns is money that cannot be falsified. The world model was folded, and it had the current splendor. If you make a game, say you are a model of the world; if you make a short story, you say you are a model of the world; if you make a video tool, you say you are a model of the world; if you make a simulation robot, you say you are a model of the world. Further on, those who make advertising materials, educational courseware, metaphysics fortune-telling, and virtual people to chat with will be able to enter the world model circle as long as they dare to blow it. Earlier, at an event, I heard investors share in a round table. Each field of medicine, finance, and law can be viewed as an independent world. According to this usage, my car repair master downstairs also has a world model in his mind. It specifically predicts when the Third Ring Road will be blocked. The accuracy rate is higher than most assisted drivers I've ever used. Not long ago, I also saw a robotics company announce that it is building a model of the industrial world and a model of the home world at the same time. I realized that the world used to be countable terms; they can be sold individually. It's not that no one has anticipated this grand event; the identity of the person who anticipated it is quite special. In March of this year, on the day AMI Labs funded that billion dollars, the company's CEO told the media that I predict “world model” will be the next buzzword. Within six months, every company will call itself World Model to finance. He was right. The only thing that wasn't accurate was the time; it didn't take six months at all. Habitual narratives Chasing narratives has long been a habitual act of investors and entrepreneurs. On May 10, 2015, a listed company whose main business is floor tiles and real estate issued an announcement saying that it wants to become the first internet finance company in China, it wants to change its name to “Pitumpi”, and the English name is directly registered as P2P Financial Information Ser...

47d agoburnking#AI #financing

World Labs, founded by Li Feifei, raised $1 billion, and Nvidia participated

Comparatively, according to investment reports, World Labs, a startup founded by Li Feifei, has raised $1 billion in a new round of financing to pursue a novel approach to AI development. As part of this funding round, Autodesk Inc. invested $200 million in World Labs. Other supporters include Andreessen Horowitz, Nvidia, and AMD, according to World Labs. World Labs focuses on developing “World Models (World Models)” to enable AI to understand and make decisions about the three-dimensional physical world.

183d ago

Li Feifei's AI startup World Labs raised $1 billion, with NVIDIA and others participating

Comparing news, according to official sources, Li Feifei's AI startup World Labs raised $1 billion, and AMD, Autodesk, Emerson Collective, Fidelity Management & Research Company, NVIDIA, and Sea participated. World Labs is committed to accelerating the development mission of spatial intelligence by constructing models of the world to innovate narratives, stimulate creativity, promote robotics, and promote scientific discovery. World Labs says its first product, Marble, allows anyone to create spatially coherent, high-fidelity, and long-lasting 3D worlds from images, videos, or text.

184d ago