Dialogue with Zhang Lei of IDEA:"A Genuine World Model Must Be Action-Conditioned"
.png)
Outperforming Google and Meta to deliver world-class object understanding capabilities
Author: Xing Lijuan
Editor: Lin Juemin
In the first half of 2026, capital and talent across China’s AI industry converged toward one core track, world models.
In the first six months alone, disclosed financing for China’s world model sector reached approximately 300 billion RMB. The figure exceeded 900 billion RMB when counting the integrated track of embodied intelligence and world models.
Year-on-year financing growth surged more than five times.
Massive capital and top-tier talent flooded into this emerging field.
An inevitable question arises. With countless practitioners rushing into the track, how many have truly clarified the essential definition and value of world models?
Practitioners are exploring technical directions. Investors are screening outstanding teams. Headhunters are sorting industrial layouts. Corporate strategy departments are assessing long-term survivors in the sector.
A widespread industry confusion remains consistent. No consensus has been reached on how to define and implement world models.
Clear answers do exist.
This in-depth interview features Zhang Lei, Distinguished Scientist at IDEA Research Institute and Founder of Visincept, to systematically clarify the essence, technical routes and industrial landing logic of world models.
Zhang Lei is an IEEE Fellow with over 77,000 Google Scholar citations. His DINO series models once topped the COCO benchmark for five consecutive months. Grounding DINO and its upgraded version DINO-X have outperformed mainstream models from Google and Meta. They have become standard reference tools adopted by top global institutions including Fei-Fei Li’s research team, NVIDIA and Galaxy General Robotics.
In 2025, he spun off his research team from IDEA Research Institute to found Visincept. The company secured nearly 100 million RMB in angel funding within just one month.
He is guided by two preeminent AI scholars. His academic advisor is Academician Zhang Bo, one of the founding pioneers of Chinese artificial intelligence. His industrial mentor is Academician Harry Shum, former Global Executive Vice President of Microsoft and former Managing Director of Microsoft Research Asia.
These remarkable credentials are not the core reason for this interview.
The core motivation is straightforward. Amid chaotic industry definitions and heated debates over world models in 2026, Zhang Lei is one of the very few scholars and entrepreneurs who has put forward a set of clear, executable and non-compromising industrial standards for world models.
From Principal Researcher at Microsoft Research to Distinguished Scientist at IDEA Research Institute and now AI entrepreneur, Zhang Lei’s career path has transformed multiple times. However, his core research direction has never changed, which is empowering AI with the ability to understand physical objects.
World models represent the core industrial answer to his two-decade-long research proposition of object-centric AI understanding.
For every practitioner and investor following the world model track, this interview is worthy of in-depth reading. It helps distinguish technological substance from industrial hype, clarify the winning routes of technical competition, and identify high-value teams worthy of long-term layout.
The following is the edited interview transcript polished by AI Tech Review while retaining all original viewpoints and core information.

Dr. Zhang Lei. IEEE Fellow, Founder and CEO of Visincept, Distinguished Scientist at the Center for Computer Vision and Robotics (CVR) of IDEA Research Institute, Adjunct Professor at Hong Kong University of Science and Technology (Guangzhou), former Principal Researcher at Microsoft Research Asia and Microsoft Redmond Research Center. He has published more than 200 academic papers in computer vision and related fields. His Google Scholar citations exceed 77,000 with an H-index of 107. He holds more than 60 authorized U.S. patents. He was elected as an IEEE Fellow for his outstanding contributions to large-scale image recognition and multimedia information retrieval.
01 The Biggest Bottleneck for Embodied Intelligence Lies in the Brain Instead of Hardware Bodies
AI Tech Review: Embodied intelligence has boomed since last year. Robotics companies have made remarkable progress in motion control. Unitree robots can complete marathon runs and Agibot robots can perform professional backflips. Nevertheless, large-scale commercial landing still faces obvious obstacles. What core challenges need to be solved for embodied intelligence to move from laboratory demos to real industrial scenarios?
Zhang Lei: Robots need to solve two fundamental problems to achieve industrialization, task success rate and generalization capability. Success rate refers to the stable execution capability of a single task. Generalization capability means the robot can complete the same task stably under changed environments or different hardware carriers.
Four core challenges restrict the industrialization of embodied intelligence.
First, insufficient generalization in single scenarios. Many teams can achieve a task success rate of 70 percent or even 80 percent in fixed laboratory environments. This seems to be a major breakthrough. An 80 percent success rate in real industrial scenarios means one failure in every five operations. It still requires continuous manual intervention and cannot support fully autonomous deployment.
Second, poor cross-scenario generalization capability. The industry should first maximize the performance of single-task capabilities and then expand to diverse complex scenarios. This enables robots to execute richer and more complex tasks in variable environments. This is the biggest gap between laboratory verification and industrial landing.
Third, insufficient high-level cognitive capabilities of the intelligent brain. Current deep learning technologies mainly learn intelligent behaviors rather than the underlying operational mechanisms of real intelligence. A qualified embodied intelligent brain needs real physical causal reasoning, common sense understanding and adaptive capabilities. It can identify logical correlations behind phenomena instead of merely statistical rules. The industry is still in the initial exploration stage in this field.
Fourth, far from terminal-level independent application capabilities. Most existing robots rely on massive external computing power. The core problem to be solved is getting rid of full-time manual intervention. Only in this way can robots achieve large-scale scenario promotion and form a positive closed loop of application, data accumulation and iterative optimization.
AI Tech Review: The industry holds divided views on this dilemma. Hardware teams believe China’s mature supply chain advantages can prioritize reducing robot body costs. Algorithm teams insist that the intelligent brain is the core bottleneck restricting industrial progress. What is your judgment?
Zhang Lei: The intelligent brain is more difficult to break through and more strategically important.
China’s hardware iteration capabilities are impressive. Enterprises including Unitree and Agibot optimize robot motion control through simulation reinforcement learning. They train control algorithms in virtual environments and migrate mature policies to physical robots. Continuous optimization of motors, transmission structures and other core components has greatly reduced hardware costs and improved overall stability. This is a very solid industrial path. Even so, the flexible movements displayed by robot hardware in marathons and performance scenarios only simulate the basic athletic abilities of animals. Robot bodies still cannot match the natural reflexes, obstacle avoidance and environmental interaction capabilities of living creatures.
AI Tech Review: You emphasize that building an intelligent brain is more challenging. Does the difficulty lie in insufficient data accumulation or the industry’s superficial understanding of intelligent mechanisms?
Zhang Lei: The core difficulty is that we have not yet fully mastered the underlying mechanisms of real intelligence. The human brain’s high-level cognitive ability is fundamentally different from that of other creatures. All current deep learning breakthroughs only simulate the external behavioral characteristics of human intelligence through algorithms. AI essentially learns human behaviors rather than the essence of human intelligence. As Academician Zhang Bo often mentions, large language models can communicate in human language, but their language comprehension logic is completely different from human cognitive patterns.
02 Why Have World Models Gained Explosive Popularity After Two Decades of Silence?
AI Tech Review: World models have become a core hot track alongside the outbreak of embodied intelligence. Industry insiders know that the concept of world models was proposed as early as the 1990s. Why did this cutting-level technological concept remain silent for two decades and suddenly become the core outlet of AI innovation in recent years?
Zhang Lei: This is a valuable question. The world model is never a new concept. In the early 1990s, researchers tried to build environmental simulation models for agent interaction to optimize reinforcement learning algorithms. The concept lacked effective technical verification and landing conditions due to the immaturity of deep learning technology at that time.
The real technological verification of world models took place around 2018 and 2019. The landmark paper on World Models realized effective model application in game scenarios. Game environments are the most ideal verification carriers for reinforcement learning algorithms.
After years of iterative development of reinforcement learning, AI agents were already able to outperform human players in video games. Researchers then began to explore new possibilities. Can AI agents autonomously learn and build environmental world models through continuous interaction with virtual environments to realize autonomous evolution of environmental prediction capabilities? This technical assumption was fully verified between 2018 and 2020.
AI Tech Review: What is the essential connection between world models and the current embodied intelligence boom? Is embodied intelligence the core trigger for the popularization of world models?
Zhang Lei: We can clarify the correlation starting from the iteration logic of language models. The industry first launched Vision-Language-Action models for robotic scenarios. It follows the imitation learning paradigm of large language models. Language models complete next-token prediction tasks. Embodied VLA models observe real-time visual frames and predict the next action of robots.
This paradigm has quickly hit a development bottleneck. Insufficient data is only a minor factor. The core limitation is the underutilization of reinforcement learning capabilities. I have a simple analogy. Imitation learning enables large language models to speak like humans. Reinforcement learning enables them to speak correctly and logically.
Similarly, the intelligence level and action success rate of robots rely heavily on reinforcement learning iteration.
However, applying reinforcement learning to embodied scenarios faces two fatal flaws. First, data collection efficiency is far lower than that of language models. Robots need to execute physical actions in real environments and wait for real physical feedback. Second, the trial-and-error cost is extremely high. A robot learning to wash dishes may break a large number of tableware. Autonomous driving models need to experience massive real accident scenarios to learn avoidance logic. This is completely different from the low-cost trial-and-error iteration of language models in pure digital environments.
The virtual simulation environment constructed by world models perfectly solves these pain points. If a model can predict environmental changes triggered by specific actions, robots can complete countless virtual trial-and-error iterations in the digital world. It no longer needs to consume massive physical resources for real-scene verification.
(Author’s note: The core logical chain is clear. Robots rely on massive trial and error to master operational skills. Real-world trial and error is costly, inefficient and risky. World models build a virtual training system for robots to pre-deduce action consequences. Robots only execute optimal strategies in real scenarios after completing sufficient virtual iterations. This explains why world models have become the core strategic layout of the global AI industry.)
03 A Qualified World Model Must Take Actions as Input Conditions
AI Tech Review: The industry currently has extremely confused definitions of world models. Video generation models, 3D spatial reconstruction models and LeCun’s JEPA architecture are all labeled as world models. What core criteria must a model meet to be recognized as a standard world model?
Zhang Lei: I insist on one core standard. A real world model must be action-conditioned. It needs to take specific actions as input and predict corresponding environmental changes. It must answer the core question of how the environment evolves after executing a certain action.
Action is the core premise of interaction between intelligent agents and the physical world. Agent actions change environmental states. Updated environmental states feed back to the agent to judge whether the current behavior is close to the target. This forms a complete closed-loop logic. Every environmental state transition depends on the specific actions executed in the previous step.
Many people define world models as upgrading language model next-token prediction to next-state prediction. This definition is incomplete and lacks core constraints. The accurate definition is action-conditioned next-state prediction. Only with action constraints can state prediction have practical industrial value and meet the essential needs of reinforcement learning.
AI Tech Review: Many investors and industry practitioners regard video generation tools such as Sora as world models. Obviously, these models do not have action input dimensions according to your definition. How do you view this industry misunderstanding?
Zhang Lei: The world models we discuss originate from model-based reinforcement learning systems. Their core value is to assist intelligent agents to interact efficiently with the physical environment. World models have multiple application scenarios, and embodied intelligence is currently the most valuable and landing-oriented track.
From this perspective, pure video generation models represented by Sora cannot be defined as world models. They only generate logically consistent pixel sequences and predict visual changes. They do not model the interactive actions between agents and the environment at all. Such models can learn partial physical change rules indirectly, but they cannot effectively assist robots in real environmental interaction tasks.
AI Tech Review: The recently popular World Action Model also incorporates action modeling. Is it a new category of standard world models?
Zhang Lei: WAM does not conform to the definition of world models in reinforcement learning. It uses future frame pixel supervision to optimize action prediction and achieves better performance than traditional VLA models in partial tasks. However, its logical sequence is reversed. It generates future pictures first and then deduces corresponding actions. This reverse modeling logic cannot support reinforcement learning closed-loop iteration, so it is not a standard world model.
04 Pixel Domain versus Latent Domain. Two Technical Routes Converge in Essence
AI Tech Review: The competition between pixel-domain and latent-domain routes is the hottest topic in the current world model track. LeCun is the representative of the latent-domain school and Sora represents the pixel-domain school. The two routes seem completely opposed. What is your judgment on this technical dispute?
Zhang Lei: We should first recognize the essential consistency between the two routes.
The industry’s perception of opposition mainly comes from LeCun’s clear academic propositions. Scholars usually express their views extremely clearly to form industry recognition.
In fact, pixel-domain models also rely on latent-domain processing. Both Stable Diffusion and Sora compress original images and videos into low-dimensional latent feature spaces for subsequent calculation. The optimization of feature space is the core essence of technological iteration. The real divergence of the two routes lies in the retention and filtering rules of latent space features instead of the simple choice of whether to use latent space.
The latent-domain route filters out irrelevant pixel interference such as light and texture. It focuses on extracting essential physical change rules. Its core risk is representation collapse. Different physical scenes may be mapped to the same latent feature, causing the model to lose the ability to distinguish environmental states. The pixel-domain route pursues ultra-high visual fidelity. It essentially adapts to human visual perception habits but ignores physical rationality.
AI Tech Review: Human brain simulation also relies on latent space deduction. Humans will not perform pixel-by-pixel rendering when imagining action consequences. Does this mean latent-domain routes are more in line with real intelligent logic?
Zhang Lei: Exactly. If the human brain has an internal world model, all counterfactual reasoning and environmental deduction are completed in abstract latent space instead of detailed pixel simulation. Therefore, human visual fidelity cannot be used as a core evaluation standard for world models. The core optimization goal of pixel-domain world models is not picture fidelity, but the physical rationality of generated content.
(Author’s note: This paragraph clarifies the core technical divergence in the industry. Pixel-domain models such as Sora pursue visual realism and often produce physical errors including floating objects and penetrating collisions. Latent-domain models represented by LeCun and Zhang Lei abstract scenes to deduce physical causality. Zhang Lei’s core innovation is introducing object structure modeling into latent space. The model first identifies independent physical objects and then learns interactive rules, which avoids visual errors and adapts to real robot operation scenarios.)
AI Tech Review: Your team adheres to the latent-domain technical route consistent with LeCun. Your research adds object structure modeling on this basis. What is the essential difference between your technical system and LeCun’s framework?
Zhang Lei: Our route shares the same first principle with LeCun. Human counterfactual reasoning is action-conditioned and completed in abstract latent space rather than original pixel space.
LeCun has verified this principle through I-JEPA, V-JEPA and LeJEPA series research. He confirmed that predictive reasoning should be carried out in latent feature space. However, physical scenarios for embodied tasks are extremely complex. Large-scale industrial landing requires structured feature constraints in latent space.
This advantage comes from our team’s verified technological accumulation. Our self-developed DINO-X model has solved the open-world object perception problem. Therefore, we take independent physical objects as the basic prediction and planning units of world models. Physical rules act on individual objects. Latent space with object awareness can learn physical laws more efficiently and accurately.
AI Tech Review: LeCun’s AMI Labs has obtained billions of US dollars in financing to layout the same latent-domain track. Why do you believe your object structure enhanced route has better industrial prospects?
Zhang Lei: Our DINO-X series has achieved world-leading object understanding capabilities. We have in-depth and systematic cognition of visual model mechanisms and physical rationality logic. This technical accumulation is highly compatible with world model iteration.
In short, LeCun provides the basic latent-domain theoretical framework. Visincept innovatively adds object structure modeling and realizes real-world closed-loop verification. This industrial closed-loop capability is our unique core barrier accumulated through years of technical iteration.
05 How Does the DINO Series Surpass Google and Meta’s Technical Solutions?
AI Tech Review: Your team’s visual foundation models have achieved industry-leading results. Grounding DINO and DINO-X have surpassed Google and Meta’s mainstream models and become the universal reference standard for top institutions including NVIDIA and Fei-Fei Li’s team. It is rare for Chinese AI teams to achieve comprehensive overtaking of overseas giants. What is your core secret?
Zhang Lei: Our technical advantages are mainly reflected in two dimensions.
In terms of paradigm innovation, we are the first team in the world to realize SOTA-level object detection based on Transformer architecture. Grounding DINO introduces large language model logic into visual detection and realizes open-world detection breakthroughs. The model is no longer limited to fixed training categories and can identify any unseen new objects. It remains the most cited and downloaded open-source detection model globally, representing a milestone paradigm innovation.
In terms of engineering iteration and data strategy, we switched to closed-source independent iteration after Grounding DINO. We continue to scale and verify innovative technologies. I have always maintained high standards for team technological output. Our young team insists on releasing only version-level disruptive breakthroughs instead of trivial optimizations. We spent one year polishing Grounding DINO 1.5 and another six months completing DINO-X iteration. The highly polished iterative versions have been widely recognized by the global industry.
AI Tech Review: What is the internal connection between your previous visual foundation model research and the current world model layout?
Zhang Lei: Our research has always maintained a coherent technical context. Previous object perception research laid the core perceptual foundation for robot embodied intelligence. A large number of industry teams apply Grounding DINO to solve environmental understanding problems in embodied scenarios.
Technically, world models are visually native systems rather than language native systems. Almost all mainstream world model research does not rely on pure language architectures such as GPT.
Our in-depth understanding of visual model mechanisms is highly compatible with the technical logic of general object understanding and world model iteration. This is the core confidence for us to continuously iterate world model technologies.
Our core team of more than 20 researchers was incubated from IDEA Research Institute with extremely high stability. During closed-source model iteration, we continuously absorb top algorithm and engineering talents. This stable core team has become the foundational strength of Visincept’s world model research.
06 How to Enable AI to Truly Understand Physical Objects
AI Tech Review: You mentioned that most current latent-domain models cannot truly understand physical objects. They cannot distinguish independent objects or judge interactive relationships between objects. How can we verify whether a model has real object understanding capabilities instead of superficial statistical fitting?
Zhang Lei: Intelligent understanding itself cannot be directly quantified and detected. The industry cannot fully explain the language understanding mechanism of large language models. We can only judge model capabilities through external behavioral performance. We recognize model comprehension when its output logic is close to human expression habits.
World models face the same verification dilemma. We cannot directly observe whether the model masters physical laws and object structures. We mainly analyze model capabilities through intermediate visualization results. Highly consistent output with human intuition proves that the model has captured effective object structural features.
AI Tech Review: Are there standardized experimental verification methods to test model object understanding capabilities?
Zhang Lei: Counterfactual testing is the most effective verification method. We input different hypothetical action instructions to observe whether the model can feed back physically reasonable environmental changes. Stable and correct prediction under slight trajectory changes proves that the model has learned essential physical rules. Continuous prediction errors indicate that the model only fits surface statistical correlations.
This requires extremely acute research insight. As Richard Sutton said, researchers always tend to artificially implant empirical rules into models during research. Early NLP researchers tried to hard-code syntactic analysis and word segmentation rules, but finally lost to the simple and efficient next-token prediction paradigm.
I firmly believe in the iterative power of concise and original technical paradigms. Excessive manual rule setting can bring short-term effects but will hinder long-term technological iteration and breakthroughs.
AI Tech Review: Special scenes such as flowing sand, running water and smoke have no clear object boundaries. Will your object-centered technical framework be limited by discrete object definitions?
Zhang Lei: Our team has long considered this problem. Technically, we adopt mask segmentation as the basic entry point to extract object pixel areas. The industry’s mature semantic understanding and mask prediction technologies can already handle non-rigid objects such as flowing water and quicksand.
Mask-based processing is a pragmatic and feasible technical path. Our long-term accumulation in general object perception gives us unique advantages in processing complex non-rigid physical scenes.
07 EgoTwin. Empowering Robot Manipulation via Human Hand Demonstrations
AI Tech Review. A core innovation of your technical framework is action alignment. It maps motion data from diverse hardware systems including human hands, two-fingered grippers and dexterous multi-fingered hands into a unified latent space. This mechanism essentially acts as a translation system for robot action logic. Would you agree with this analogy?
Zhang Lei. World models rely on two core input modalities, visual state observations and actionable control signals. Action understanding therefore serves as a foundational capability for embodied intelligence.
Action alignment solves the compatibility challenges across heterogeneous hardware devices. The most difficult scenario lies in the alignment between human hands and robot manipulators. Robotic arms output precise XYZ coordinate data in physical space, while most human hand demonstration data is collected from 2D daily videos without accurate 3D ground truth. Direct mixed training of such heterogeneous data leads to extremely low efficiency and poor generalization.
Our solution extracts 3D hand keypoints from 2D demonstration videos. It converts ordinary human hand motion footage into standardized 3D action sequences, which can be perfectly aligned with robot arm data within a unified physical space. This addresses the widespread industry challenge of collecting and processing egocentric human demonstration data.
Our self-developed EgoTwin platform precisely targets these data acquisition and alignment pain points. We are currently conducting in-depth technical cooperation with the data team of Baidu Cloud to iterate and upgrade this system.
AI Tech Review. EgoTwin effectively solves the data acquisition problem for physical hand movements. How does your framework capture high-level human intent behind these movements?
Zhang Lei. High-level intent reasoning is undertaken jointly by world models and VLA models. High-quality standardized action data remains an indispensable prerequisite for accurate intent understanding. The industry invests heavily in large-scale data collection yet frequently suffers from low data utilization efficiency.
Our experience in developing Grounding DINO and DINO-X proves that algorithm researchers must participate deeply in data curation. Purely collected datasets often fail to match algorithm iteration requirements and become unusable during formal training. Algorithm optimization and data refinement must advance iteratively and synchronously to form a positive closed loop.
08 Two Mentors, One Consistent Academic Ethos
AI Tech Review. Two distinguished scholars have profoundly shaped your academic and entrepreneurial journey. Academician Zhang Bo is a foundational pioneer of Chinese AI focusing on academic research and talent cultivation. Academician Harry Shum is a global industry leader with rich experience in technological industrialization. What core lessons have you learned from their respective guidance?
Zhang Lei. Academician Zhang Bo has always been my lifelong mentor and role model. I first met him during my undergraduate studies at Tsinghua University, when I joined his laboratory focused on automated guided vehicles. I spent long hours debugging hardware and conducting repeated experiments. I often observed him guiding PhD students in robotics research with rigorous patience and sincere curiosity. I was fortunate enough to become his doctoral student later.
What impressed me most is his eternal curiosity and openness toward emerging technologies. Whenever I presented new experimental results, he would carefully review every detail and propose targeted parameter adjustment suggestions to explore unknown possibilities. He embodies the pure curiosity-driven spirit of scholarly research.
One principle I have upheld ever since is thorough understanding of underlying mechanisms. Superficial cognition always leads to misleading judgments. True technological mastery comes from grasping essential principles rather than memorizing superficial phenomena. This has become my core criterion for research and student training.
Academician Harry Shum’s influence on my career trajectory is equally profound. He was my senior leader during my tenure at Microsoft and has shaped my entire research career. He recruited me to conduct research in the United States and later encouraged my return to Shenzhen to engage in original innovative research. He possesses sharp industrial insight and has provided crucial practical guidance for my team building, technological layout and fundraising throughout my entrepreneurial journey.
AI Tech Review. We are deeply touched to learn that Academician Zhang Bo, now in his nineties, still visits your team regularly for technical exchanges.
Zhang Lei. That is true. Last November, Tsinghua Graduate School invited him to deliver lectures in Shenzhen. After finishing his morning lecture, he devoted the entire afternoon to in-depth technical discussions with our team and joined us for further exchanges during dinner. He maintains an extremely broad academic vision and has offered valuable feedback on our visual understanding and robotic intelligence research.
Despite his advanced age, his logical thinking remains rigorous and sharp. His analytical depth is comparable to active young researchers in the industry.
09 Embrace Technological Waves. Opportunities Favor the Prepared Mind
AI Tech Review. You have deeply participated in multiple transformative AI waves including deep learning, foundation models, embodied intelligence and world models. Is there a defining moment that made you firmly determined to focus on world model research?
Zhang Lei. My research on object detection brought me a strong sense of achievement, as it drove tangible progress across the entire computer vision field. The ongoing world model research brings me the same passion and sense of mission.
Decades of iterative progress in reinforcement learning and artificial intelligence, together with the mature paradigm of large language models, have accumulated abundant reusable technical insights. I firmly believe that world model technology will reshape the entire AI industry and unlock the next stage of embodied intelligence industrialization.
AI Tech Review. The industry widely celebrates researchers’ breakthrough achievements yet rarely discusses their setbacks and hardships. Have you encountered critical technical bottlenecks during your research journey?
Zhang Lei. Absolutely. Around 2018, I dedicated extensive efforts to weakly supervised object detection. I aimed to build detection models trained only on image-level category labels without manual bounding box annotations. The technical path was logically feasible, but I invested tremendous time and energy without obtaining fully satisfactory results.
Later at IDEA Research Institute, we broke through this bottleneck from an innovative perspective by integrating language understanding and detection capabilities within Grounding DINO. This breakthrough was built on the solid technical accumulation of our DINO detection series and benefited greatly from the booming iteration of large language models.
Grounding DINO matured around late 2022 and early 2023, coinciding with the official launch of ChatGPT. This verifies a core industry rule. Sufficient data scale and iterative accumulation will inevitably trigger intelligent emergence, and mature algorithms will become increasingly concise and efficient.
Setbacks are inevitable in innovative research. What matters most is persistent dedication to core problems. If one technical path fails, researchers should adjust their thinking and explore new directions. Core research questions deserve long-term and continuous exploration until qualitative breakthroughs emerge naturally.
AI Tech Review. A new generation of young entrepreneurs is emerging in the world model track, including many undergraduate founders. What is your view on these emerging young practitioners?
Zhang Lei. This phenomenon indicates that the entry threshold for AI research has gradually lowered. Before the deep learning era, newcomers needed years of systematic foundational training to participate in cutting-edge research. Today’s open-source technical ecosystem and universal Transformer architecture enable young researchers to quickly enter the track with first-principles thinking and innovative perspectives.
Nevertheless, tackling fundamental and hard industrial problems still requires experienced researchers with profound accumulation, precise directional judgment and in-depth technical insight.
Industry trends keep changing rapidly. Blindly chasing hot tracks rarely leads to sustained success. Regardless of technological iterations and trend shifts, outstanding talents with solid foundational capabilities and continuous learning ability will always be sought after by the industry. Cutting-edge opportunities always favor well-prepared practitioners.
10 Epilogue. The Power of Thorough Understanding
The most touching takeaway from this interview is not Zhang Lei’s impressive technical achievements or his influential academic heritage. It is a simple principle he has adhered to for three decades, thorough understanding.
This philosophy has accompanied his entire research career.
Back in the early 1990s, artificial intelligence was still a niche and overlooked discipline in China. As an undergraduate student majoring in computer science at Tsinghua University, Zhang Lei devoted most of his spare time to debugging AGV hardware and conducting repeated experiments in the laboratory.
At that time, Academician Zhang Bo served as the laboratory director and supervised doctoral research on robotics. His rigorous academic attitude and sincere curiosity deeply influenced young Zhang Lei, who later became his doctoral student.
Whenever Zhang Lei delivered new experimental results, Academician Zhang Bo would carefully examine every detail and propose tentative adjustments. He never treated student research as routine assessment. Instead, he explored unknown technical possibilities with genuine curiosity.
He passed down a lifelong belief to Zhang Lei. Superficial cognition leads to misleading judgments. Real technological mastery lies in the thorough grasp of underlying principles.
This embodies the precious craftsmanship of China’s older generation of scientists.
In 2018 at Microsoft, Zhang Lei encountered a persistent technical bottleneck in weakly supervised object detection. Facing repeated failures, he never abandoned this valuable research problem.
He always believes that failed attempts are valuable accumulation. Researchers must keep thinking about core problems day and night. Meaningful technological breakthroughs never come from accidental attempts, but from long-term precipitation and persistent exploration.
His persistence finally paid off at IDEA Research Institute. The launch of Grounding DINO opened up a new technical path, coinciding with the industrial outbreak of large language models. The expansion of data scale and the maturity of algorithm paradigms triggered the emergence of new capabilities.
Zhang Lei describes this process as natural intelligent emergence.
Thirty years separate his early AGV hardware exploration and today’s world model research. From large-scale image retrieval and general visual understanding to embodied intelligent world models, every stage of his research focuses on exploring essential principles rather than chasing superficial trends. As Steve Jobs once said, you can only connect the dots looking backward. Every earnest exploration and persistent effort will eventually generate unique value.
At the end of 2025, ninety-year-old Academician Zhang Bo traveled to Shenzhen for academic lectures. Immediately after his morning speech, he spent the entire afternoon communicating with young researchers at Visincept.
He shared his research experience and unique insights into industry trends with the team. His thinking remained rigorous, clear and forward-looking. Zhang Lei systematically introduced the team’s latest progress and technical layout in world model research to his mentor.

Academician Zhang Bo visits Visincept. Zhang Lei presents the team’s world model research progress to his mentor
Thirty years ago on the Tsinghua campus, a senior scholar accompanied a young student to explore unknown technical boundaries and pursue essential principles.
Thirty years later in a Shenzhen technology office, the grown-up student explains cutting-edge world model innovation to his elderly mentor.
This full-circle inheritance scene is deeply touching.
Talents grow old, industrial bubbles burst, and technological trends fade away. Yet the persistent pursuit of Chinese researchers for essential principles and core technological breakthroughs will never cease.