
A new survey in the journal Robot systematically reviews how large language models (LLMs) and multimodal large models are reshaping intelligent robot perception, navigation, and manipulation. The authors, affiliated with Huazhong University of Science and Technology and other institutions, argue that these models drive a paradigm shift from “perception-driven” to “cognition-driven” robotics, enabling deeper environmental understanding and autonomous decision-making.
In perception, multimodal fusion and language-spatial joint reasoning allow robots to grasp both semantic and geometric attributes of scenes. For navigation, chain-of-thought task decomposition and commonsense reasoning help parse ambiguous instructions and support autonomous exploration in unknown environments. In manipulation, vision-language-action (VLA) models coupled with physical commonsense improve dexterity and adaptability in complex interactive tasks.
The survey identifies four key technologies in embodied perception: multimodal fusion, VLA decoding and representation, language-space active reasoning, and Sim2Real generalization. It also categorizes fusion methods based on point cloud, image, and heterogeneous representations, noting trade-offs in accuracy, computational cost, and sensor synchronization.
Key takeaways
- Large models enable robots to move beyond pre-programmed logic toward context-aware, adaptive behavior in unstructured settings.
- VLA models integrate visual, linguistic, and action data to support flexible task execution and real-time strategy adjustment.
- Core challenges remain: cross-modal alignment accuracy, real-time performance, safety and reliability, and Sim2Real generalization.
- The paper provides a technical reference and development roadmap for building general-purpose, cognition-enhanced robotic systems.
Source: 《机器人》期刊 (robot.sia.cn) · Published 2026-07-15 · “大模型与智能机器人的融合:智能感知、导航与操作”
