Abstract:
Dynamic unstructured construction environments are characterized by open sites, time-varying working conditions, discrete material properties, and complex interactions among humans, robots, and construction objects. These characteristics make it difficult for construction robots to directly follow the development path of industrial robots, which typically operate in fixed workstations with deterministic processes. To address the transition of construction robots from automatic execution of single tasks to autonomous operation in complex on-site environments, this paper focuses on skill learning methods for direct construction manipulation tasks. It reviews recent progress, task adaptation logic, and engineering bottlenecks of deep reinforcement learning, imitation learning, transfer learning, and multi-agent learning. The analysis shows that skill learning for construction robots has shifted from fixed program execution to feedback-based policy optimization, demonstration-driven skill acquisition, cross-scenario skill reuse, and multi-agent collaboration. Differentiated methodological pathways have gradually emerged across typical construction tasks: masonry and assembly tasks are more suited to learning pathways that integrate imitation-learning-based initialization, local optimization through reinforcement learning, and semantic priors derived from building information modeling; continuous concrete construction operations rely more on reinforcement-learning-based or learning-enhanced control with multimodal process feedback; and demolition, maintenance, and earthwork tasks require tighter integration of safety constraints, demonstration priors, and mechanism-based models. However, existing studies are still generally in the transitional stage from proof-of-concept validation to engineering deployment, and commonly face challenges such as low sample efficiency, difficulties in sim-to-real transfer, insufficient skill generalization, and a lack of long-term validation. Future research should focus on embodied-AI-driven task modeling, few-shot safe learning, multimodal state representation, and semantic–action collaborative reasoning, so as to support stable applications in complex construction sites.