语义有效性:生成样本属于目标设计对象的判断(目标类别、功能暗示与主体可识别性)及语义有效率定义
灵巧手与功能性抓取合成研究
这些文献专注于灵巧手(Dexterous Hands)的建模、抓取合成算法、功能性操作及相关数据集的构建,核心在于如何实现类人的复杂抓取与操作能力。
- A Multi-modal Hand Imitation Dataset for Dexterous Hand(Shaochen Wang, Qilin Wu, Kang Chen, Qing Huang, Zhuo Cheng, Beihao Xia, 2025, 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS))
- DexDiffuser: Generating Dexterous Grasps With Diffusion Models(Zehang Weng, Haofei Lu, Danica Kragic, Jens Lundell, 2024, IEEE Robotics and Automation Letters)
- HandX: Scaling Bimanual Motion and Interaction Generation(Zimu Zhang, Yuchen Zhang, Xiyan Xu, Ziyin Wang, Sirui Xu, Kaimao Zhou, Bing Zhou, Chuan Guo, Jian Wang, Yu-Xiong Wang, Liangxin Gui, 2026, arXiv.org)
- Grasp as You Say: Language-guided Dexterous Grasp Generation(Mark R. Cutkosky, Jian-Jian Jiang, Hao Li, Xian-Tuo Tan, Yi-Lin Wei, Xiaoming Wu, Chengyi Xing, Wei‐Shi Zheng, 2024, Advances in Neural Information Processing Systems 37)
- ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning(Kailin Li, Puhao Li, Tengyu Liu, Yuyang Li, Siyuan Huang, 2025, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- Towards Natural Prosthetic Hand Gestures: A Common-Rig and Diffusion Inpainting Pipeline(Seungyup Ka, Taemoon Jeong, Sunwoo Kim, Sankalp Yamsani, Joohyung Kim, Sungjoon Choi, 2024, 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC))
- DexFuncGrasp: A Robotic Dexterous Functional Grasp Dataset Constructed from a Cost-Effective Real-Simulation Annotation System(Jinglue Hang, Xiangbo Lin, Tianqiang Zhu, Xuanheng Li, Rina Wu, Xiaohong Ma, Yi Sun, 2024, Proceedings of the AAAI Conference on Artificial Intelligence)
- A Sensorless, Inherently Compliant Anthropomorphic Musculoskeletal Hand Driven by Electrohydraulic Actuators(Misato Sonoda, R. Hinchet, Amirhossein Kazemipour, Yasunori Toshimitsu, Robert K. Katzschmann, 2026, arXiv.org)
- Toward Human-Like Grasp: Dexterous Grasping via Semantic Representation of Object-Hand(Tianqiang Zhu, Rina Wu, Xiangbo Lin, Yi Sun, 2021, 2021 IEEE/CVF International Conference on Computer Vision (ICCV))
- Humanoid dexterous hands from structure to gesture semantics for enhanced human–robot interaction: A review(Xin Li, Wenfu Xu, Zaiqiao Ye, Han Yuan, 2025, Biomimetic Intelligence and Robotics)
- Design and Construction of a Cost-Effective Dexterous Robotic Hand for Research and Development(Shahram Mohsini, Meaghan Charest-Finn, R. Dubay, 2025, 2025 IEEE International systems Conference (SysCon))
- Grasping a Handful: Sequential Multi-Object Dexterous Grasp Generation(Haofei Lu, Yifei Dong, Zehang Weng, Florian T. Pokorny, Jens Lundell, Danica Kragic, 2025, IEEE Robotics and Automation Letters)
- Toward Human-Like Grasp: Functional Grasp by Dexterous Robotic Hand Via Object-Hand Semantic Representation(Tianqiang Zhu, Rina Wu, Jinglue Hang, Xiangbo Lin, Yi Sun, 2023, IEEE Transactions on Pattern Analysis and Machine Intelligence)
机器人形态演化与生成式设计
这些文献侧重于利用生成式人工智能(如扩散模型)和计算方法自动优化机器人(包括夹持器、软体机器人等)的物理结构与形态设计,以提升特定任务的表现。
- Structural Synthesis and Optimisation of a Robotic Gripper Using Generative AI Design(Hamid Isakhani, S. Nefti-Meziani, Steve Davis, Amir M. Hajiyavand, Xiazhen Xu, 2024, 2024 IEEE-RAS 23rd International Conference on Humanoid Robots (Humanoids))
- Generative-AI-Driven Jumping Robot Design Using Diffusion Models(Byungchul Kim, Tsun-Hsuan Wang, Daniela Rus, 2025, 2025 IEEE International Conference on Robotics and Automation (ICRA))
- Towards Controllable Generative Design: A Conceptual Design Generation Approach Leveraging the FBS Ontology and Large Language Models(Liuqing Chen, H. Zuo, Zebin Cai, Yuan Yin, Yuan Zhang, Lingyun Sun, Peter R. N. Childs, Boheng Wang, 2024, Journal of Mechanical Design)
- Exploration of Service Robot Morphology Through Generative AI Applications(Y. Ghim, 2024, AHFE International)
- Morphology Evolution for Embodied Robot Design With a Classifier-Guided Diffusion Model(Shulei Liu, Junchi Yan, Handing Wang, Yaochu Jin, 2026, IEEE Transactions on Evolutionary Computation)
机器人通用操作学习与任务基准
这些文献讨论了机器人学习中的通用操作技能、任务自动化、基准测试(Benchmark)以及感知到执行的闭环控制框架,关注点在于提升机器人解决复杂长程任务的泛化能力。
- Bab_Sak Robotic Intubation System (BRIS): A Learning-Enabled Control Framework for Safe Fiberoptic Endotracheal Intubation(Saksham Gupta, Sarthak Mishra, A. Ayub, K. Farooque, Spandan Roy, B. Gupta, 2025, arXiv.org)
- FMB: A functional manipulation benchmark for generalizable robotic learning(Jianlan Luo, Charles Xu, Fangchen Liu, Liam Tan, Zipeng Lin, Jeffrey Wu, Pieter Abbeel, Sergey Levine, 2024, The International Journal of Robotics Research)
- A Survey on Deep Generative Models for Robot Learning From Multimodal Demonstrations(Julen Urain, A. Mandlekar, Yilun Du, Nur Muhammad “Mahi” Shafiullah, Danfei Xu, Katerina Fragkiadaki, G. Chalvatzaki, Jan Peters, 2026, IEEE Transactions on Robotics)
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning(Ran Gong, Xiaohan Zhang, Jinghuan Shang, M. Minniti, J. Patel, Valerio Pepe, Riedana Yan, A. Gundogdu, Ivan Kapelyukh, A. Abbas, Xiaoqiang Yan, Harsh Patel, Laura Herlant, Karl Schmeckpeper, 2025, arXiv.org)
- Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation(Liang Heng, Jiadong Xu, Yiwen Wang, Xiaoqi Li, Mu Cai, Yan Shen, Juan Zhu, Guanghui Ren, Hao Dong, 2025, arXiv.org)
本次梳理的文献主要涵盖了灵巧操作与功能性抓取、机器人结构的生成式设计、以及机器人任务操作的通用学习三大领域。研究趋势表现为从单一抓取任务向多模态、语义驱动、长程操作技能迁移的深入探索,同时强调了生成式AI模型在形态结构优化与动作合成中的核心作用。
总计23篇相关文献
As a problem-solving activity, engineering design is usually iterative involving multiple proposed solutions that are tested against a predefined set of constraints. Human designers usually rely on their knowledge, experience, and intuition, which is a drawback when dealing with certain unknown problems. This is easily overcome by an AI that can generate and test several thousand alternative solutions to a design problem iteratively in the form of a parametric computational model. This paper seeks to present one such automated design process involving the development and testing of a low-maintenance robotic gripper featuring underactuation and reduced weight for missions in extreme environments. This is achieved by considering the computer as a collaborative partner in the design process, where the cloud computing engines generate thousands of mechanically improved designs in response to our rigorous and robust input computational model. Generated solutions include uniquely synthesised structures designed to achieve the aforementioned objectives. Notable contributions of this paper are presented through a comparative study confirming the gripper’s improved component accessibility, structural resilience, and doubled weight-to-power ratio achieved through 73% crude weight reduction compared to its predecessor.
… of research in dexterous hand mechanics, gesture semantics, and user experience evaluation, … Furthermore, we propose a closed-loop framework—“perception–cognition–generation–…
Intelligent robotic manipulation is a challenging study of machine intelligence. Although many dexterous robotic hands have been designed to assist or replace human hands in executing various tasks, how to teach them to perform dexterous operations like human hands is still a challenge. This motivates us to conduct an in-depth analysis of human behavior in manipulating objects and propose an object-hand manipulation representation. This representation provides an intuitive and clear semantic indication of how the dexterous hand should touch and manipulate an object based on the object's own functional areas. At the same time, we propose a functional grasp synthesis framework, which does not require real grasp label supervision, but relies on the guidance of our object-hand manipulation representation. In addition, in order to obtain better functional grasp synthesis results, we propose a network pre-training method that can make full use of easily obtained stable grasp data, and a network training strategy to coordinate the loss functions. We conduct object manipulation experiments on a real robot platform, and evaluate the performance and generalization of our object-hand manipulation representation and grasp synthesis framework.
In recent years, many dexterous robotic hands have been designed to assist or replace human hands in executing various tasks. But how to teach them to perform dexterous operations like human hands is still a challenging task. In this paper, we propose a grasp synthesis framework to make robots grasp and manipulate objects like human beings. We first build a dataset by accurately segmenting the functional areas of the object and annotating semantic touch code for each functional area to guide the dexterous hand to complete the functional grasp and post-grasp manipulation. This dataset contains 18 categories of 129 objects selected from four datasets, and 15 people participated in data annotation. Then we carefully design four loss functions to constrain the model, which successfully generates the functional grasp of dexterous hand under the guidance of semantic touch code. The thorough experiments in synthetic data show our model can robustly generate functional grasp, even for objects that the model has not see before.
We introduce DexDiffuser, a novel dexterous grasping method that generates, evaluates, and refines grasps on partial object point clouds. DexDiffuser includes the conditional diffusion-based grasp sampler DexSampler and the dexterous grasp evaluator DexEvaluator. DexSampler generates high-quality grasps conditioned on object point clouds by iterative denoising of randomly sampled grasps. We also introduce two grasp refinement strategies: Evaluator-Guided Diffusion and Evaluator-based Sampling Refinement. The experiment results demonstrate that DexDiffuser consistently outperforms the state-of-the-art multi-finger grasp generation method FFHNet with an, on average, 9.12% and 19.44% higher grasp success rate in simulation and real robot experiments, respectively.
… However, these approaches often lack the specific semantic context or corresponding … We evaluate the grasp success rate in Issac Gym simulation environment. To simulate the …
We introduce the sequential multi-object robotic grasp sampling algorithm SeqGrasp that can robustly synthesize stable grasps on diverse objects using the robotic hand’s partial Degrees of Freedom (DoF). We use SeqGrasp to construct the large-scale Allegro Hand sequential grasping dataset SeqDataset and use it for training the diffusion-based sequential grasp generator SeqDiffuser. We experimentally evaluate SeqGrasp and SeqDiffuser against the state-of-the-art non-sequential multi-object grasp generation method MultiGrasp in simulation and on a real robot. The experimental results demonstrate that SeqGrasp and SeqDiffuser reach an 8.71%–43.33% higher grasp success rate than MultiGrasp. Furthermore, SeqDiffuser is approximately 1000 times faster at generating grasps than SeqGrasp and MultiGrasp.
Learning from demonstrations, the field that proposes to learn robot behavior models from data, is gaining popularity with the emergence of deep generative models. Although the problem has been studied for years under names, such as imitation learning, behavioral cloning, and inverse reinforcement learning, classical methods have relied on models that do not capture complex data distributions well or do not scale well to large numbers of demonstrations. In recent years, the robot learning community has shown increasing interest in using deep generative models to capture the complexity of large datasets. In this survey, we aim to provide a unified and comprehensive review of the last year’s progress in the use of deep generative models in robotics. We present the different types of models that the community has explored, such as energy-based models, diffusion models, action value maps, and generative adversarial networks. We also present the different types of applications in which deep generative models have been used, from grasp generation to trajectory generation or cost learning. One of the most important elements of generative models is the generalization out of distributions. In our survey, we review the different decisions the community has made to improve the generalization of the learned models. Finally, we highlight the research challenges and propose a number of future directions for learning deep generative models in robotics.
Astract-Recent advances in foundation models are significantly expanding the capabilities of AI models. As part of this progress, this paper introduces a robot design framework that uses a diffusion model approach for generating 3D mesh structures. Specifically, we focus on generating directly fabri-cable robot structures that require no post-processing guided by human-imposed design constraints. Our approach can find the optimal design of the robot by optimizing or composing embedding vectors of the model. The efficacy of the framework is validated through an application to design, fabricate, and evaluate a jumping robot. Our solution is an optimized jumping robot with a 41% increase in jump height compared to the state-of-the-art design. Additionally, when the robot is augmented with an optimized foot, it can land reliably with a success ratio of 88% in contrast to the 4% success ratio of the base robot.
Recent research in the field of design engineering is primarily focusing on using AI technologies such as Large Language Models (LLMs) to assist early-stage design. The engineer or designer can use LLMs to explore, validate and compare thousands of generated conceptual stimuli and make final choices. This was seen as a significant stride in advancing the status of the generative approach in computer-aided design. However, it is often difficult to instruct LLMs to obtain novel conceptual solutions and requirement-compliant in real design tasks, due to the lack of transparency and insufficient controllability of LLMs. This study presents an approach to leverage LLMs to infer Function-Behavior-Structure (FBS) ontology for high-quality design concepts. Prompting design based on the FBS model decomposes the design task into three sub-tasks including functional, behavioral, and structural reasoning. In each sub-task, prompting templates and specification signifiers are specified to guide the LLMs to generate concepts. User can determine the selected concepts by judging and evaluating the generated function-structure pairs. A comparative experiment has been conducted to evaluate the concept generation approach. According to the concept evaluation results, our approach achieves the highest scores in concept evaluation, and the generated concepts are more novel, useful, functional, and low-cost compared to the baseline.
Automatic design of intelligent robots plays a central role and has been a trending topic in embodied intelligence. A promising approach is the co-design framework, wherein evolutionary algorithms (EAs) are employed to optimize the robot’s morphology while reinforcement learning algorithms are utilized to refine its control strategies. However, the discrete morphology design space introduces significant challenges for EAs, to efficiently identify optimal morphologies. In this work, we propose a novel morphology optimization method driven by a classifier-guided diffusion model to enhance the search efficiency of EAs. Prior to the design process, a universal diffusion model is trained using a set of randomly sampled feasible morphologies, prompting that the generated structures satisfy physical constraints. In each iteration of the EA for morphology optimization, a classifier is trained using previously evaluated morphologies and is then used to condition the diffusion model to generate quality solutions. Subsequently, the generated morphology is further refined based on the voxel distribution to incorporate features from the current high-quality morphology. Extensive experiments on a large-scale benchmark for co-designing the morphology and control of voxel-based soft robots demonstrate that our method significantly improves search efficiency in the morphological design space, outperforming both traditional EAs and surrogate-assisted EAs.
Advancements in robotics and artificial intelligence (AI) are bringing service robots into various aspects of our lives, and the appearance of service robots has diversified along with their increase. While anthropomorphic design has been extensively discussed in human-robot interaction (HRI) as a way of making robots more understandable and acceptable, much remains to be investigated about the desired level of human-likeness in a robot’s design, which is also dependent on the specific context of use. This paper proposes a visual mapping method as a means of guiding the appearance and degree of human-likeness of a service robot in the corresponding use context. A service robot context map, comprising the robot’s task nature and operation environment, is constructed and translated into a morphology map regarding the level of human-likeness and aesthetic qualities. Based on these mappings, two service robot contexts were selected to create evaluation materials to measure the desired degree of human-likeness. Variations of a service robot design were created and visualized in photorealistic rendering through the utilization of generative image AI tools. Though with some unintended design changes, generative image AI is an efficient way of creating robot representations in context for an evaluation study.
Existing works on prosthetic hands focus on increasing dexterity by carrying out functional tasks. Achieving specific hand movements, such as pointing the index finger, are desired but research on generating the hand movement itself has yet to be widely explored. In this work, we propose a pipeline for generating hand motion from body motion via using the Common-Rig, a kinematic rig representation for effective motion representation, and a diffusion-based inpainting method, which has shown strengths in generalization and stability. Common rigging is applied to a motion capture dataset with both body and hands information, and hand motions are generated while conditioned on the body motions of a hand- zeroed test set. The generated results of our proposed method, compared to two baseline methods, attain smaller fingertip positional errors and diversity closer to that of the ground truth. In addition, the generated motions are implemented on a real robotic system with prosthetic hands for evaluation.
Human hands play a central role in interacting, motivating increasing research in dexterous robotic manipulation. Data-Driven embodied AI algorithms demand precise, large-scale, human-like manipulation sequences, which are challenging to obtain with conventional reinforcement learning or real-world teleoperation. To address this, we introduce ManipTrans, a novel two-stage method for efficiently transferring human bimanual skills to dexterous robotic hands in simulation. ManipTrans first pre-trains a generalist trajectory imitator to mimic hand motion, then fine-tunes a specific residual module under interaction constraints, enabling efficient learning and accurate execution of complex bimanual tasks. Experiments show that ManipTrans surpasses state-of-the-art methods in success rate, fidelity, and efficiency. Leveraging ManipTrans, we transfer multiple hand-object datasets to robotic hands, creating DexManipNet, a large-scale dataset featuring previously unexplored tasks like pen capping and bottle unscrewing. DexManipNet comprises 3.3K episodes of robotic manipulation and is easily extensible, facilitating further policy training for dexterous hands and enabling real-world deployments.
Multimodal data is indispensable for advancing imitation learning, particularly in the context of dexterous hands. However, existing datasets predominantly rely on single-modality inputs, such as RGB images, which inherently lack the capacity to capture the spatial and temporal dynamics essential for achieving human-like dexterity. To address this limitation, we introduce Multi-Modal Dex, a dataset that integrates multimodal sensory data to enable the effective learning of dexterous skills from human demonstrations. By combining visual, point cloud, and kinematic modalities, our dataset provides a richer representation of hand interactions, thereby facilitating a more nuanced understanding of dexterous imitation. Our framework leverages neural rendering and kinematic optimization to align human and robotic hand poses in a shared canonical space, enabling geometrically consistent skill transfer. Furthermore, we analyze the dataset’s potential to advance dexterous robots in perception, imitation learning, and real-world dexterous skill transfer. The data is available at https://github.com/WangShaoSUN/MutliDex.
Robot grasp dataset is the basis of designing the robot's grasp generation model. Compared with the building grasp dataset for Low-DOF grippers, it is harder for High-DOF dexterous robot hand. Most current datasets meet the needs of generating stable grasps, but they are not suitable for dexterous hands to complete human-like functional grasp, such as grasp the handle of a cup or pressing the button of a flashlight, so as to enable robots to complete subsequent functional manipulation action autonomously, and there is no dataset with functional grasp pose annotations at present. This paper develops a unique Cost-Effective Real-Simulation Annotation System by leveraging natural hand's actions. The system is able to capture a functional grasp of a dexterous hand in a simulated environment assisted by human demonstration in real world. By using this system, dexterous grasp data can be collected efficiently as well as cost-effective. Finally, we construct the first dexterous functional grasp dataset with rich pose annotations. A Functional Grasp Synthesis Model is also provided to validate the effectiveness of the proposed system and dataset. Our project page is: https://hjlllll.github.io/DFG/.
Human-Robot Interaction (HRI) is increasingly becoming commonplace in various fields, including service robotics, industrial automation, and healthcare. Public acceptance and positive interaction with these robots are heavily influenced by the appearance and human likeness of these robots. To enable interaction with human beings, robots need manipulators that support high dexterity, cost-efficiency, and ease of development. This study presents the design and development of EvoGrip, a humanoid dexterous robotic hand aimed at supporting research and development in HRI. EvoGrip is a cost-effective and easy-to-build robotic hand, adapted from the open-source project Inmoov, enhanced with actuators, sensors and other added value features. The mechanical design changes to the hand include improvements for movement repeatability and the addition of a degree of freedom in the thumb to enable greater dexterity. The development and integration of a custom string potentiometer system, enables precise finger position tracking while simultaneously reducing design complexity. Furthermore, the study develops two distinct modeling approaches for EvoGrip's finger dynamics: a mathematical model based on first principles and a data-driven model using system identification techniques. Both modeling strategies demonstrated high accuracy, with the system identification model showing superior performance in compensating for complex, nonlinear behaviors. This work establishes a foundation for future research and advancements in robotic hands, focusing on real-time control, advanced pressure excursion and applications in human-centric tasks.
In this paper, we propose a real-world benchmark for studying robotic learning in the context of functional manipulation: a robot needs to accomplish complex long-horizon behaviors by composing individual manipulation skills in functionally relevant ways. The core design principles of our Functional Manipulation Benchmark (FMB) emphasize a harmonious balance between complexity and accessibility. Tasks are deliberately scoped to be narrow, ensuring that models and datasets of manageable scale can be utilized effectively to track progress. Simultaneously, they are diverse enough to pose a significant generalization challenge. Furthermore, the benchmark is designed to be easily replicable, encompassing all essential hardware and software components. To achieve this goal, FMB consists of a variety of 3D-printed objects designed for easy and accurate replication by other researchers. The objects are procedurally generated, providing a principled framework to study generalization in a controlled fashion. We focus on fundamental manipulation skills, including grasping, repositioning, and a range of assembly behaviors. The FMB can be used to evaluate methods for acquiring individual skills, as well as methods for effectively combining and ordering such skills in order to solve complex, multi-stage manipulation tasks. We also offer an imitation learning framework that includes a suite of policies trained to solve the proposed tasks. This enables researchers to utilize our tasks as a versatile toolkit for examining various parts of the pipeline. For example, researchers could propose a better design for a grasping controller and evaluate it in combination with our baseline reorientation and assembly policies as part of a pipeline for solving multi-stage tasks. Our dataset, object CAD files, code, and evaluation videos can be found on our project website: https://functional-manipulation-benchmark.github.io.
Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing, and inter-hand coordination, and existing resources lack high-fidelity bimanual sequences that capture nuanced finger dynamics and collaboration. To fill this gap, we present HandX, a unified foundation spanning data, annotation, and evaluation. We consolidate and filter existing datasets for quality, and collect a new motion-capture dataset targeting underrepresented bimanual interactions with detailed finger dynamics. For scalable annotation, we introduce a decoupled strategy that extracts representative motion features, e.g., contact events and finger flexion, and then leverages reasoning from large language models to produce fine-grained, semantically rich descriptions aligned with these features. Building on the resulting data and annotations, we benchmark diffusion and autoregressive models with versatile conditioning modes. Experiments demonstrate high-quality dexterous motion generation, supported by our newly proposed hand-focused metrics. We further observe clear scaling trends: larger models trained on larger, higher-quality datasets produce more semantically coherent bimanual motion. Our dataset is released to support future research.
Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to capture complex geometric constraints or generate goal-state observations to capture semantic and geometric knowledge, but fail to explicitly couple object transformation with action prediction, resulting in errors due to generative noise. To address these limitations, we propose Imagine2Act, a 3D imitation-learning framework that incorporates semantic and geometric constraints of objects into policy learning to tackle high-precision manipulation tasks. We first generate imagined goal images conditioned on language instructions and reconstruct corresponding 3D point clouds to provide robust semantic and geometric priors. These imagined goal point clouds serve as additional inputs to the policy model, while an object-action consistency strategy with soft pose supervision explicitly aligns predicted end-effector motion with generated object transformation. This design enables Imagine2Act to reason about semantic and geometric relationships between objects and predict accurate actions across diverse tasks. Experiments in both simulation and the real world demonstrate that Imagine2Act outperforms previous state-of-the-art policies. More visualizations can be found at https://sites.google.com/view/imagine2act.
Robotic manipulation in unstructured environments requires end-effectors that combine high kinematic dexterity with physical compliance. While traditional rigid hands rely on complex external sensors for safe interaction, electrohydraulic actuators offer a promising alternative. This paper presents the design, control, and evaluation of a novel musculoskeletal robotic hand architecture powered entirely by remote Peano-HASEL actuators, specifically optimized for safe manipulation. By relocating the actuators to the forearm, we functionally isolate the grasping interface from electrical hazards while maintaining a slim, human-like profile. To address the inherently limited linear contraction of these soft actuators, we integrate a 1:2 pulley routing mechanism that mechanically amplifies tendon displacement. The resulting system prioritizes compliant interaction over high payload capacity, leveraging the intrinsic force-limiting characteristics of the actuators to provide a high level of inherent safety. Furthermore, this physical safety is augmented by the self-sensing nature of the HASEL actuators. By simply monitoring the operating current, we achieve real-time grasp detection and closed-loop contact-aware control without relying on external force transducers or encoders. Experimental results validate the system's dexterity and inherent safety, demonstrating the successful execution of various grasp taxonomies and the non-destructive grasping of highly fragile objects, such as a paper balloon. These findings highlight a significant step toward simplified, inherently compliant soft robotic manipulation.
Endotracheal intubation is a critical yet technically demanding procedure, with failure or improper tube placement leading to severe complications. Existing robotic and teleoperated intubation systems primarily focus on airway navigation and do not provide integrated control of endotracheal tube advancement or objective verification of tube depth relative to the carina. This paper presents the Robotic Intubation System (BRIS), a compact, human-in-the-loop platform designed to assist fiberoptic-guided intubation while enabling real-time, objective depth awareness. BRIS integrates a four-way steerable fiberoptic bronchoscope, an independent endotracheal tube advancement mechanism, and a camera-augmented mouthpiece compatible with standard clinical workflows. A learning-enabled closed-loop control framework leverages real-time shape sensing to map joystick inputs to distal bronchoscope tip motion in Cartesian space, providing stable and intuitive teleoperation under tendon nonlinearities and airway contact. Monocular endoscopic depth estimation is used to classify airway regions and provide interpretable, anatomy-aware guidance for safe tube positioning relative to the carina. The system is validated on high-fidelity airway mannequins under standard and difficult airway configurations, demonstrating reliable navigation and controlled tube placement. These results highlight BRIS as a step toward safer, more consistent, and clinically compatible robotic airway management.
Generalist robot learning remains constrained by data: large-scale, diverse, and high-quality interaction data are expensive to collect in the real world. While simulation has become a promising way for scaling up data collection, the related tasks, including simulation task design, task-aware scene generation, expert demonstration synthesis, and sim-to-real transfer, still demand substantial human effort. We present AnyTask, an automated framework that pairs massively parallel GPU simulation with foundation models to design diverse manipulation tasks and synthesize robot data. We introduce three AnyTask agents for generating expert demonstrations aiming to solve as many tasks as possible: 1) ViPR, a novel task and motion planning agent with VLM-in-the-loop Parallel Refinement; 2) ViPR-Eureka, a reinforcement learning agent with generated dense rewards and LLM-guided contact sampling; 3) ViPR-RL, a hybrid planning and learning approach that jointly produces high-quality demonstrations with only sparse rewards. We train behavior cloning policies on generated data, validate them in simulation, and deploy them directly on real robot hardware. The policies generalize to novel object poses, achieving 44% average success across a suite of real-world pick-and-place, drawer opening, contact-rich pushing, and long-horizon manipulation tasks. Our project website is at https://anytask.rai-inst.com .
本次梳理的文献主要涵盖了灵巧操作与功能性抓取、机器人结构的生成式设计、以及机器人任务操作的通用学习三大领域。研究趋势表现为从单一抓取任务向多模态、语义驱动、长程操作技能迁移的深入探索,同时强调了生成式AI模型在形态结构优化与动作合成中的核心作用。