机器人适应受数据收集瓶颈的困扰

2 分•作者: aunmesh•2 个月前
如何缓解机器人适应性问题,即可靠部署机器人策略的“最后一公里”? 视觉-语言-动作策略(如 Gr00T、Pi0、OpenVLA)的设计需要一个适应阶段,才能在特定的应用环境和机器人任务组合上可靠部署(可靠意味着极高的成功率)策略模型。 这个适应阶段,通常被称为微调或后训练,针对的是在特定实体(如单臂机器人、双臂机器人或人形机器人等)上,在精确的部署环境中完成的一系列狭窄任务。 最近发布的“世界-动作模型”(WAMs),虽然在泛化到新颖任务和环境方面有所改进,但也需要适应阶段才能在给定环境中可靠地执行给定任务。 机器人适应性的关键要求是高质量的适应数据,通常是人类遥操作机器人执行给定环境中所需任务的数据。通常需要 50-200 个示例的人类遥操作数据,具体取决于被微调的基础策略、期望的性能水平以及任务的复杂性。 因此,对于我们希望机器人可靠执行的任何环境或任务的变更,都需要一次又一次地重新收集此类遥操作数据。 因此,机器人适应性,即可靠部署机器人策略的关键“最后一公里”,受制于数据收集的瓶颈——这个问题需要得到缓解。 因此,一个需要解决方案的重要研究问题是: 有哪些有效的方法可以缓解机器人策略适应性所需的人类遥操作数据收集工作?
查看原文
How can we ease Robotic Adaptation, or the critical last mile of reliable robotic policy deployment?<p>Vision-Language-Action Policies (such as Gr00T, Pi0, OpenVLA) are designed so that an adaptation stage is needed before reliable deployment (reliable meaning with very high success rate) of the policy model on a specific combination of application environment and robotic task.<p>This adaptation stage, often called fine-tuning or post-training, caters to the narrow set of tasks, done in the exact deployment environment, atop a specific embodiment (such as single arm robots, or bi manual robots, or humanoids among others).<p>The recently released World-Action-Models (WAMs), while having improved generalization to novel tasks and environments, also need adaptation stage to reliably execute a given task on a given environment.<p>The key requirement of robotic adaptation is high quality adaptation data, typically human teleoperation data of the robot performing the required task in the given environment. Typically 50-200 examples of human teleoperation are required, depending on the base policy being finetuned, as well as the desired performance levels and the complexity of the task.<p>Thus, given any change in the environment or task where we want robot to perform reliably, there will be a need to collect such teleoperation data again and again.<p>Thus, robotic adaptation, which is the critical last mile of reliable robotic policy deployment - suffers from the data collection bottleneck - which needs to be eased.<p>Thus, an important research question needing solutions is -<p>What are effective approaches to ease human-teleoperation data collection effort required for robotic policy adaptation?