Assistance games formalize human-robot collaboration under asymmetric information, where the human knows the goal but the robot must infer it from observations and interactions. However, computing optimal strategies online is generally intractable because it requires planning in a POMDP.
The authors identify a class of assistance games where a pragmatic-pedagogic best response can be computed efficiently. This strategy balances the robot's pragmatic reasoning about human behavior with pedagogic signaling to elicit useful feedback.
In a single round of interaction, the robot can achieve corrigible assistance, meaning it can adjust its behavior based on human corrections without requiring full POMDP solving. This provides a practical approximation to optimal assistance.
The approach is validated through theoretical analysis and simulations, demonstrating that it outperforms baseline methods in terms of task success and correction efficiency.
This work opens the door to deploying assistance game frameworks in real-time robotic systems, enhancing adaptability and safety in collaborative tasks.
Source: arXiv (2607.27508).