Jenga-playing robot gains skills with faster machine-learning approach
Amy J. Born | February 06, 2019MIT engineers have developed a robot that can learn to complete tasks through both touch and vision. The robot developed its skills by playing the block-stacking game Jenga.
The Jenga-playing robot demonstrates something that’s been tricky to attain in previous systems: the ability to quickly learn the best way to carry out a task, not just from visual cues, as it is commonly studied today, but also from tactile, physical interactions. Source: MIT / CC BY-NC-ND 3.0The robot's soft-pronged gripper, a force-sensing wrist cuff and external camera allow it to see and feel the tower. Each time the robot attempts to remove a block, it compares the visual and tactile feedback with the visual and force measurements of its previous attempts as well as the outcomes (successful or unsuccessful). The robot "learns" which course of action is most likely to keep the tower standing and decides whether to continue extracting the piece or to try a different piece.
A robot requires interactive perception and manipulation to master both the physical and cognitive skills necessary to successfully play the game, said Alberto Rodriguez, the Walter Henry Gale Career Development Assistant Professor in the Department of Mechanical Engineering at MIT. The researchers determined that in order to do this, the robot needed real-world experience with the actual game.
This machine-learning approach is unique because it eliminates the need for the robot to experience every possibility of interacting with the tower and an individual block. Instead of gathering data from potentially tens of thousands of tries, the researchers were able to accomplish their goal with only 300 attempts.
"The key challenge is to learn from a relatively small number of experiments by exploiting common sense about objects and physics,” said Rodriguez. The team's more efficient approach draws from human cognition and the ways people learn to play the game. Similar results are grouped in clusters to represent types of block behavior, such as whether the block was hard or easy to move, or if the tower fell when the block was moved. These data clusters allow the robot to develop a model using the visual and tactile measurements of each block it attempts to move to predict the block's behavior.
Industries that use robots for a variety of applications, such as assembling, gluing, sealing and material handling, could benefit from the tactile learning system developed by Rodriguez and his colleagues. “There are many tasks that we do with our hands where the feeling of doing it ‘the right way’ comes in the language of forces and tactile cues,” he said. “In a cellphone assembly line, in almost every single step, the feeling of a snap-fit, or a threaded screw, is coming from force and touch rather than vision.”
[Learn more about industrial robots on Engineering360.]
The research is published in the journal Science Robotics.