Most people associate artificial intelligence directly with robots-as an inseparable pair. In fact, the term ‘artificial intelligence’ is rarely used in research laboratories as the terminologies for specific technologies are more relevant. When I receive the question “Is this robot operated by AI?” during external lectures, I remember myself hesitating to answer every time, thinking whether it would be appropriate to call the algorithms we develop artificial intelligence. First used by scientists such as John McCarthy and Marvin Minsky in the 1950s, and frequently appearing in sci-fi novels or films for decades, AI is now being used in smartphone virtual assistants and autonomous vehicle algorithms. As far as its history goes, AI can mean many different things, and therefore, cause confusion.
However, there seems to be a preconception in people’s minds regarding AI. It is the belief that AI is an artificially realized version of ‘human intelligence.’ Even though there is no reason for artificial intelligence to resemble that of humans, this preconception might come from our cognitive bias as human beings.
Then, where does this bias come from? I believe it comes from our psychological tendency to subconsciously anthropomorphize the subjects we see. Humans have evolved as social animals, probably developing the ability to understand and empathize with each other in the process. Our tendency to anthropomorphize subjects would have come from the same evolutionary process. The same goes even for researchers in the engineering field. When engineers update robots with a new program, they tend to use the expression ‘teaching’ robots. Although there are machine learning techniques similar to teaching humans, it is not identical to actually teaching humans. Nevertheless, we are used to using anthropomorphized expressions.
"There is a universal tendency among mankind to conceive all beings like themselves.
We find human faces in the moon, armies in the clouds..."
- David Hume
Of course, we not only anthropomorphize subjects’ appearance but also their state of mind. For example, when Boston Dynamics released a video of its engineers kicking a robot, many viewers reacted by saying ‘this is cruel,’ and that they ‘pity the robot.’ A comment saying, “one day, robots will take revenge on that engineer” received ‘likes.’ In reality, the engineer was simply testing the robot’s balancing algorithm. However, before any thought process to comprehend this situation, the kicking combined with the struggling of the animal-like robot is instantaneously transmitted to our brains, leaving a strong impression. Like this, such instantaneous anthropomorphism has deep effect in our cognitive process.
Problems may occur if policy makers or government research program managers fail to recognize such bias, as this may cause unexpected errors in understanding a new technology or setting research directions. This is what I’ve discovered on such problems, through various media channels and conversations.
We tend to judge difficulty of robots or AI’s tasks in comparison to humans
How did you feel when AlphaGo, AI developed by DeepMind, defeated 9-dan Go player Lee Sedol? You may have been surprised or terrified, thinking that AI has surpassed the ability of geniuses. Still, winning a game with an exponential number of possible moves like Go only means that AI has exceeded a very limited part of human intelligence. The same goes for IBM’s AI, Watson, which competed in ‘Jeopardy!’, the television quiz show. Though memory is an important ability for humans, we do not consider Wikipedia, containing much more knowledge, an intelligent being. I believe many were impressed to see the Mini Cheetah, developed in my MIT Biomimetic Robotics Laboratory, perform a backflip. While jumping backwards and landing on the ground is very dynamic and eye-catching, the algorithm for the particular motion is incredibly simple compared to one that enables walking.
Likewise, robots performing tasks that are difficult for humans seems quite impressive. On the contrary, people aren’t as interested when robots perform tasks easily done by humans, such as walking, tying shoelaces, or putting on clothes. However, realizing these tasks with algorithms is either extremely difficult and complicated, or still mostly impossible. This gap occurs because we tend to think of a task’s difficulty based on human standards.
While humans can learn to do 5 different tasks from one thing,
robots need to be taught 10 different things to do a single task.
We tend to generalize AI functionality after watching a single robot demonstration. When we see someone on the street doing backflips, we tend to assume this person would be good at walking and running, and also be flexible and athletic enough to be good at other sports. Of course, such judgement on this person would not be very wrong. However, can we also apply this judgement on robots? It’s easy for us to generalize and determine AI performance based on an observation of a specific robot motion or function, just as we do with humans. By watching a video of a robot hand solving Rubik’s Cube at OpenAI, an AI research lab, we think that the AI can perform all other simpler tasks because it can perform such a complex one. We overlook the fact that this AI’s neural network was only trained for a limited type of task; solving the Rubik’s Cube in that configuration. If the situation changes—for example, holding the cube upside down while manipulating it-the algorithm does not work as well as you expected. Unlike AI, humans can combine individual skills and apply them to multiple complicated tasks. Once we learn how to solve a Rubik’s Cube, we can quickly work on the cube even when we’re told to hold it upside down, though it may feel strange at first. Human intelligence can naturally combine the objectives of not dropping the cube and solving the cube. Most robot algorithms will require new data or reprogramming to do so. A person who can spread jam on bread with a spoon can do the same using a fork. It is obvious. We understand the concept of spreading jam, and can quickly get used to using a completely different tool. Also, while autonomous vehicles require actual data for each situation, human drivers can make rational decisions based on pre-learned concepts to respond to countless situations. These examples show one characteristic of human intelligence in stark contrast to robot algorithms, which cannot perform tasks with insufficient data.
Mammals have continuously been evolving for more than 65 million years. The entire time humans spent on learning math, using languages, and playing games would sum up to a mere 10,000 years. In other words, humanity spent a tremendous amount of time developing abilities directly related to survival, such as walking, running, and using our hands. This would have been sufficient time for us to freely combine different abilities and apply them to complex tasks. Therefore, it may not be surprising that computers can compute much faster than humans, as they were developed for this purpose in the first place. Likewise, it is natural that computers cannot easily obtain the ability to freely use hands and feet for various purposes as humans do. These skills have been attained through evolution of over 10 million years.
This is why it is unreasonable to compare robot or AI performance from demonstrations to that of an animal or human’s abilities. It would be rash to believe that robot technologies on walking and running like animals are complete, while watching videos of the cheetah robot run across fields at MIT and leaping over obstacles. Numerous robot demonstrations still rely on algorithms set for specialized tasks in specific situations. There indeed is a tendency for researchers to select demonstrations that seem difficult, as it can give a very challenging and strong impression. However, this level of difficulty is from the human perspective, which may be irrelevant to the actual algorithm performance. Humans are easily influenced by instantaneous and reflective perception before any logical thought. And this cognitive bias is strengthened when the subject is very complicated and difficult to theoretically analyze, for example, machine learning.
There is a high probability that tasks humans can easily perform and are often seen are highly advanced for AI or robots. It would be wrong to think animalistic abilities that are easy for us, are also easy for computers. Perhaps it may be easier for us to find a new frontier of robot research in skills that seem too easy for ourselves and are repetitive.
Humans process information qualitatively, and computers, quantitively
Looking around, our daily lives are filled with algorithms, as can be seen by machines and services that run on these algorithms. Most algorithms operate on numbers. We use the term objective function, which literally means a numerical function that describes a certain objective. Many algorithms have the sole purpose of reaching the maximum or minimum value of this function, and an algorithm’s characteristics differ based on how it achieves this.
The goal of tasks such as winning a game of Go or chess are relatively easy to quantify. This may come as a surprise, but this is what differentiates human intelligence from computer algorithms. The easier quantification is, the better the algorithms work. On the contrary, humans often make decisions without quantitative thinking. Of course, as the digital world has become popularized, we meet more people accustomed to quantitative thinking. Still, even those people make dozens of decisions in a day based on qualitative thinking.
How about cleaning a room as an example? The organization style differs considerably from when one cleans his/her own room to when a different person does so. Even one’s own way of cleaning differs subtly from day to day, depending on the situation or how one feels. Were we trying to maximize a certain function in this process? We did no such thing. The act of cleaning has been done with an abstract objective of “clean enough.” Besides, the standard for how much is “enough” changes easily. This standard may be different among people, causing conflicts particularly among family members or roommates.
There are many other examples. When you wash your face every day, which quantitative indicators do you intend to maximize with your hand movements? How hard do you rub? Which quantitative indicators are considered when making friends? When choosing what to wear? When choosing what to have for dinner? When choosing which dish to wash first? The list goes on. We are used to making decisions that are good enough by putting together information we already have. However, we often do not check whether every single decision is optimized. Most of the time, it is impossible to know because we would have to satisfy numerous contradicting indicators. When selecting groceries with a friend at the store, we cannot each quantify standards for groceries and make a decision based on these numerical values. Usually, when one picks something out, the other will either say “OK!” or suggest another option. This is very different from saying this vegetable “is the optimal choice!” It is more like saying “this is good enough”
Many people believe that the evolution process of animals, including humans, is mathematical optimization. For example, it’s easy to think that a cheetah is optimized to run fast. However, in the evolutionary process where survival itself is the objective, running fast is just one of the many essential survival functions. In order to survive, a cheetah needs to evolve to have the appropriate complex and diverse functions that are almost impossible to quantify. I personally do not believe evolution is the process of optimizing a specific quantitative value. Animals that adapted well enough are the ones who survive. In other words, for survival, there is no need for one function to reach its optimal state. Optimization of a single function may debilitate other functions and lead to flawed evolution. If all animals had gone through optimization for millions of years, would it be possible for so many species to coexist in one place now? Only a few species that achieve the optimal state—if this is definable—would have survived. As this is not the case, it seems more plausible that animals that have adapted ‘well enough’ to survive through evolution are the ones remaining.
AI tries to find the optimal solution.
Humans on the other hand are skilled at quickly finding ‘good-enough’ methods.
The difference between people, who only need to be ‘good enough’, and computers that operate based on functions for optimization may be problematic when designing work or services we expect robots to perform. This is because while robots perform tasks based on quantitative values, humans’ satisfaction, the outcome of the task, cannot be quantified. It is not easy to quantify tasks that must adapt to individual preferences or changing circumstances like the aforementioned room cleaning or dishwashing tasks. That is, to coexist with humans, robots may have to evolve not to optimize particular functions, but to achieve “this is good enough.” Of course, the latter is much more difficult to achieve in a real-life situation where you need to manage so many conflicting objectives.
Actually, we do not know what we are doing
Try to recall the most recent meal you had before reading this. Can you remember what you had? Then, can you also remember the process of chewing and swallowing the food? Do you remember what exactly your tongue was doing at that very moment? Our tongue does so many things for us that in Asia, we even have an idiom “tongue in the mouth (如口之舌: doing things thoroughly and promptly without explicit instructions).” It helps us put food in our mouths, distribute between our teeth, swallow the finely chewed pieces, or even send large pieces back toward our teeth, if needed. We can naturally do all of this, even while talking to a friend. Oh! Come to think of it, the tongue is also in charge of pronouncing during conversations. How much do our conscious decisions contribute to the movement of our tongues that accomplish so many complex tasks simultaneously? It may seem like we are moving our tongues as we want, but in fact, there are more moments when the tongue is moving automatically, irrelevant to our consciousness. This is why we cannot remember detailed movements of our tongues during a meal. We know little about their movement in the first place.
We may assume that the hand is the most consciously controllable organ, but many hand movements also happen automatically and unconsciously, or subconsciously at most. For those who disagree, try putting something like keys in your pocket and take it back out. In that short moment, countless micromanipulations instantly and seamlessly occur in a continuum. We cannot perceive each action separately. We do not even know what units we should divide them into, so we collectively express them as organize, wash, apply, rub, wipe, etc. These verbs are qualitative definitions that refer to the aggregate of fine movements and manipulations. Of course, it is easy even for children to understand and think of this concept, but from the perspective of algorithm development, these words are endlessly vague and abstract.
Let’s try to teach how to make a sandwich by spreading peanut butter on bread. We can show how this is done and explain with a few simple words. Even a child could easily understand and learn how to do this in a few minutes. Let’s assume a slightly different situation. Say there is an alien who uses the same language as us, but knows nothing about human civilization or culture (this assumption is already contradictory, but bear with me). Can we explain how to make a peanut butter sandwich over the phone? We will probably get stuck trying to explain how to scoop peanut butter out of the jar. Even grasping the slice of bread is not so simple. We have to grasp the bread strongly enough so we can spread the peanut butter, but not so much so as to ruin the shape of the soft bread. At the same time, we should not drop the bread either. It is easy for us to think of how to grasp the bread, but it will not be easy to express this through speech or text. Even if it is a human learning a task, can we learn a carpenter’s work over the phone? Can we precisely correct tennis or golf postures over the phone? It is difficult to discern to what extent the details we see are done either consciously or unconsciously.
My point is that not everything we do with our hands and feet can directly be expressed with our language. Things that happen in between successive actions often automatically occur unconsciously, and thus we explain our actions in a much simpler way than how they actually take place. This is why our actions seem very simple, and why we forget how incredible they really are. The limitations of expression lead to underestimation of actual complexity.
We should recognize the fact that difficulty of language depiction can hinder research progress in fields where words are not well developed. Many engineers and machine learning experts have already noticed the importance of this technological research in the future robot market. They are trying various approaches to create new research methods. We’ve started the challenge of training AI and robots to do things that we do well, but do not understand how.
Tasks that need to be physically trained instead of being taught with words,
and intelligence on areas such as unconscious movement of the eyes and tongue,
are difficult to understand and cannot be completely expressed through language.
Until recently, AI has been practically applied in information services related to data processing. Examples include voice recognition, facial recognition, etc. Now, we are entering a new era of AI that can effectively perform physical service beside us. That is, the time is coming in which automation of complex physical tasks becomes imperative. Particularly, the upcoming aging society poses a huge challenge to us. Shortage of labor is no longer a vague social problem. It is urgent that we discuss how to develop technologies that augment humans’ capability, allowing us to focus on more valuable work and pursue lives uniquely human. This is why not only engineers but also members of society from various fields should improve their understanding of AI and unconscious cognitive biases. It is easy to misunderstand artificial intelligence as it is unlike human intelligence. Things that are very natural among humans may be cognitive biases for AI and robots. Without a clear understanding on this, we cannot set the appropriate directions for technology research, application, and policy. In order for all of us to be free from these biases, we need constant awareness and debate to promote the development and application of technology.
▶︎Professor Sangbae Kim, NAVER LABS technical consultant and the director of MIT Biomimetic Robotics Laboratory, is a world-renowned robotics expert with over 9,000 accumulated citations. He has been drawing worldwide attention with Stickybot, named one of 2006’s best inventions in Time magazine, and the MIT Cheetah robots, quadruped robots with an innovative mechanism.
▶︎ Hyung Taek Yoon, the creator of Forward Thinking series’ illustrations, is an illustrator and spatial storyteller. He has designed concepts and illustrations of various spatial projects, with works such as ‘Space Begins from Stories.’
▶︎Forward Thinking is a recurring online publication focusing on today’s major technological trends, such as AI, robotics, autonomous driving, and metaverses, containing stories of outstanding researchers collaborating with NAVER LABS. www.naverlabs.com/en/forwardthinking