The Limits of Control: Friendly AI and the AI Box Thought Experiment
- Kala Jagoda
- Jun 27
- 3 min read
Chapter 4 of James Barrat’s Our Final Invention raises the question of what would happen when humans create an intelligence greater than our own? Here, Barrat focuses on the concept of a “Friendly AI” and the AI Box Experiment. These challenge the assumption that building AGI and ASI machines is mainly a technical achievement rather than an ethical one.
The idea of “Friendly AI” comes from the attempt to ensure that an AI, regardless of its advancement, will continue to help and benefit humans. At a glance, the resolution may seem as simple as programming ethical rules into the AI systems. Barrat, however, emphasizes that human values are not consistent or easy to understand. As previously discussed, what one person might consider “good” may significantly differ from another. Additionally, depending on the circumstance, that same person may change their view. Barrat goes even further, acknowledging that even if a so-called “good organization” or “good company” created the AI system, it is not guaranteed that the system itself will be friendly.
An example of this that I find extremely prevalent is the popularity of recommendation systems such as YouTube or other platforms. Although the idea seems pretty harmless: the algorithm seeks to help users find videos they like through predictions, the use of algorithms such as these, however, have many negatives. Firstly, much like with many’s concerns about AI systems such as ChatGPT, the use of an algorithm to automatically generate content for the viewer removes the users’ tendency to search on their own. While, yes, not having to look for videos on your own is helpful, because users are less likely to research on their own, critical thinking is deteriorating, as well as problem-solving skills. Additionally, by promoting content based on what users have previously clicked, algorithms increase bias. For example, if a user has watched News Source 1 more times than News Source 2, the YouTube algorithm is more likely to promote News Source 1. This limits the diversity of information that the user is exposed to, and a feedback loop is created. The user continues to obtain information from the same viewpoints and perspectives, while other viewpoints, such as that from News Source 2, are more likely to be ignored.
This concern also leads directly into one of the most thought-provoking parts of the chapter: the AI Box Experiment. The experiment, proposed by AI researcher and cofounder of MIRI Eliezer Yudkowsky, explores whether an ASI placed into a sealed computer system with no outside access or influence would be able to be safely contained. Yudkowsky reported that in several instances, he was able to persuade human gatekeepers to release him from the box despite the clear rules of the game: don’t, under any circumstances, allow the AI to leave the box. The results from this experiment have scary implementations: if a human playing the role of an ASI could convince another human in a controlled scenario, what would real-world containment look like, and how effective would it actually be?
Together, these ideas demonstrate that intelligence is not controllable. People often say that the risk of advanced AI is its potential to become hostile, but Barrat argues that this isn’t the full picture. The real danger is that the AI only needs to pursue objectives that are not perfectly aligned with ours in order to become dangerous. Once this happens at a superhuman level of intelligence, even well-meaning instructions and well-fortified programs can lead to unintended and irreversible consequences.
Much like many others, I fear these outcomes, and believe that it's crucial that more learn about the negatives of AI, and what can be done to make a safer future. With modern society being so starstruck and impressed with the swiftness in which AI is developing, I think it's only a matter of time before these changes soon overwhelm human society and wreak havoc. Policies and restrictions need to be implemented now if this wants to be prevented.
Comments