“The computer can help us be better, but we must be careful not to pass our biases onto it”

With the word “gato” (cat) in different languages and images of felines, Jorge Pérez, Computer Science academic at Universidad de Chile and researcher at the Millennium Institute for Foundational Research on Data (IMFD), kicked off his presentation at the VI International Conference on Scientific Culture at Universidad Andrés Bello, an event open to the general public that seeks to bring scientific knowledge closer to society at large.

The different meanings and ways of depicting the concept of “gato” served as the hook to explain how computers understand human language and represent the meaning of a word through codes. Amid laughter and spontaneous participation, attendees came to understand the logic behind the association between words and the meaning conveyed by the computer through codes.

The researcher gave relatable examples of the applications of his research in the field of political data analysis. One such example was the 2016 constitutional process, for which he designed a Constitutional Explorer that analyzed the most recurring topics in citizen assemblies by municipality.

Galaxy of concepts

His most recent work along these lines is the construction of a “galaxy where each star is a phrase said by Michelle Bachelet during the previous government. I computed the representation of each one of them and placed them in three dimensions so they could be explored,” he said, while projecting a complex three-dimensional data cloud.

What is remarkable about this analysis is that it is not a summary, but rather encompasses all of the former president’s speeches, which is possible because all of her addresses have been digitized, Jorge Pérez explained.

The researcher used as an example Bachelet’s phrase “everyone knows where their shoe pinches”: the system was able to identify every time it was said during her second term in office. This galaxy is available online for anyone who wants to explore it, at the link bit.do/galaxia-presidencial

The scope of artificial intelligence

A question arose from the audience about the use of data in the context of artificial intelligence. “The way we train artificial intelligence today is with data. We have so much data that we can have a machine learn from it. But it will also learn the biases in the data. We scientists have to take responsibility for that: an artificial intelligence is not going to learn something I am not showing it,” the researcher stated, while clarifying that these biases occur across all languages, not just Spanish.

Jorge Pérez left an open challenge: that in this type of research -based on data or aimed at better understanding the user- “ethics should be present at all times. We scientists should reward work that advances ethics,” and he concluded by stating that “computers can understand human language up to a certain point. They can help us become better, but we must be careful not to pass on our own biases to them.”