Study Using CEP Survey Data Tests Surveys Answered by Artificial Intelligence: They Are Less Accurate in Chile Than in the U.S.
The study, published in EPJ Data Science, evaluated the "silicon sampling" technique, which replaces actual respondents with responses simulated using artificial intelligence. In Chile, the models showed lower accuracy among women, older adults, People , those with lower levels of education, and lower-income groups.
Large language models (LLMs), such as those behind ChatGPT, Gemini, and Claude, began being used a few years ago to simulate survey responses without consulting People . The technique, known as silicon sampling, involves asking an artificial intelligence (AI) model to answer a questionnaire as if it were a real person.
In 2023, a study published in the journal *Political Analysis* by researchers at Brigham Young University showed that this method was capable of reproducing the responses obtained in surveys conducted in the United States with considerable accuracy, opening up the possibility of using AI to study public opinion much more quickly and cost-effectively.

Is the same thing happening outside the United States?
That was the question that inspired the work of Andrés Abeliuk and Vanessa Gaete, researchers in the Department of Computer Science at the University of Chile and at the National Center for Artificial Intelligence (CENIA), along with Naim Bro, a professor at the School of Public Policy at Adolfo Ibáñez University (UAI) and a researcher at the Millennium Institute Foundational Research on Data IMFD).
Their study, published in EPJ Data Science, used data from the CEP Survey to assess whether four language models—GPT-3.5, GPT-4, Llama-13B, and Mistral—were capable of accurately representing the opinions of the Chilean population. In addition , it compared those results with those obtained using the American National Election Studies (ANES) survey.
Lower accuracy among historically underrepresented groups
The results showed that the models perform with less accuracy in Chile than in the United States. In particular, the loss of accuracy was greater among women, older adults, People , those with lower levels of education, those in lower-income groups, and those who do not declare a political identity. The models also had difficulty representing less common profiles, such as People who are also religious or have low levels of education.
In the United States, however, the main errors were concentrated in variables related to race and political identity
According to the researchers, these differences can be explained by the fact that the models are better at predicting People for People profiles appear more frequently in the data used to train them. Since most of that information comes from developed countries—especially the United States— the models’ performance declines when they are asked to represent societies that are less represented in that data, such as Chile.
Among the models evaluated, Llama-13B achieved the best overall performance, followed by GPT-4. GPT-3.5 showed more inconsistent results, and Mistral had the greatest difficulty adapting to the Chilean context. Even retraining Llama with more data from the CEP survey only improved a few specific results, without resolving the underlying problem.
Can they replace polls?
For Naim Bro, the growing interest in this technology has a practical explanation. “Surveys are very expensive, and this appears to be a very inexpensive alternative. Instead of conducting a larger survey and spending much more money, I could use a language model to get a sense of public opinion,” he explained in an interview with La Tercera.
However, he adds that the study shows that this promise still has significant limitations: “Our question was whether this would also work in Chile, given that these models have been trained primarily using data from the United States and other developed countries.”
According to the study, the answer is that simulated surveys are not yet 100% effective in replacing traditional surveys, and their errors are not random: they systematically affect certain population groups—precisely those that are often relevant to public policy design.
Looking beyond the Chilean case, the researchers conclude that a language model’s performance depends on the context in which it is used, not just on its technical quality. Therefore, they recommend validating these tools with local data and evaluating their performance across different sociodemographic groups before using them to study public opinion or support decision-making.
The study “
:‘Auditing Socio-Demographic and Cross-Societal Fairness in LLM-Simulated Public Opinion,’”by Andrés Abeliuk (University of Chile and CENIA), Vanessa Gaete (University of Chile and CENIA), and Naim Bro (Adolfo Ibáñez University and Millennium Institute Foundational Research on Data), was published in EPJ Data Science.
