Home / Business and Politics / Synthetic Population: Simulation is Confirmed by Simulation

Synthetic Population: Simulation is Confirmed by Simulation

Karlo Knežević Sofascore
Karlo Knežević Sofascore / Image by: foto

In the modern business environment, organizations are under pressure to cut costs while simultaneously needing to increase efficiency. In such a world where market trends change every week, decisions must be made faster than ever, and to make good decisions, strong data is required. However, this is where the paradox arises, as traditional market research is losing the battle against time and budgets, especially in fragmented and niche areas, and the methodology reaches its limits. Strict data protection standards remain an unavoidable obstacle that further complicates rapid hypothesis testing.

It is precisely in this gap between the need for speed and budget constraints synthetic populations emerge as a necessary and valuable companion to classical market research methods. And today, when artificial intelligence has reached a level where it can convincingly simulate human behavior, attitudes, and reactions, doors are opening to entirely new approaches in market research, from digital twins that statistically mimic real consumers to synthetic populations that replace traditional surveys and focus groups.

However, as Dr. Karlo Knežević, head of AI at Sofascore, warns, it is this very ability to simulate that raises some of the most important questions of the digital economy: Where does useful approximation end, and dangerous illusion begin?

– Answers to this question will determine not only the future of marketing but also the quality of the information on which we base our business, regulatory, and social decisions – claims Knežević.

Photocopy of a Photocopy

It is precisely where technology opens new possibilities that new limitations also emerge. One of the most important questions being raised today is about the long-term quality of the data on which such systems rely. That is, what happens when models begin to learn from content generated by other models, rather than from real human responses?

Knežević states that this question strikes at the core of one of the most serious technical problems of today’s artificial intelligence, known as ‘model collapse’. He warns that the technical risk is increasing as generative systems flood the digital space. As an example, he cites research published in 2024 in the journal Nature, which showed that generative models that iteratively learn from their own results gradually lose diversity and reproduce an increasingly narrow spectrum of patterns.

– Simply put, a model that learns from another model resembles a photocopy of a photocopy. Each subsequent generation is fainter than the previous one – explains Knežević.

In the context of market research, the consequences of such a process can be quite concrete. A system originally trained on real consumer data may over time lose the ability to recognize unconventional preferences, niche market segments, or cultural specifics. If the entire ecosystem begins to recycle synthetic data without regular input of fresh human signals, warns Knežević, market research can turn into a closed loop in which the system confirms its own assumptions instead of discovering something new about real people.

Fundamental Truth

Such a risk is not merely theoretical, Knežević states. Studies like that of Bisbee et al. published in 2024 in the journal Political Analysis have shown that synthetic responses generated by large language models have significantly less variability than real human responses. Average values sometimes align, but the range of opinions, intensity of attitudes, and behavior of specific groups often deviate significantly from reality.

– In other words, data may appear convincing, but conclusions about specific consumer segments can be wrong by ten percentage points or more – emphasizes Knežević, adding that the solution is not to discard synthetic data, but to use it in a disciplined manner.

– Real human data must remain the fundamental reference framework, what is referred to in the profession as ground truth. Synthetic datasets must be regularly calibrated against new field research, and systems should have built-in mechanisms that automatically alert when data diversity begins to fall below a critical level. Furthermore, market research professionals must maintain an active role in assessing quality and interpreting results instead of blindly accepting what the algorithm generates – he explains.

Perverse Economic Mechanism

When discussing synthetic data, the question of authenticity must also be raised, especially in the advertising industry, where the success of campaigns is often measured solely by digital metrics, making it increasingly difficult to distinguish real human reactions from automated signals that imitate them.

Knežević warns that such an approach can create a dangerous closed loop.

– In such an environment, brands that measure campaign success solely by digital metrics such as impressions, clicks, and interactions risk that a large part of what they consider engagement actually comes from automated systems, not from people who would ever buy their product. This creates a perverse economic mechanism. If the advertising industry relies on synthetic profiles to test campaigns and then measures their performance on platforms filled with bots, a closed loop is created in which simulation is confirmed by simulation. A campaign that is ‘successful’ according to synthetic metrics may be completely invisible to real consumers. Paradoxically, all stakeholders in that chain, from the agency to the platform, can report positive results while actual sales stagnate – warns Knežević.

Change of Mentality

To avoid such a trap, companies must diversify their data sources. Digital metrics need to be combined with physical indicators such as actual sales, visits to sales points, or feedback from customer service. It is equally important to use tools that differentiate human traffic from automated interactions and to demand clear data on the share of verified users from platforms. In the European regulatory framework, such issues are increasingly coming to the forefront. The EU AI Act, which is gradually coming into force, introduces the obligation to label synthetic content in a machine-readable way. According to Knežević, this should in the future enable campaign tracking systems to distinguish interactions of real users from synthetic profiles. However, regulation alone is not enough.

– A change of mentality in the industry is also needed. It must stop fetishizing digital metrics and return to the fundamental question: Have we actually reached the real person who has a real need? – he says.

Lost Minorities

Another dilemma accompanying the development of synthetic populations concerns their ability to truly represent society. Statistical models, no matter how sophisticated, always operate on the principle of generalization. They learn dominant patterns from available data and reproduce them.

– What is rare, unexpected, or statistically marginal, which are precisely minority voices, irrational impulses, and cultural specifics, is the first candidate for loss in the process of synthetic reproduction – says Knežević.

Research suggests that this problem is not negligible. A study by Paxton and Yang from 2024 showed that the median correlation between attitudes generated by large language models and actual human attitudes is only 0.10. In other words, models reproduce only a small part of the actual patterns of human thought. An even more dramatic example comes from the analysis of the 2024 European parliamentary elections, where language models estimated turnout at as high as 83 percent, while actual turnout was 49 percent.

Such discrepancies, Knežević warns, are not just statistical errors but signs that models sometimes fundamentally misunderstand social reality. The problem is partly in the structure of the data on which large language models are trained, which disproportionately reflect the Anglophone, urban, and technology-oriented part of the population. When such a model attempts to simulate a consumer from a local or culturally specific environment, it often projects dominant patterns using very superficial demographic variables. For this reason, synthetic populations, he emphasizes, may be useful for testing hypotheses about average behavior, but they are not sufficient for decisions that affect real people.

Tagged: