In the modern business environment, organizations are under pressure to cut costs while simultaneously needing to increase efficiency. In such a world where market trends change every week, decisions must be made faster than ever, and to make good decisions, strong data is required. However, this is where the paradox arises, as traditional market research is losing the battle against time and budgets, especially in fragmented and niche areas, and the methodology reaches its limits. Strict data protection standards remain an unavoidable obstacle that further complicates rapid hypothesis testing.
It is precisely in this gap between the need for speed and budget constraints synthetic populations emerge as a necessary and valuable companion to classical market research methods. And today, when artificial intelligence has reached a level where it can convincingly simulate human behavior, attitudes, and reactions, doors are opening to entirely new approaches in market research, from digital twins that statistically mimic real consumers to synthetic populations that replace traditional surveys and focus groups.
However, as Dr. Karlo Knežević, head of AI at Sofascore, warns, it is this very ability to simulate that raises some of the most important questions of the digital economy: Where does useful approximation end, and dangerous illusion begin?
– Answers to this question will determine not only the future of marketing but also the quality of the information on which we base our business, regulatory, and social decisions – claims Knežević.
Photocopy of a Photocopy
It is precisely where technology opens new possibilities that new limitations also emerge. One of the most important questions being raised today is about the long-term quality of the data on which such systems rely. That is, what happens when models begin to learn from content generated by other models, rather than from real human responses?
Knežević states that this question strikes at the core of one of the most serious technical problems of today’s artificial intelligence, known as ‘model collapse’. He warns that the technical risk is increasing as generative systems flood the digital space. As an example, he cites research published in 2024 in the journal Nature, which showed that generative models that iteratively learn from their own results gradually lose diversity and reproduce an increasingly narrow spectrum of patterns.
– Simply put, a model that learns from another model resembles a photocopy of a photocopy. Each subsequent generation is fainter than the previous one – explains Knežević.
In the context of market research, the consequences of such a process can be quite concrete. A system originally trained on real consumer data may over time lose the ability to recognize unconventional preferences, niche market segments, or cultural specifics. If the entire ecosystem begins to recycle synthetic data without regular input of fresh human signals, warns Knežević, market research can turn into a closed loop in which the system confirms its own assumptions instead of discovering something new about real people.
Fundamental Truth
Such a risk is not merely theoretical, Knežević states. Studies like that of Bisbee et al. published in 2024 in the journal Political Analysis have shown that synthetic responses generated by large language models have significantly less variability than real human responses. Average values sometimes align, but the range of opinions, intensity of attitudes, and behavior of specific groups often deviate significantly from reality.
– In other words, data may appear convincing, but conclusions about specific consumer segments can be wrong by ten percentage points or more – emphasizes Knežević, adding that the solution is not to discard synthetic data, but to use it in a disciplined manner.
– Real human data must remain the fundamental reference framework, what is referred to in the profession as ground truth. Synthetic datasets must be regularly calibrated against new field research, and systems should have built-in mechanisms that automatically alert when data diversity begins to fall below a critical level. Furthermore, market research professionals must maintain an active role in assessing quality and interpreting results instead of blindly accepting what the algorithm generates – he explains.
