ODISSEI Conference 2025

ODISSEI is hosting its next live conference in Jaarbeurs, Utrecht, on 4 November 2025. ODISSEI, the research infrastructure for social science in the Netherlands, connects researchers with the necessary data, expertise and resources to conduct ground-breaking research and embrace the computational turn in social enquiry. This conference seeks to bring together a community of computational social scientists to discuss data, methods, infrastructure, policy application, ethics and theoretical work related to digital and computational approaches in social science research.  

The two keynotes will be delivered by Merce Crosas and Melinda Mills. Mercè Crosas is the Head of the SSH Division at the Barcelona Supercomputing Center (BSC).  She also serves on the CODATA board and co-founded Dataverse. Melinda Mills is a Professor of Demography and Population Health and Director of the Demographic Science Unit at the University of Oxford.

If you want to receive conference updates, please sign up for the ODISSEI newsletter and follow us on Bluesky and LinkedIn

Picture by Dasha Nelina

Registration

Registration for the ODISSEI Conference is free and now open. Click here to sign up as a participant.

Date 4 November 2025
Time09:00-18:00
LocationJaarbeurs Utrecht, Jaarbeursboulevard, Utrecht, Netherlands

Programme

Below is the preliminary programme and abstracts.

Walk-in and registration with coffee and tea

The ODISSEI Marketplace is an initiative to provide ODISSEI partners and infrastructure providers with a platform to showcase their facilities and resources. This format enables researchers to explore a wide range of tools, services, and opportunities available through ODISSEI, fostering connections and facilitating access to valuable research infrastructure. The Marketplace will be operating throughout the day and will host the following partners and facilities: 

  • CBS
  • SURF & SANE
  • Centerdata & LISS panel
  • EYRA
  • Data Donation
  • FIRMBACKBONE Infrastructure
  • ASReview
  • ODISSEI SoDa Team
  • Netherlands eScience Center
  • TDCC-SSH
  • DANS & ODISSEI Portal
  • CLARIAH
  • Ethical Reflection and Consultation Tool
  • KiM Explorer

Room: Progress

The ODISSEI Updates session will explore the current developments of ODISSEI, as well as the future. 

Keynote presentation: The Challenges and Opportunities of New Data Frontiers – Melinda Mills (University of Oxford).

Abstract: The exponential growth of digital, administrative, and linked population data is transforming the social sciences, offering unprecedented opportunities to understand human behaviour, social inequality, and demographic change. Yet, navigating these “new data frontiers” poses profound methodological, ethical, and institutional challenges. In this lecture, Prof Mills will explore how researchers can harness these emerging data ecosystems while maintaining scientific rigour, transparency, and public trust.
Drawing on examples from genetics, demography, and computational social science, Prof Mills will illustrate how combining large-scale administrative records, digital trace data, and longitudinal surveys enables richer causal inference and population-level insights than ever before. These opportunities, however, come with significant hurdles: the increasing complexity of data integration, issues of representativeness and bias, and the growing gap between data availability and researcher capacity. Ethical governance will also be central to the future of social data. The talk will address how frameworks like the GDPR, the EU Data Governance Act, and emerging AI regulations shape what is possible, and how we can balance data protection with scientific discovery. Finally, Prof Mills will consider the skills and collaborations required to thrive in this new environment: investing in data science literacy, interdisciplinary training, and cross-sector partnerships that link academia, government, and industry. Ultimately, the “new data frontiers” demand not only technical innovation but also social imagination. By embracing open, ethical, and inclusive data practices, Social Scientists can ensure that these transformative resources serve both scientific progress and the public good.

 

Large-scale international surveys: innovations and future collaborations

Room: Mission 1

Chair: Peter Lugtig

In this panel, with contributions by the coordinators of the large-scale international surveys in the Netherlands, such as ESS, SHARE, ISSP, EVS, WVS, DPES and GGP. Several large-scale surveys will start with a short talk to illustrate current issues in data collection in the Netherlands for these surveys. After the introductions, a panel discussion and Q&A with the audience will allow for discussion on issues such as how to ensure that hard-to-reach groups are being represented in the large surveys, how to deal with the combination of survey modes in data collection, how to ensure comparability, and how to move towards a future where the different surveys collaborate more in fieldwork.

The International Social Survey Programme (ISSP) – Ruud Luijkx (TiU), Piet Gotlieb (TiU), Tim van Meurs (TiU)

Abstract: In this presentation, we will pay attention to the strict rules imposed by the ISSP secretariat to use an online probability panel and how ISSP collaborates with LISS. Another important issue to be covered is the use of an extensive set of socio-demographic variables.

A Self-Completion Approach for the European Social Survey Tim van Meurs (TiU), Tim Reeskens (TiU), Ruud Luijkx (TiU)

Abstract: Faced with declining response rates and evolving survey environments, the European Social Survey (ESS) has moved toward incorporating online data collection. As part of this strategy, the 2025 ESS will adopt a self-completion approach: alongside traditional CAPI (face-to-face) fieldwork, the survey will also be administered via a sequential self-completion design. Participants will first be invited to complete the questionnaire online (CAWI), with a paper version (PAPI) sent to non-respondents. In addition, as part of broader efforts to harmonize international social surveys, an experimental study is currently underway that integrates the ESS into the LISS Panel, in parallel with the other two data collection strategies (CAPI and CAWI/PAPI). While data collection is ongoing, this presentation will reflect on the challenges encountered during the preparation of this mixed-mode approach (i.e., translation processes, sampling designs, timing of fieldwork, and so on), and explore the implications of self-completion modes for the future of international survey research.

Large-scale RIs in transition: Integration of multiple modes and data sources across countries in SHARE 2.0 data collection – Bella Struminskaya (UU)

Abstract: The SHARE survey is facing the challenge of moving the survey to a furture of mixed-modes. This is challenging in a setting where the oldest old are being interviewed, and physical measurements are at the core of SHARE

European and World Values in the Netherlands: Towards A Renewed Marriage? Mark Boukes (UvA), Ruud Luijkx (TiU), Tim Reeskens (TiU)

Abstract: Fielded for the first time in 1981, the European Values Study (EVS) aimed to investigate the influence of secularization and other postwar changes on individual belaief systems using a cross-national and longitudinal survey design. To broaden this scope, the World Values Survey (WVS) launched a global comparative survey at the end of the 1980s on values and cultural change, which builded on the EVS core methodologies and survey items. Although the two initiatives were originally well integrated, they gradually diverged, eventually resulting in situations where both surveys, using nearly identical questionnaires, were fielded almost simultaneously in the same countries. This presentation shifts focus from historical developments to current preparations and future challenges; outlining the fieldwork underway and the broader goals for data collection and knowledge generation in this next phase of EVS – WVS collaboration.

GGP’s Mixed-Mode Journey – Yuri Pettinicchi (NIDI)

The Generations and Gender Programme (GGP) is at the forefront of implementing mixed-mode data collection in cross-national research. This presentation outlines the progress made so far in transitioning to a mixed-mode approach, drawing on experiences from multiple countries. It reflects on the operational and methodological challenges encountered, the strategies adopted to address them, and the valuable lessons learned that can inform future large-scale, comparative social surveys. 

Workshop – Large Language Models in Social Sciences

Room: Mission 2

Workshop leader: Qixiang Fang (Universiteit Utrecht, ODISSEI SoDa Team)

This hands-on tutorial introduces social science researchers to practical applications of large language models (LLMs) for data annotation and analysis. As LLMs such as OpenAI’s ChatGPT become increasingly accessible, they offer powerful tools for automating tasks traditionally done manually—such as coding open-ended survey responses, conducting sentiment or thematic analysis, and extracting structured information from unstructured text. However, effective use of LLMs requires an understanding of both their capabilities and their limitations. This tutorial bridges that gap by guiding participants through the end-to-end process of using LLMs for common social science workflows. Participants will learn how to formulate effective prompts, construct annotation schemes, validate model outputs against human coding, and integrate LLM-assisted analysis into broader research pipelines. The session emphasizes transparency, reproducibility, and critical reflection on model bias and error. Prior programming experience is recommended; tools and examples are accessible as simple Python and R notebooks (see this link). By the end of the tutorial, participants will have the skills to prototype LLM-powered annotation and analysis tasks in their own research, along with a conceptual understanding of when and how to use these models responsibly in empirical social science work.

Building Blocks for FAIR in SSH

Room: Progress

Chair: Angelica Maineri

Finding Data for Reuse – Showcasing SSH Data Discovery Platforms in the Netherlands, Ricarda Braukmann (Data Archiving and Networked Services), Jetze Touber (Data Archiving and Networked Services)

Abstract: The first step to reuse existing data is to become aware of its existence. While there are many great and relevant datasets available in the Netherlands, data discovery can still be challenging, limiting the potential that existing data has. Assessing how the Findability of data in the Dutch Social Sciences and Humanities (SSH) community can be improved is hence one of the aspects that the FAIR workstream of the SSHOC-NL project aims to address. In this talk, we would like to showcase the existing data discovery platforms available for researchers in the SSH domain. One of such platforms, the ODISSEI Portal, allows social science researchers to find data from different data providers including survey data, experimental research, and register data available at CBS in one search interface. Similarly, for the humanities, CLARIAH has developed Ineo to provide one entrance to search across available datasets. We present the current features of these data discovery platforms, as well as providing an outlook into steps that the project wants to take to further improve the Findability and discoverability of available data.

Standardizing Access to Sensitive Data with Machine-Actionable Policies, Lucas van der Meer (Erasmus Universiteit Rotterdam), Ahmad Hesam (SURF)

Abstract: The reuse of research data is widely acknowledged as a driver of scientific progress. However, access to sensitive datasets—such as those containing personal, health-related, or proprietary information—remains a significant challenge. In the Netherlands, researchers and data holders face a fragmented and inconsistent landscape of data access policies. Each major data provider, from Statistics Netherlands (CBS) to RIVM and various health registries, enforces its own procedures, legal terms, and review processes. This lack of standardisation creates confusion, lengthy delays, and administrative burdens for all parties involved. Although many data holders are willing to share sensitive datasets under specific conditions, the current system fails to clearly communicate these access conditions in a structured and machine-readable way, as mandated by the Accessibility principle (the A in FAIR). Metadata often lacks critical details about eligibility, permitted uses, and user obligations, making it difficult for researchers to judge whether and how they can gain access. This presentation offers a vision for improving access to sensitive data through the standardisation of access conditions using a machine-actionable language such as the Open Digital Rights Language (ODRL). Based on community consultation with Dutch researchers and data holders, the authors argue that encoding access policies in a digital, interoperable format—supported by user-friendly tools and a coordinating broker service—would lead to a more efficient, secure, and FAIR-compliant system. Such an approach could significantly increase the availability and responsible reuse of sensitive data for research.

FAIR Vocabularies and where to find them, Angelica Maineri (Erasmus Universiteit Rotterdam), Liliana Melgar ( KNAW Humanities Cluster), Menzo Windhouwer (KNAW Humanities Cluster), André Valdestilhas (Vrije Universiteit)

Abstract: FAIR vocabularies help make data interoperable – i.e., they make it possible to exchange and reuse data across different sources. They can be used to structure the data consistently, or to point to meanings of terms as defined by wider disciplinary communities. However, FAIR vocabularies are difficult to find and reuse, especially in the social sciences, where the uptake is still low. During the presentation, we focus on FAIR vocabularies as resources enabling the I principle of FAIR (Interoperability), and explain the difference between the main types of vocabularies (e.g., controlled vocabularies, thesauri, ontologies, etc…) by means of concrete examples. Moreover, we demonstrate the use of the SSH FAIR vocabulary registry, a platform developed by CLARIAH to facilitate the retrieval and reuse of FAIR vocabularies, by making it possible to search, query and inspect vocabularies that are relevant for the field. Next to the Registry, we also present the SSH FAIR Vocabulary Guidelines, namely step-by-step tutorials which guide SSH practitioners not only in finding but also in creating, adapting, and reusing, FAIR vocabularies. By the end of the presentation, the audience will have a better understanding of FAIR vocabularies concepts and applications, and know about the SSH FAIR Vocabulary registry to search vocabularies that are relevant to the SSH field.

Assessing the Quality of Linked Data, An-Chiao Liu (Universiteit Utrecht), Peter Lugtig (Universiteit Utrecht)

Abstract: Linking data from two or more sources is widely used to answer research questions in many fields. The quality of linked data is naturally affected by each data source. These data sources may be subject to selection bias and/or measurement errors, and the quality of the linkage process also plays an important role. If we overlook potential quality issues when analyzing linked data, a misleading conclusion may be drawn. Much research assessing the quality of linked data can be found in the literature. Many of them focus on the amount of the correct linkage or try to quantify the difference between the linked records. For example, match rate or Damerau-Levenshtein distance. However, the exact quality of the inference depends on the research questions in mind. Different types of research questions may reflect different sensitivity to the data quality. In this paper, we discuss possible quality issues of the linked data by giving a general framework. Our framework aims to help individuals interested in data linkage conceptualize potential problems in their studies and provide an overview of existing methods that can potentially detect bias and enhance the quality of linked data. Examples from the humanities and social sciences are shown to illustrate common problems, and some practical suggestions are provided.

From Corporate to Employment Data – Using FIRMBACKBONE for research

Room: Quest

Chair: Peter Gerbands

Which Businesses Answer Surveys? Evidence from Dutch Administrative Data, Jack Fitzgerald (Vrije Universiteit)

Abstract: I leverage a unique administrative register covering the universe of establishments in the Netherlands to examine how characteristics differ between establishments that do and do not respond to business surveys. Only 20% of Dutch establishments responded to regional business surveys in 2022. Responsive establishments employ two fewer people, and exhibit parttime employment rates 15 percentage points higher, than unresponsive establishments. Sectoral response rates can vary by around 50 percentage points, with public-sector and white-collar occupations overrepresented amongst responsive establishments. Solo enterprises registered to residential addresses are the most common kind of establishment, yet exhibit response rates 18 percentage points lower than an average office. However, controlling for contact probability reveals that the majority of sectoral and occupational variation in response rates can be traced back to differences in contact probability rather than responsiveness. These findings highlight generalizability challenges in business surveys and opportunities to improve their design.

Flattening Hierarchies or Amplifying Control? How Frontier Technologies (Re)Shape (De)Centralization within Organizations, Muskan Achhpilia (Radboud Universiteit Nijmegen ), André van Hoorn (Radboud Universiteit Nijmegen), Imtiaz Sifat (Radboud Universiteit Nijmegen)

Abstract: Do information and communication technologies (ICTs) shift decision-making authority from managers to non-managerial employees, or do they reinforce managerial control? Decentralization – the delegation of decision-making authority to lower levels of the organizational hierarchy – is central to harnessing benefits of comparative advantage through specialization. While prior cross-sectional studies suggest that Information and Communication Technology (ICT) investments affect worker autonomy and span of control, the direction of causality remains uncertain: does technology reshape organizational structure, or does organizational structure determine how technology is used? Our paper investigates the evolving relationship between ICT and decentralization, with a focus on maturity and investments in technology over time and organizational context. Using a panel dataset of 28,691 German establishments (1998–2021), we find robust evidence that ICT investments lead to sustained increase in decentralization over time. However, this effect is smaller than prior literature reported once establishment-level fixed effects are accounted for. Notably, the positive effect weakens over time within a firm, indicating that earlier technology investments reshape workplaces and (de)centralization but frontier technology investments may continue to do so at a decreasing rate and may also have possibility in future of (re)centralizing organizations. Moreover, we find that smaller firms tend to be more decentralized, and that digital technologies do not systematically amplify decentralization in megafirms—challenging the assumption that ICT uniformly shifts authority downward across organizational scales. We conclude a nuanced relationship in which technologies can both decentralize and recentralize decision-making, depending on its use, maturity, and organizational context.

Gender diversity and industrial decarbonisation in the Netherlands, Thelma Arko (Universiteit Utrecht), Xuening Tang (Vrije Universiteit), Kattia Moreno (Universiteit Utrecht), Keling Wang (Erasmus MC)

Abstract: Industrial firms are central to the Netherlands’ goal of reducing emissions by 55% by 2030, making it crucial to understand the drivers of industrial decarbonization. While gender diversity is linked to better sustainability practices, its specific role in industrial decarbonization remains underexplored. This study examines whether higher female workforce participation influences decarbonization in Dutch heavy industries, based on the hypothesis that women’s communal traits foster environmental stewardship, especially when reaching a 30% representation threshold. Using a mixed-methods approach, we analysed firms in carbon-intensive sectors (steel, cement, chemicals, petrochemicals) from 2019 to 2024. We developed a three-stage framework: (1) Agenda Ambition Index (AAI) from web-scraped sustainability commitments, (2) Implementation Index (II) tracking concrete actions, and (3) Realised Sustainability Performance (RSP) from emissions data. The key metric, Agenda Realisation Gap (ARG), was calculated as the Mahalanobis distance between ambition and performance. Data sources included LISA employment data, KvK company records, S&P Trucost emissions data, and web-scraped text from over 3,000 companies. We used linear mixed-effects and two-way fixed-effects models with multiple confounder adjustments. Among 552 firms (898 observations), a one-standard-deviation increase in female participation was associated with a 35.5% increase in the ARG (95% CI: -8.79% to 101.4%). Firms with above-median female participation had a 20.6% higher gap than those below the median (95% CI: 5.87%–37.4%). Sectoral differences emerged: manufacturing showed smaller or negative effects (-10.1%) versus other sectors (28.6%). Contrary to expectations, higher female participation correlated with larger ambition-performance gaps, possibly due to women driving more ambitious targets rather than implementation failures. Manufacturing’s distinct pattern may reflect stricter regulatory oversight. These findings highlight the complex, context-dependent role of gender diversity in corporate decarbonization, warranting further research into its influence on both target-setting and implementation.

On the spatial distribution of headquarters across Dutch firms; an explorative analysis using FIRMBACKBONE data, Wolter Hassink (Universiteit Utrecht), Rutger Schilpzand (Universiteit Utrecht), Nicola Sabatini (Universiteit Utrecht), Peter Gerbrands (Universiteit Utrecht)

Abstract: In this presentation, we report on an explorative analysis on headquarter agglomerations in the Netherlands. We make use of data of FIRMBACKBONE, which consists of information at the establishment level on the entire range of firms over the period 2019-2022. We focus on the location of the headquarter at the municipality level, which may provide information on the prevalence of agglomeration effects. Initial findings indicate that the locational patterns of the head quarter depend on the economic sector and the size of the firm (in terms of number of employees). It suggests that headquarter agglomerations are heterogenous with respect to both dimensions.

Health and Longevity Across Generations

Room: Royal Lobby

Chair: Lina Kramer

External validation of population-based, multimodal prediction models for wellbeing in a depression and elderly cohort, Dirk Pelt (Vrije Universiteit), Philippe Habets (Amsterdam University Medical Center), Martijn Heymans (Amsterdam University Medical Center), Christiaan Vinkers (Amsterdam University Medical Center)

Abstract: Early identification of individuals at risk for mental health problems, as well as understanding the factors that promote wellbeing, is crucial for prevention and intervention strategies. To this end, we previously demonstrated (Pelt et al., 2024, Nature Mental Health) that wellbeing can be reliably predicted in a general population sample (Netherlands Twin Register; NTR) using psychosocial survey data, with modest contributions from environmental exposures, and no added value from genetic predictors. However, to enable broader, cost-effective applications, it is important to assess whether these prediction models are transportable to other populations, such as clinical or elderly cohorts. This study therefore aims to evaluate the external validity of wellbeing prediction models in the Netherlands Study of Depression and Anxiety (NESDA) and the Longitudinal Aging Study Amsterdam (LASA). This cross-cohort comparison seeks to determine whether models maintain predictive accuracy and to identify both shared and unique predictors across different life stages and mental health contexts. Objective environmental predictors were obtained by linking participants’ postal codes to registry-based exposures, providing information on, for example, air pollution, neighborhood characteristics, and greenspace availability (117 predictors). These were combined with 96 self-reported psychosocial features and 22 polygenic scores (PGS) covering a wide range of domains, and used as input for ordinal gradient boosting models predicting life satisfaction ratings.
Preliminary results on the survey data indicate that models trained on the NTR (ordinal AUC: .80 [.75 – .82] performed similarly in NESDA (o-AUC: .84 [.79 – .88]) but less well in LASA (o-AUC: .68 [.59 – .77]). In addition, the most important features were highly similar across cohorts, including neuroticism, partner status, and self-rated health. Results for models based on environmental and genetic data are forthcoming. These findings support the feasibility of transferring wellbeing prediction models across populations, with implications for developing targeted, personalized mental health interventions.

Do People Know What Breaks Their Hearts?, Fanny Tallgren (Erasmus Universiteit Rotterdam), Bram Wouterse (Radboud University Medical Center), Owen O’Donnell (Erasmus Universiteit Rotterdam)

Abstract: Accurate self-assessment of health risks is critical for optimal individual decision-making and prevention. This paper studies how accurately individuals assess their own risk of developing cardiovascular disease (CVD) and whether they appropriately weigh common risk factors. We link individual-level data from the Dutch LISS panel to longitudinal administrative microdata from Statistics Netherlands (CBS), including hospital discharge records, cause-of-death registers, and medication prescription data. This combined dataset enables us to compare subjective CVD risk perceptions with an objective risk function as well as actual CVD events occurring within five and ten years after the survey. We find that individuals substantially overestimate their CVD risk. However, subjective beliefs are positively associated with realised outcomes. Those who develop CVD on average report probabilities 5.7 percentage points higher than those who do not. We decompose this predictive power and show that most of it is explained by observable risk factors such as smoking, blood pressure, cholesterol, and diabetes. Private information such as family history plays a limited role. Most miscalibration arises from underestimating age-related risk increases. We further analyse how individuals interpret their change in CVD risk over time by comparing their five- and ten-year probability assessments. Despite actual CVD incidence doubling between the two periods, subjective probabilities increase only marginally, suggesting a “flatness bias” in risk perceptions over time. Furthermore, we find that those with lower education and numeracy skills display larger miscalibration, less appropriate weighting of risk factors, and weaker predictive power in their beliefs. Finally, we find that verbal statements about perceived risk (“”I think I am at high risk””) are also slightly misaligned with actual health outcomes. This study contributes to the literature on biases in disease risk perception by combining individual level subjective belief data with rich administrative data on realised health outcomes.”

Tracing Longevity Across Generations: Health and Disease Trajectories in Familial Longevity, Niels van den Berg (Leiden University Medical Center)

Abstract: The genetic component underlying longevity represents key mechanisms contributing to a life-long decreased mortality and morbidity risk. Identifying the mechanisms involved is challenging, mainly because of uncertainty in defining long-lived cases with heritable longevity amongst phenocopies and complex gene x environment interactions. Hence, we investigated the longevity trait and its transmission from one generation to the next. In large-scale family-tree data from Utah (UPDB) and the Netherlands (LINKS), we studied 20,360 unselected families containing index persons, their parents, siblings, spouses, and children, comprising 314,819 individuals. We found strong evidence that longevity is transmitted as a quantitative genetic trait among the top 10% survivors of their birth cohort. The survival advantage amounted to 31% for individuals with top 10% surviving first and second-degree relatives in both databases and across two generations, even in the absence of non-long-lived parents. Subsequently, we developed the Longevity Relatives Count (LRC) score as an instrument to quantify the number of long-lived family members and observed that the survival advantage of study participants increased with each additional long-lived family member. Applying the LRC score to the LLS (Netherlands) and SEDD (Sweden; register data) showed that an increasing number of long-lived ancestors associates with an increasing delay in disease incidence (Fig1). As compared to their partners, members of long-lived families have a delayed onset of medication use, multimorbidity and blood-based profiles indicating improved metabolic health and low inflammation in mid-life. Our results indicate that an increasing number of long-lived ancestors marks a decade of healthspan extension, healthier metabolomics profiles, and can be used for more optimized case definitions. Building on these findings, future work will refine and generalize the LRC score into a broader family-based survival metric, enabling wider application across existing studies and supporting the disentanglement of gene x environment interactions in healthy aging and longevity.

The Absence of Genetic Risk for Cardiovascular Disease in Exceptional Survival (Longevity), Pedro Ferreira (Leiden University Medical Center), Marian Beekman (Leiden University Medical Center), Niels van den Berg (Leiden University Medical Center)

Abstract: Aging is the major risk factor for chronic diseases. Unlike the general population, members of long-lived families maintain exceptional health as they age. Healthy survival to extreme ages (longevity) clusters within families. However, research has not yet elucidated the underlying mechanisms of longevity. There are two important reasons for this: 1) the group with the highest heritability is often not studied and 2) the hypothesis-generating nature of most genetic studies requires larger study cohorts. In our previous work, we showed that members of the longest-lived families have a 10-year delayed onset of their first chronic diseases. We therefore hypothesize that the absence of genetic predisposition to chronic diseases is one of the key-mechanisms involved in longevity delayed disease onset. We investigated this hypothesis in the Leiden Longevity Study, a cohort with data from more than 400 long-lived families in 3-generations. To analyze our data, we constructed a set of PRSs covering the top 10 causes of death in the Netherlands. We observed that descendants of long-lived families have lower genetic risk for cardiovascular disease (CVD). Using accelerated failure time modeling, we further showed that around 20% of the delayed cardiovascular disease incidence in long-lived families is explained by CVD common genetic variants. We conducted gene-annotation enrichment analysis of the SNPs in the CVD PRS using DAVID and observed seven significantly enriched clusters. Finally, we constructed a novel cholesterol PRS based on the cholesterol metabolism cluster which significantly predicted time to all-cause mortality in a 90+ study population, covering 19 years of follow-up. Our study indicates that common variants related to cardiovascular diseases and cholesterol metabolism contribute to healthy aging. Furthermore, we demonstrate that investigating SNPs, identified in well-powered Genome Wide Association studies, associated with longevity-related endophenotypes can provide insight into the genetic architecture of the longevity phenotype itself.

Integrating Sources, Reducing Errors: New Frontiers in Data Science

Room: Expedition

Chair: Daniel Oberski

Citizen-in the-loop in training and updating machine learning models for smart data, Barry Schouten (Centraal Bureau voor de Statistiek), Chris Lam (TU Eindhoven), Marco Puts (Centraal Bureau voor de Statistiek)

Abstract: Accurate and timely training data are a crucial prerequisite to (supervised) AI/ML methods. For some AI/ML applications training data can be collected at relatively low cost. However, often an investment is needed in evaluating, annotating and checking events of interest. Apart from costs, annotations and corrections may not be feasible without the involvement of the data subjects. We consider official statistics AI/ML (re)training data settings where the help of ‘citizens’ is imperative to obtain accurate training data and to keep training data up to date. As a secondary motivation, especially in the official statistics context, we see added value in involving general populations in terms of transparency, engagement and trustworthiness. Citizen-in-the-loop strategies are an interplay between a statistical sampling viewpoint and a methodological annotator recruitment viewpoint. The sampling viewpoint aims at efficacy of citizen involvement, i.e. AI/ML performance. The annotator viewpoint considers annotator psychology in terms of burden and competence. Together they lead to a focus on efficiency: how to reach and maintain sufficient performance while not overloading or overburdening citizens. The main elements in our strategy are a distinction of different citizen involvement phases and a trade-off against in-house annotation by staff. We see four phases: feature selection, baseline training, targeted training and updating. Each requires different citizen involvement. The decision as to which annotation queries may be allocated to citizens is set against the proportion which are to be allocated to in-house staff. A key role in the trade-off is annotator bias, i.e. measurement errors made in the annotations. Major methodological challenges are representation and drift in features and concepts. We evaluate ideas and tactics at the hand of case studies in the area of smart surveys and machine learning applied to the resulting smart data.

Being certain about the uncertain: Prediction intervals with missing data, Florian Van Leeuwen (Universiteit Utrecht), Thom Volker (Universiteit Utrecht), Stef van Buuren (Universiteit Utrecht)

Abstract: The development and application of prediction models is often complicated by missing data. Most statistical methods do not readily allow for incorporating missing values during estimation. Consequently, the model of choice cannot be estimated. Imputation can be a solution: the missing values are replaced by values drawn from an imputation model, after which the prediction model can be estimated. Broadly, two strategies can be defined: single and multiple imputation. With single imputation, each missing cell is filled in with the most probable value given the other variables in the data, under some model. Historically, single imputation has been the method of choice in a prediction context. It is simple, readily fits in subsequent workflows and may not necessarily harm predictive accuracy. However, if the interest is in accurate predictive uncertainty quantification, single imputation is problematic. Attempting to reconstruct the missing value precisely is often naïve, and by doing so, the distribution of the data is heavily distorted. Important, the variability in the data is reduced, resulting in uncertainty estimates that are too small. We propose a multiple imputation-based workflow that yields valid uncertainty estimates Multiple imputation accounts for the uncertainty around the missing values by drawing imputations from a distribution. The uncertainty in the imputations is propagated through the prediction model, and the scale of this uncertainty can be used to quantify the prediction uncertainty using Rubin’s rules. We show that, using this approach, valid prediction intervals can be constructed in a linear model setting, regardless of whether missing data occurs in training or testing data, or both. Moreover, the prediction interval is valid regardless of the amount of missing data, and is individualized, in the sense that it scales with the amount of missing data an individual record has. We implemented the required estimation method in the R-package MICE.

Filling the Gaps Through Amplified Asking: Combining Register and Survey Data to Understand Carbon-Intensive Time Use, Maike Weiper (Universiteit Utrecht), Javier Garcia-Bernardo (Universiteit Utrecht), Erik-Jan van Kesteren (Universiteit Utrecht), Jiamin Ou (Universiteit Utrecht)

Abstract: Statistical institutions collect rich register data and administer valuable surveys, but the latter often cover only a small portion of the population. While combining register and survey data already holds great promise, a new range of opportunities emerges when we can also combine multiple surveys – even when they lack (substantial) overlap. This research project explores how we can leverage Dutch register data to impute survey variables for individuals outside the original sample, thereby effectively expanding the scope of survey-based research.
Our initial focus is on time-use data, a uniquely informative but resource-intensive type of data, typically gathered through detailed diaries that are impractical to collect at scale. We are developing a method to generate realistic estimates of how individuals across the Netherlands spend their time, using demographic and household-level register data. This approach may make it possible to link time-use data to other survey domains, such as income, health, or consumption, and thus study previously inaccessible patterns, including intersectional forms of gender inequality or individual-level carbon emissions. In this presentation, we will introduce a preliminary approach for producing such large-scale time-use estimates by combining various sources of register data. As the project develops, particular attention will be given to addressing methodological challenges, including the generalizability of predictions across social groups, potential biases in imputed data, and the broader implications of applying such methods in survey research. Ultimately, the project aims to develop a scalable and responsible framework for expanding the reach of survey research. By enabling new combinations of survey content and increasing representativeness, we hope to contribute to more inclusive, data-driven understandings of societal dynamics in the Netherlands.

A Two Step Approach for Modeling Total Error with Multi-source Data, Santiago Gómez-Echeverry (Vrije Universiteit), Arnout van Delden (Centraal Bureau voor de Statistiek), Ton de Waal (Centraal Bureau voor de Statistiek), Dimitris Pavlopoulos (Vrije Universiteit)

Abstract: The expansion of administrative and big data and the increase in the survey’s non-responses have highlighted the relevance of assessing the quality of non-probability samples. To tackle this issue, people usually resort to the Total Error (TE) framework, which divides the error into a measurement and a representation component. Extensive literature focuses on measurement error, often using a combination of data from different sources to evaluate whether the observed variables adequately capture the concept intended to be measured. Another branch of the literature has centered on the representation error, assessing how respondents are selected in the sample, leading to systematic differences between the population and the observed units. However, research modeling both of these components simultaneously is still scant. In the present study, we address this gap by jointly modeling the measurement and the representation errors, combining recent advances in both areas. We conducted a simulation study to evaluate our TE model under different specifications of measurement and representation errors. Additionally, we performed a case study analysis using a combination of Italian administrative registers and the Labor Force Survey (LFS) to evaluate the TE in the income variable. Our preliminary results show that our model adequately captures the different error sources and provides a good strategy for assessing the TE when using a combination of probability and non-probability data.

Poster presentations

  1.  An Ai approach to Investors narratives shaping biodiversity as an asset class, Catalina Papari (Universiteit Utrecht), Qixiang Fang (Universiteit Utrecht), Helen Toxopeus (Universiteit Utrecht)
    Abstract: The biodiversity crisis requires significant financial investment, yet private finance remains limited and underdeveloped. For biodiversity to be integrated into financial markets, it must be framed as an investable asset with identifiable risks, returns, and impacts. This paper develops a theoretical framework defining the core features of biodiversity as an asset class and applies it to analyze investor narratives using prompt engineering with large language models (LLMs). We examine sustainability-related reports from 2020 – 2025 of 254 financial institutions committed to biodiversity finance. After mining the PDF texts and extracting all the paragraphs, we apply a three-step script to assess the degree of ‘assetization’: (1) using a biodiversity keyword list and the Schimanski et al. (2023) LLM to extract only relevant paragraphs about biodiversity; (2) filtering for investor-specific biodiversity activities via prompt-based classification; (3) applying a second prompt to detect three theoretical elements of assetization (current value – cash flow, ownership, and future value generation). Findings reveal that biodiversity is often framed in asset-like terms—particularly by asset managers—with stronger emphasis on risk-return than on ownership or rent-sharing. Two asset types emerge: adjacent assets (e.g., green roofs) and stand-alone assets (e.g., regenerative agriculture), with land-based assets more clearly positioned as investable than water-related ones. These insights are critical for shaping future financial instruments and policies that move biodiversity finance beyond offsets toward long-term ecological and economic value creation.
  2. Human centred explainable AI decision-making in healthcare, Catharina van Leersum (Open Universiteit), Clara Maathuis (Open Universiteit)
    Abstract: Human-centred AI (HCAI1) implies building AI systems in a manner that comprehends human aims, needs, and expectations by assisting, interacting, and collaborating with humans. Further focusing on explainable AI (XAI2) allows to gather insight in the data, reasoning, and decisions made by the AI systems facilitating human understanding, trust, and contributing to identifying issues like errors and bias. While current XAI approaches mainly have a technical focus, to be able to understand the context and human dynamics, a transdisciplinary perspective and a socio-technical approach is necessary. This fact is critical in the healthcare domain as various risks could imply serious consequences on both the safety of human life and medical devices. A reflective ethical and socio-technical perspective, where technical advancements and human factors coevolve, is called human-centred explainable AI (HCXAI3). This perspective sets humans at the centre of AI design with a holistic understanding of values, interpersonal dynamics, and the socially situated nature of AI systems. In the healthcare domain, to the best of our knowledge, limited knowledge exists on applying HCXAI, the ethical risks are unknown, and it is unclear which explainability elements are needed in decision-making to closely mimic human decision-making. Moreover, different stakeholders have different explanation needs, thus HCXAI could be a solution to focus on humane ethical decision-making instead of pure technical choices. To tackle this knowledge gap, this article aims to design an actionable HCXAI ethical framework adopting a transdisciplinary approach that merges academic and practitioner knowledge and expertise from the AI, XAI, HCXAI, design science, and healthcare domains. To demonstrate the applicability of the proposed actionable framework in real scenarios and settings while reflecting on human decision-making, two use cases are considered. The first one is on AI-based interpretation of MRI scans and the second one on the application of smart flooring.
  3. Bias in Italian Open Government Data ecosystems: disparities and impacts on SME innovation, Davide Serioli (University of Milan-Bicocca)
    Abstract: Open Government Data (OGD) holds significant potential to drive social and economic progress by fostering transparency, innovation, and public participation in governance. However, the impact of OGD can be limited if the data remains underutilized. In this order, this research could be interesting as it explores the use and quality of OGD, highlighting its potential economic and social benefits. Adopting a critical data studies perspective, this study examines the Italian OGD ecosystem to uncover usage trends, metadata quality, and their role in perpetuating socio-economic inequalities, particularly for SMEs’ innovation potential. Drawing on frameworks from Neumaier et al. (2016) and Quarati (2023), we analyze regional and municipal OGD portals (e.g., dati.gov.it, dati.lombardia.it) to map disparities and biases. Through computational social science methods, we explore three key dimensions: (1) the distribution of OGD across regions and themes, highlighting underrepresented areas and topics; (2) the quality and representativeness of OGD, including datasets like ANAC’s public procurement data and innovation-related data (e.g., funding opportunities); and (3) the societal impact of OGD biases, particularly on SMEs’ access to innovation resources. Using data from ISTAT (e.g., socio-economic indicators, SME metrics) and OGD portals, I employ quantitative and geospatial analyses to reveal “data deserts” and their implications. Preliminary findings confirm significant underutilization of OGD and variations in metadata quality, with no clear correlation between quality and usage. Moreover, OGD coverage is skewed toward wealthier regions, with fewer datasets on social issues or rural areas, potentially limiting SMEs’ innovation opportunities in marginalized regions. This exploratory study underscores how OGD biases reinforce disparities, offering insights for inclusive data policies to support equitable innovation. By combining large-scale data analysis with a critical lens, this work informs strategies for a more equitable OGD ecosystem.
  4. Drawing the Line: An Experimental Study on Tolerance and the Limits of Religious Expression in the Netherlands, Guido Priem (KU Leuven), David Kretschmer (University of Oxford), Eva Jaspers (Universiteit Utrecht)
    Abstract: Many people in Western Europe hold prejudices against Muslims and, therefore, oppose their inclusion in any way. However, even those with positive attitudes may still reject specific Muslim practices that conflict with their moral beliefs, such as gender equality or secularism. Unwillingness to tolerate Muslim practices can therefore reflect prejudice, a principled disapproval of the specific expression, or both, making it empirically difficult to distinguish the two. To disentangle these effects, this study argues that we should examine how tolerance decisions towards religious expressions are shaped by a range of different attributes simultaneously, including the actor’s background, the message and the setting where it takes place. Moreover, this should be tested for religious expressions by majority groups (Christians) and minority groups (Muslims) alike, who are, in theory, equally entitled to claim their rights to express their religion in public. This study, funded by the LISS Panel grant of 2024, will combine a single-profile conjoint experiment with a between-subject experiment embedded within the LISS panel infrastructure. First, we randomly assign the respondents to see either Christian or Muslim religious expressions. We then present respondents with a series of instances of religious speakers with this background and ask respondents to judge if they should be allowed to speak or not. We randomly vary who is speaking, what is said, and where it takes place, allowing us to separately estimate the impact of each attribute on decisions to tolerate religious expressions. Through a Bayesian Hierarchical Model, we also test if the strength of the effects of these attributes differs depending on the religious background of the speaker. This allows us to answer two related questions: 1) What contextual attributes of a religious expression shape tolerance decisions, and 2) are these effects consistent between expressions by religious majority and minority groups?
  5. Changing media representations of delinquent youth: a topic model analysis of Dutch newspaper coverage from 1995 to 2024, Inge Aarts (Erasmus Universiteit Rotterdam)
    Abstract: Misconduct among young people often raises societal concern. In the news media, juvenile delinquency is repeatedly ‘rediscovered’ as a societal issue, shaped by the socio-political and cultural contexts of the time. Previous research has primarily analyzed this concern through qualitative content analyses of specific key events, resulting in studies that focus mainly on relatively short time periods. However, mapping potential changes and patterns in the way newspapers write about juvenile delinquency requires a longitudinal research perspective. In light of this, this study explores developments in both the quantity and thematic focus of Dutch newspaper coverage from 1995 to 2024. The corpus includes approximately 50.000 articles related to juvenile delinquency from multiple Dutch newspapers, allowing for a focus of potential differences in the coverage between national and regional newspapers and quality and tabloid newspapers. The corpus will be analyzed with (dynamic) topic modelling, a computational text analysis technique that has gained traction within social sciences due to its ability to offer insights into the textual ‘aboutness’ of a corpus. Therefore, this study will first consider whether, and in what ways, the messages about juvenile delinquency conveyed by Dutch newspapers have become more frequent or have shifted in character. Second, building on Entman’s (1993) theory of framing and Becker’s (1963) concept of moral entrepreneurs, this study explores the value of topic modeling for identifying framing within the reporting of juvenile delinquency. Understanding which topics are reflected in this newspaper discourse allows for clearer insight into the role of news media in shaping collectively recognized meanings and assumptions about delinquent youth.
  6. The effect of the network of children of migrants on their Dutch language proficiency, Jan van der Laan (Centraal Bureau voor de Statistiek), Fijnanda van Klingeren (Centraal Bureau voor de Statistiek), Marjolijn Das (Centraal Bureau voor de Statistiek)
    Abstract: Even though the school results of children of migrants are improving in the Netherlands, they still are lower than those of native Dutch children. Language proficiency plays an important role in the school career of children. Therefore, it is important to understand which factors contribute to Dutch language proficiency of these children. Besides formal education, exposure to the language plays an important role. Children not only get exposed to a language via their parents, but the entire social network plays a role. This study examens to what extent the language proficiency of children with two parents born outside of the Netherlands) is affected by the network composition of those children. Using the Person Network of Statistics Netherlands the isolation of each child is calculated. The Persons Network contains family members, household members, neighbours, co-workers and class mates for each person in the Netherlands. The isolation measures the extent to which the local network around a person, the ego network, consists of persons with the same migration background. It takes into account both direct contacts and indirect contacts. The effect of isolation on language proficiency is modelled using regression with taking into account possible confounding variables. We find that higher isolation score are related to lower scores on the language component of the final exam of the 8th grade of the primary school (around the age of 12). This in turn affects the track recommendation for the lower secondary education.
  7. Prosocial persuasion at scale? Effects of personalization and content source in social media donation appeals, John Caffier (Tilburg University), Bennett Kleinberg (Tilburg University), Olga Stavrova (Tilburg University)
    Abstract: A 3 × 2 within-subjects study was run online via Prolific with 658 U.S. participants. Each participant saw six charity appeal posts that varied by personalization (generic, personalized, counterfactual) and source (human vs. LLM). After viewing each post, participants rated their engagement (–1 = Dislike to +1 = Like), judged persuasiveness on three 1–7 scales (α = .93), and chose how much of a $0.10 bonus to donate. Donation amounts did not differ between human-authored and LLM-generated posts, nor between personalized and generic messages. Counterfactual personalization yielded a small drop in donations. In contrast, LLM-generated content drove higher engagement and persuasiveness ratings than human content. Personalized posts gave a modest boost to engagement but did not affect persuasiveness, while counterfactual personalization reduced both. No source × personalization interaction occurred for donations, but LLM gains in engagement and persuasiveness held across all personalization types. These findings show that LLMs can match human writers in collecting real donations and outperform them in attracting attention and shaping persuasive impressions. LLMs thus offer a scalable way to craft prosocial messages. Future work will isolate which linguistic features drive these effects and test smarter personalization strategies.
  8. The determinants of social distancing behavior during COVID-19: quantifying behavioral influences for transmission models, Joshua Chevalier (University Medical Center Utrecht), Leonard Stellbrink (University of Lubeck), Senne Wijnen (GGD Zuid-Limberg), Lisanne Steijvers (GGD Zuid-Limberg)
    Abstract: There are dynamic feedback loops between the progression of respiratory infection epidemics, their (perceived) severity, and human behaviors that seek to mitigate transmission. Quantitative estimates of these relationship dynamics are needed to incorporate human behavior into epidemic transmission models. Our aim was to understand and quantify the influences of social distancing behavior during the COVID-19 pandemic. For three different countries (Canada, Japan, Netherlands) we linked different sources of data—the Imperial College London YouGov COVID-19 Behavior Tracker survey data on pandemic perceptions and behaviors, to the Oxford COVID-19 Stringency Index on country-specific responses to COVID-19, and to COVID-19 hospitalization data from Our World in Data. This was done by date from June 2020 to January 2021. From the survey data, we elicited self-reported social distancing behavior from four questions—frequency of avoiding small, medium, and large gatherings, or crowds on a 5-point Likert scale. Social distancing was considered a composite score of 16 or higher. We performed logistic regression using stepwise AIC with both forward and backward selection and calculated the standardized odds ratios (SOR; per standard deviation increase in the predictor). We then confirmed predictor importance using machine learning decision tree gradient boosting (XGBoost). The top predictors of social distancing behavior, across models and countries, were COVID-19 hospitalizations per 100,000 population (SOR:1.16–1.25), perceived severity of COVID-19 (1.19–1.20), willingness to self-isolate (1.12–1.21), and age (1.05). The Area Under the Curve (AUC) of the logistic regression models was between 0.68–0.74 and 0.62–0.76 for XGBoost, indicating similar fits to the data. Results agreed across model types and country contexts, signifying social distancing behavior was largely influenced by COVID-19 hospitalizations and perceived disease severity. In transmission models calibrated to country-specific epidemiological data, these results can be used to ground modelled behavior changes in real-world data.
  9. Data Quality for Computational Social Science: From Concepts to Infrastructure, Jun Sun (GESIS), Thomas Knopf (GESIS), Fabienne Kraemer (GESIS), Anne Stroppe (GESIS)
    Abstract: Valid conclusions can only be drawn when data meets rigorous quality standards, and when researchers are equipped with the appropriate tools and expertise to assess them. Ensuring data quality has become particularly important within the social sciences due to increasingly complex research settings, a declining willingness to participate in surveys, and the growing reliance on automated and scalable evaluation of available data. To this end, we present two infrastructures, the KODAQS Data Quality Toolbox and the KODAQS Data Quality Academy, developed as part of the BMFTR/EU-funded Competence Center for Data Quality in the Social Sciences (KODAQS). Both the Toolbox and the Academy address data quality issues arising from two fundamental error sources: Representation and Measurement, across three main data types that are prominent in computational social science: 1) survey data, 2) digital behavioral data, and 3) survey data linked with data from other sources. The KODAQS Toolbox provides tools and tutorials on conducting key data quality analysis to help social scientists evaluate and improve the quality of their research data. As an open educational platform, the Toolbox lowers entry barriers for learners by providing reusable code, embedded video tutorials, sample datasets, self-assessments, and downloadable source files – all designed to support hands-on learning. In addition, it incorporates interactive execution environments that allow users to run the learning materials directly in their browser without requiring complex setup. The KODAQS Academy provides necessary theoretical foundations and applied skills to detect, improve, and report data quality aspects for researchers from various disciplines and all career levels who work with social science data. It offers a comprehensive training program on data quality through three components: 1) a six-month blended Certificate Program, 2) a flexible self-paced Individual Training option, and 3) a Train-the-Trainer module providing open resources for integrating data quality into university teaching.
  10. An individual’s own diabetes status and their exposure to diabetes using a whole population network of the Netherlands, Karen van Hedel (Centraal Bureau voor de Statistiek), Edwin de Jonge (Centraal Bureau voor de Statistiek), Marjolijn Das (Centraal Bureau voor de Statistiek)
    Abstract: Diabetes is one of the most common chronic diseases in the Netherlands: approximately 1.2 million people have diabetes, of which 1.1 million have diabetes type 2. Type 2 diabetes is shown to be related to lifestyle factors such as unhealthy eating habits, being overweight and smoking, but also to ageing and it may be genetic. An individual’s social environment may influence their lifestyle (choices) and subsequently also their chance of having diabetes. This study examines the relationship between an individual’s social environment and diabetes with data from the Person Network of Statistics Netherlands. The Person Network contains all household members, (extended) family members, colleagues, neighbours and classmates for each inhabitant of the Netherlands in 2022. Individual’s diabetes status was based on whether they were prescribed diabetes medication in 2022. An exposure score was calculated for each individual indicating to what extent diabetes was present in their (local) network. Results are shown for individuals aged 40 years and older as the prevalence of diabetes, in particular for type 2 diabetes, increases rapidly after age 40. The prevalence of diabetes and the exposure to diabetes in one’s network is presented by gender, 10-year age groups and regional area (on NUTS-3 level). The exposure score is then decomposed for the different network layers (household, family, colleagues, neighbours and classmates) to assess their distinct contribution to the overall exposure score. Furthermore, predicted probabilities are calculated from a logistic regression to examine whether clustering on background variables (e.g. socioeconomic status, household composition) and/or lifestyle factors (e.g. body mass index, smoking history) might explain some of the observed variation in diabetes status by exposure to diabetes. Preliminary results using the Person Network of 2016 suggest that individuals who have diabetes, on average, have a higher exposure to diabetes than those who do not have diabetes.
  11. Predicting adverse birth outcomes based on information known at conception, during pregnancy and immediately after birth in the Netherlands, Kebede Haile Misgina (Erasmus MC), Anton Schreuder (Erasmus MC), Wessel Kraaij (Leiden University), David van Klaveren (Erasmus MC)
    Abstract: Background: Adverse birth outcomes are the leading causes of perinatal mortality and morbidity with consequences that extend to later age and future generations. Early identification of at-risk pregnancies could help develop more personalized preventive interventions and improve outcomes. Objectives: We aimed to develop risk assessment models based on information available at conception, during pregnancy, and immediately after birth. Method: We conducted a prospective cohort of over 600,000 singleton births ≥24 months of gestation in the Netherlands from 2016 to 2019. Using logistic regression analysis with backward selection, we developed and internally validated parsimonious models at conception, right before birth, and immediately after birth. Models were stratified by primi-and multiparity. Predictive performance of the models was evaluated using the area under the receiver operating characteristics curve (AUC), sensitivity, and specificity. Results: Overall, 16.6% births experienced at least one adverse outcome. Predictive performance varied by parity and outcome. Before birth, small-for-gestational age was the most predictable outcome, with multiparous women demonstrating superior model discrimination compared to primiparous followed by stillbirth. The third most predictable outcome was preterm birth. In contrast, congenital anomalies, low Apgar score at 5 minutes, neonatal intensive care unit admission, and neonatal mortality were the least predictable. The main predictors included history of adverse birth outcome, age, parity, absent father, health insurance debt, maternal cardiovascular disease, maternal psychiatric disease, and parental education.For some outcomes, we refitted the models with congenital anomalies and small-for-gestational-age included as predictors, which led to substantial improvement in predictive performance. For example, for neonatal mortality, we observed an AUC of 0.749 in the model right before birth. Conclusions: While some adverse birth outcomes can be well predicted before birth, others remain moderately predictable. These risk models could support early interventions to optimize birth outcomes, child health, and growth. Early detection of congenital anomalies and intra-uterine growth retardation may enhance the prediction of other adverse outcomes.
  12. On the roles of gender and sex multidimensionality in causal inference practice: bias from messy use of dichotomized indicators, Keling Wang (Erasmus MC)
    Abstract: Sex and gender, two distinct concepts, are critical for causal inference practice in social sciences, epidemiology and public health. Dichotomized sex and gender indicators, often messed up with each other, have long been practiced as stratifying variables, effect measure modifiers, potential confounders, or often-ill-defined treatment/exposures. As two distinct, co-constituted, multidimensional, bio-socio-political, dynamical and culture-dependent constructs, sex and gender’s dimensionality and complexity is almost always overlooked. The common practice, using dichotomized indicators and collapsing/discarding minor groups, mismeasures underlying mechanisms, providing bad and malfunctioning proxies. Once sex and gender have any roles in a causal question and mechanisms, these bad proxies can make causal assumptions less plausible and introduce bias. Sometimes the magnitude of the bias could be argued as trivial, but there exist many situations where bias is not negligible and may lead to wrong conclusions and do harm. Therefore, we aimed to examine the misconception or mismeasurement of sex and gender under two situations: (a) when they act as eligibility criteria or treatments, concerning the formulation of a causal research question; and (b) when they act as confounders, effect measure modifiers and selection indicators, concerning the plausibility of causal assumptions. For each of the roles of sex and gender in these situations, we will re-analyze published epidemiologic and social science studies utilizing open trial and cohort data using causal inference and bias analysis methods. We consider sex and gender as classes of traits (e.g. meta-perceived gender, genital status, gender identity, administrator gender indicator assigned at birth), and justify for each case the concerning traits in causal mechanisms using subject-matter knowledge; we then make assumptions about the distribution of traits based on census data and previous research. Possible magnitude of bias that could be introduced by using dichotomous sex/gender indicators or messing them up is reported as output of this work.
  13. Misleading Deception Classifiers With Model-Based and Human Paraphrasing Attacks, Lucca Pfründer (Universiteit van Amsterdam), Riccardo Loconte (IMT Lucca), Bruno Verschuere (Universiteit van Amsterdam), Bennett Kleinberg (Tilburg University)
    Abstract: Background: Automated models often outperform humans at detecting deception but remain vulnerable to adversarial attacks. These are modifications (i.e., paraphrases) of for instance deceptive statements crafted with the intention of fooling the model into a wrong classification (in this case truthful). Importantly, the meaning of the statement is not changed. Method: A DistilBERT classifier was trained on 80 percent of the Hippocorpus dataset (n statements = 4542) which contains autobiographical truths and lies. Following, humans and GPT-4o each rewrote 153 test statements (of the remaining 20%: n = 505) up to 10 times (iterations), attempting to flip the model’s prediction. For each statement, the modification that induced the greatest change in class probability (i.e., classifier’s predicted probability of a statement being truthful (or deceptive)) relative to the original statement was selected for the main analysis. We then investigated who (humans or LLM) was more effective (i.e., larger change) and efficient (i.e., in fewer iterations) at fooling the classifier. Results: Nearly 70% of modified statements succeeded in changing the model’s prediction (i.e., lie to truth or truth to lie). While humans and LLM were similarly effective and efficient overall, humans were more effective and efficient at modifying truthful statements to appear deceptive. This might be because they selectively removed named entities more than GPT-4o which made larger, consistent changes. Conclusion: Overall, we found that deception classifiers can be tricked through adversarial paraphrases from humans and GPT-4o warraning caution of using these models in real life.
  14. Adaptive Tests and Questionnaires Using Online Survey Tools, Lydia Dekkker-Klein Nibbelink (Universiteit van Amsterdam), Goan Booij (Universiteit van Amsterdam), Andries Van der Ark (Universiteit van Amsterdam)
    Abstract: Time constraints are a well-known problem in online surveys: There is often insufficient time to administer all the items necessary for optimal measurement. Measuring constructs using only a few items is undesirable, as these measurements generally have low reliability, which attenuates the estimates of the effects hypothesized by the researcher. Fixed-length computerized adaptive testing (CAT) has been proposed as a solution to mitigate this problem, as it allows for relatively precise measurement with relatively few items. However, many social and behavioral scientists use online survey tools such as Google Forms, LimeSurvey, Qualtrics, SurveyMonkey, or Typeform for data collection, none of which support CAT. We propose a simplified version of fixed-length CAT that can be implemented in online survey tools, called Approximate Multi-Stage Testing (Approximate MST). Depending on the amount of information obtained prior to data collection, a more or less sophisticated form of Approximate MST can be employed. Using the linear administration of all items and fixed-length CAT as upper benchmarks, and the linear administration of a reduced test as a lower benchmark, we will investigate the bias and variance of measurements obtained using Approximate MST with simulated data. We expect that Approximate Multi-Stage Testing will yield the most precise measurements if all estimated item information functions are available prior to data collection. We also expect that if limited information is available (e.g., only item means reported in the literature), Approximate MST will still be more precise than the linear administration of a reduced test. In the presentation, we will present the results and demonstrate Approximate MST using Qualtrics.
  15. Reputation, Trust, and Coalition Formation: Insights from an Interactive Behavioral Game, Merve Timuroğulları (Tilburg University), Seger Breugelmans (Tilburg University), Thorsten Erle (Tilburg University), Frans Cruijssen (Tilburg University)
    Abstract: This study used a recently introduced feature of the LISS panel to conduct a real-time negotiation experiment, combining pre-survey data with interactive group decision-making in a collaborative transport setting. It examined how company reputation, operationalized via Social Value Orientation (SVO), influences trust and coalition behavior. Participants (N = 681; 227 triads) first completed an incentivized SVO measure two weeks before the experiment, indicating their degree of prosocial or individualistic orientation. They were then matched in triads to play an incentivized simple weighted majority game simulating coalition formation between three transport companies. Coalitions required at least two players, who negotiated how to divide a monetary incentive. In the experimental condition, participants saw their counterparts’ SVO scores in a simplified visual gauge as a reputation signal. The control condition included only company information. To measure trust, we developed a two-dimensional scale capturing predictability (assurance, knowledge) and benevolence (prosocial expectations), assessed before and after the information phase, and after negotiation. Against expectations, not all players were equally affected by the reputation manipulation. Player A, the strongest actor in the game, was more frequently included in coalitions, which was linked to a significantly higher frequency of grand coalitions (ABC) when reputation was visible. Player A was also perceived as more benevolent at the second trust measurement, when reputation was revealed. Prosocial participants were more likely to be included, especially when their reputation was shown. They proposed less self-serving offers, formed more inclusive coalitions, and earned more. This may be partly due to the unexpectedly high share of prosocial participants in the sample (78.6%), which likely increased trust and willingness to collaborate under the visible reputation condition. This study contributes to computational social science by combining survey data with interactive group decision-making to examine how social information shapes trust dynamics and negotiation behavior.
  16. A panel analysis of changes in individuals’ perceived accessibility, Milan Moleman (Kennisinstituut voor Mobiliteitsbeleid (KiM)), Iris Roeleven (Kennisinstituut voor Mobiliteitsbeleid (KiM)), Maarten Kroesen (TU Delft), Marije Hamersma (Kennisinstituut voor Mobiliteitsbeleid (KiM))
    Abstract: Perceived accessibility, i.e. the perceived ease of reaching destinations, has received growing attention in transportation research. Unlike spatially inferred accessibility levels, this person-based measure allows individuals living in the a similar geographical area to differentiate in their accessibility levels. Socio-demographics, mobility means, and travel attitudes drive variations in perceived accessibility. The factors contributing to within-person changes in perceived ease of reaching destinations have yet to be explored. In this study, we assess changes in perceived accessibility over time and the factors contributing to these changes. To do so, panel analyses have been conducted using data from 543 individuals of the Netherlands Mobility Panel in 2020 and 2023. Specifically, a longitudinal latent class analyses is employed to identify trajectories in perceived accessibility for different groups. Following this, we estimated a fixed effects regression to assess the factors explaining changes in perceived accessibility. The longitudinal latent class analysis revealed six trajectories in perceived accessibility. Three of these trajectories (61% of the sample) remained relatively stable over time, while the other trajectories indicate notable transitions. Whereas one cluster (12% of the sample) developed a higher perceived accessibility, two clusters (27% of the sample) reported a decline. The fixed effects model identified key determinants of these changes. Interestingly, changes in travel means had substantial effects on changes in perceived accessibility. Additionally, changes in the distance to nearest amenities such as the grocery store, secondary school, and train station contributed to changes in perceived accessibility, yet played a less prominent role than mobility means. The link between spatial accessibility and perceived accessibility underscores that individuals’ perceived accessibility will worsen as amenities become scarcer. However, when amenities are located farther away, having the means to travel becomes essential to maintaining access. It is this interplay between nearby amenities and travel options that should not be overlooked.
  17. A latent-profile approach in modelling individual COVID-19 vaccination decision-making during the COVID-19 pandemic, Mitchell Matthijssen (Vrije Universiteit), Taymara C. Abreu (Universiteit Utrecht), Javier Garcia-Bernardo (Universiteit Utrecht), Vincent Buskens (Universiteit Utrecht)
    Abstract: Background. The World Health Organization has identified vaccine hesitancy as one of the 10 major threats to public health. Previous work defined individual decision making as the core concept of vaccine hesitancy. The heterogeneity in vaccination decision-making hampers implementation of more effective vaccination strategies. Previous work, by applying a variable-centred approach, has shown which factors people use in their vaccination decision-making (e.g., risk-perceptions, values). However, this approach neglects individual differences in decision-making and related risk beliefs and values, i.e., differences in vaccination decision profiles. Taking these differences in vaccination decision-making profiles into account might improve the effectiveness of interventions to address vaccination hesitancy. Therefore, in our project, we use a novel person-centered approach to explore COVID-19 vaccination decision-making profiles in order to be able to model the distribution of these profiles at the regional level to find potential targets for interventions. We used a combination of beliefs (e.g., risk perceptions of COVID-19, safeness and effectiveness of the vaccine) and values (e.g., purity, autonomy, freedom) to develop the profiles. Method. We collected data with the LISS-panel (N = 2,567) and used an exploratory Latent Profile Analysis approach. Results. We inspected the most optimal statistical solutions and determined whether the profiles are theoretically meaningful. Our first analyses suggests that there are 7 meaningful vaccination decision-making profiles, among others based on how people value autonomy and safety and effectiveness of vaccination. Conclusion. In future work, we will use the key feature of the LISS panel to connect to CBS micro-data to link the profiles to demographic, geographical, and social network features of people in these profiles. We aim to model the distribution of these vaccination decision-making profiles in the population in order to help professionals to better target vaccination hesitancy.
  18. Who’s Being Addressed? Predicting Receivers in Multi-Party Conversations with Large Language Models, Myrthe Prins (Universiteit Utrecht), Daniel Oberski (Universiteit Utrecht), Mahdi Shafiee Kamalabad (Universiteit Utrecht)
    Abstract: This study investigates the ability of large language models (LLMs) to infer the intended addressee of a message in multi-party conversations. This task is crucial for enhancing the utility of Relational Event History (REH) datasets, which consist of speaker, time, and receiver information. However REH data frequently lack information on who is being addressed. If LLMs can be used to accurately predict the receiver, incomplete REH data can be made suitable for more comprehensive analyses in the social and behavioral sciences. We evaluate this approach on two real-world datasets: Dutch parliamentary debate transcripts and meeting transcripts from a student association. These represent more formal and informal conversational settings, respectively. Using a few-shot cascade prompting strategy, we test GPT-4.1 via the OpenAI API. The model receives three sequential utterances (two preceding and the current one) and must predict the numeric ID of the intended receiver of the current utterance. Model performance is assessed using standard classification metrics and graph-based comparisons between predicted and ground-truth interaction networks. These results are above chance level, suggesting that LLMs can make plausible addressee predictions based on transcripts. To explore the reliability of the model’s predictions, we extracted log probabilities and confidence scores for each predicted receiver; however, these measures proved to be weak indicators of prediction correctness. This work addresses a key gap in REH data and demonstrates a scalable method for enriching incomplete datasets. It offers practical benefits for research domains such as political communication, sociolinguistics, and organizational behavior, where access to complete interaction networks is essential. While results are encouraging, further research is needed to fully understand the limitations and generalizability of this approach.
  19. How Robust Is Embedding-Based Topic Modelling? A Multiverse Analysis of BERTopic on COVID-19 Survey Responses, Ngoc-Anh Tran (Tilburg University), Bennett Kleinberg (Tilburg University)
    Abstract: Embedding-based topic modelling techniques such as BERTopic offer a powerful, context-sensitive approach to extracting latent themes from large-scale textual data. However, the flexibility of its modular design introduces substantial researcher degrees of freedom. BERTopic includes five core components (embedding models, dimensionality reduction, clustering algorithms, vectorizers, and term-weighting schemes), each offering several alternatives to the default option. Some applied studies and NLP practitioners have observed that this flexibility can lead to inconsistent topic structures and challenges in reproducibility. This study presents the first systematic investigation of this issue through a multiverse analysis, using the short and long open-ended responses from the COVID Real World Worry Dataset. We applied 162 plausible combinations of embedding models (MiniLM, MPNet, COVID-TweetBERT), dimensionality reduction methods (UMAP, PCA, t-SNE), clustering algorithms (HDBSCAN, K-means), text vectorizers (unigram to trigram), and term-weighting schemes (c-TF-IDF and sublinear variants). Results revealed substantial variability in all modelling outcomes, including the number of topics generated, the proportion of documents classified as outliers, and topic coherence and diversity. Only 5.5% of specifications in the long-text condition met thresholds for both high topic coherence and diversity; none did so in the short-text condition. Pipelines using trigram vectorizers and sublinear term-weighting schemes outperformed default configurations, whereas those with unigram vectorizers and the default c-TF-IDF produced low-quality, overly generic topics. Concerningly, even pipelines with strong coherence and diversity scores produced widely different topic structures (e.g., one yielded two broad themes of anger/concern and emotional impact, while another revealed over ten, including fear of infection, economic worries, and conspiracy beliefs). Our findings caution against treating any single BERTopic pipeline as definitive. Instead, researchers should critically reflect on their modelling choices and, where feasible, adopt multiverse analyses to assess robustness, as BERTopic is increasingly applied in high-stakes contexts such as clinical risk prediction and disaster monitoring.
  20. Local modeler: The next step for data donation?, Niek de Schipper (Universiteit Utrecht), Laura Boeschoten (Universiteit Utrecht), Daniel Oberski (Universiteit Utrecht)
    Abstract: The process of data donation involves participants sharing their digital trace data for academic research. In the data donation framework by Boeschoten, Ausloos, et al. (2022), participants request their Data Download Package (DDP) from platforms such as Google or Meta. Participants then download this data onto their personal devices, where they process it locally to extract only the data points relevant to the research project. After reviewing the extracted data, participants provide informed consent to share the processed data with researchers. However, this process poses challenges, particularly regarding the handling of privacy-sensitive data. Despite local processing and informed consent, even the transfer of processed data to researchers could introduce potential risks to data security and privacy, especially when the processed data is highly sensitive. To address this limitation, we propose a federated learning approach to data donation: local model training on participants’ devices, which eliminates the need for direct data transfer. Instead of donating data, participants train models locally using their personal data and share only the resulting model parameters. This ensures that sensitive data remains private while still enabling researchers to derive meaningful insights using the estimated model parameters. We extend the framework (Boeschoten, Ausloos, et al., 2022) by incorporating this local modeling approach. We demonstrate this extension using Instagram data from participants in the Netherlands. As a case study, we apply Latent Dirichlet Allocation (LDA) to aggregated Instagram data from participants, which results in the clustering of the top 1,000 Instagram accounts in the Netherlands into topics. We compare the local modeling approach to the same analysis performed on the full dataset. Our approach enhances GDPR compliance, reduces data security risks, and builds participants’ trust. While it limits researchers’ ability to perform direct data quality checks, it represents a promising step toward privacy-preserving academic research.
  21. Daily Mobility and the Persistence of Income Segregation in Mixed-Income Areas, Olena Holubowska (KU Leuven University), Anirudh Govind (KU Leuven University)
    Abstract: Income segregation across neighborhoods often reinforces inequalities by shaping access to opportunities and resources. Recently, researchers have increasingly used the lens of activity spaces—the locations people visit in their daily routines—to study how daily mobility mediates peoples’ experience of segregation. Currently, studies typically model a neighborhood’s residents using a hypothetical ‘average’ individual and assume that their income and socioeconomic status can be represented as the neighborhood’s average. However, this approach overlooks the experience of individuals whose personal income differs substantially from the neighborhood’s. These individuals—referred to in this study as income outliers—may have varied mobility patterns and thus, different experiences of segregation. Evaluations of such income outliers remain underexplored in existing analyses. This paper investigates the daily mobility patterns of income outliers to assess the influence of residential income mixing on activity space income segregation. Specifically, we ask whether income outliers, i.e., relatively lower-income individuals living in higher-income neighborhoods and vice-versa, experience segregation differently to their neighbors based on their mobility behaviors? The goal is to understand whether income diversity in residential neighborhoods leads to intergroup exposure in shared urban spaces. We analyze anonymized mobility and individual-level income data from 3,000 individuals in London, England, over a three-month period. By comparing the spatial trajectories of income outliers to those of residents whose income closely matches the neighborhood average, we examine differences in visited locations and exposure to neighborhoods of varying income levels. Our findings suggest that income outliers tend to maintain distinct mobility patterns and are less likely to share activity spaces with their neighbors. This indicates that even within mixed-income residential neighborhoods, individuals may experience greater income segregation. In some cases, activity space segregation appears more pronounced than residential segregation, raising questions about the effectiveness of mixed-income housing policies in addressing segregation.
  22. Mapping geographical connectivity from colleague networks: Associations with regional transmission dynamics of COVID-19 , PingPing Song (Erasmus MC), Luc Coffeng (Erasmus MC), Tom Emery (Erasmus Universiteit Rotterdam), Sake de Vlas (Erasmus MC)
    Abstract: Mathematical models enable understanding of the regional transmission dynamics of novel pathogens like SARS-CoV-2 and support geographically targeted interventions that are effective yet less disruptive. A key concern is how accurately the population mixing patterns are incorporated, as they shape the type and frequency of contacts through which diseases spread. These patterns are commonly synthesized from diary-based surveys in which, however, geographical information is poorly captured due to limited survey size. Among various social settings, workplace connections contribute substantially to disease spread across regions due to the high volume and wide spatial reach. We used administrative data from Statistics Netherlands to construct networks for the 8 million Dutch working population annually from 2019 to 2022. We derived a structured dataset quantifying colleague connections by triplets of locations: two residential municipalities and one workplace municipality. Based on provincial surveillance data of SARS-CoV-2 variant sequences and total cases, we estimated Omicron onset as the first date when Omicron incidence exceeded a predefined threshold. We then investigated the association between regional colleague connectivity and Omicron onset timing using a Bayesian framework. Aggregated at the provincial level by residential locations, all provinces showed high colleague connection densities within province borders. However, the total connection volumes varied substantially, from 3.4 million in Zeeland to 32.2 million in Zuid-Holland. The geographical distribution of these connections also differed. Among inhabitants of Flevoland, 71.2% of colleague connections involved cross-province commuting, while Limburg only had 31.4%. Based on surveillance data, Noord-Holland experienced the earliest Omicron onset. A tenfold increase in within-province connections (living and working in the same province) was associated with 12 days earlier Omicron onset (95% CI: 5 to 20 days). Similarly, a tenfold increase in between-province connections corresponded to 11 days earlier onset (95% CI: 2 to 21 days). Furthermore, provinces with more colleague connections to Noord-Holland, the likely origin of the Omicron variant, showed more synchronized outbreak timing. Our results indicate that regional colleague connectivity is associated with local transmission dynamics, even in highly connected regions like the Netherlands. This underscores the value of incorporating regional connectivity patterns into transmission models to better support policy.
  23. KiM Explorer will be presented at the ODISSEI Marketplace
  24. Language Agnostic Metadata: Why Computational Social Science Needs Multilingual Infrastructure Now, Ramazan Turgut (Directory of Open Access Journals (DOAJ)), Sajad Sepehri (Institute for Social Sciences & Humanities)
    Abstract: Despite increasing commitments to multilingualism and Open Science in Europe, global scholarly infrastructure is still overwhelmingly Anglocentric. Non-English research, especially from the Global South and minority language communities, remains at the margins of discovery, citation, and policy impact. Most metadata systems, from Crossref to ORCID and major citation databases, are not designed for genuine multilingual interoperability. This creates persistent barriers in computational social science and risks reinforcing epistemic inequalities in the digital age. Over the past year, we (Ramazan Turgut and Sajad Sepehri) have been advocating for a new direction: Language Agnostic Knowledge (LAK). LAK is not yet a funded project, but a vision and call to action for the scholarly community. Our goal is to drive development of truly language-neutral metadata workflows, leveraging AI, persistent identifiers, and open standards like Schema.org and BCP47, so that discovery and citation linking can happen across languages and scripts without technical or policy barriers. We have presented these ideas and challenges at the Crossref Metadata Sprint in Madrid (2025) and the PKP OJS Sprint in Oslo (2025), where they sparked broad interest among publishers, infrastructure providers, and researchers. In this talk, we will share practical examples that reveal the limitations of current systems and suggest steps the ODISSEI community can take to build a more inclusive, multilingual infrastructure. Our invitation is simple: join us to pilot new solutions, attract policy attention, and ultimately co-create a pan-European, community-driven project. For computational social science to deliver on its potential, multilingual metadata must become a core component of research infrastructure, not an afterthought.
  25. How AI-related accuracy and confidence shape human reliance on AI-based veracity judgments, Riccardo Loconte (Tilburg University), Merlin Monaro (University of Padova), Bruno Verschuere (University of Amsterdam), Bennett Kleinberg (Tilburg University)
    Abstract: Artificial intelligence (AI) may offer opportunities for deception detection, given that automated verbal deception detection systems are showing higher rates of accuracy than humans. However, the extent to which humans incorporate AI judgments is understudied. This study investigates how information about accuracy and confidence in a fictitious deception detection AI model affects human reliance on AI-based veracity judgments. To this aim, we adopted a 2 (Accuracy: low vs. high) by 5 (Confidence: indecisive, poorly confident, moderately confident, confident, very confident) by 2 (Veracity: truthful vs. deceptive) mixed design. We recruited 373 participants via Prolific who were instructed to perform a traditional lie-detection task for ten statements coupled with judgments from an AI-based lie-detector. Those judgments were manipulated for the AI’s confidence and accuracy according to participants’ conditions. The results of the average absolute deviation from AI judgments across different manipulations were computed using linear mixed models, and our findings are discussed through the lens of the truth-default theory.
  26. Populists and Global Economic Governance: The Turkish Case and the IMF, Saliha Metinsoy (Erasmus Universiteit Rotterdam), Merih Angin (Koc University), Sinan Akgunay (Middle East Technical University)
    Abstract: How do populists and global governance institutions interact? In this paper, we discuss the interaction of the AKP in Turkey and the International Monetary Fund (IMF). The results of the sentiment and the automated text analysis show that the IMF conceals its criticism against the AKP government in a diplomatic language. The AKP leadership on the other hand leans towards a more negative speech. Despite being on the opposite ends of the linguistic scale in terms of sentiment, however, the themes they invoke while challenging one another largely overlap. Particularly, the IMF criticizes the widening current account deficit and external imbalance under the AKP government, the use of selective tax reliefs and pardons, selective distribution to supporters, and the use of monetary policy tools to substitute fiscal discipline. The AKP leadership counters this criticism by arguing that monetary policy difficulties are because of ‘conspiratorial attacks of global forces and power centers’ against the Turkish Lira and claims to defend the ‘national sovereignty’. We argue that the clash boils down to the political economy of the AKP’s power base. The IMF’s criticism directly challenges the trilateral dependency that the AKP built between the business, the electorate, and itself in power. In other words, the main contention between the populists and the global governance institutions in this case seems to be the actual macroeconomic choices and their political consequences beyond the ’empty signifiers’ that the existing studies often attribute to the populists. The project provides not only a theoretical and empirical contribution to the field of studying the dialogical interaction between populists and global institutions but also makes a methodological contribution to the field by using mixed methods, including machine learning in addition to text analysis.
  27. Understanding the roots of health inequality: The impact of social mobility on longevity utilising the intergenerational HSN-SSD, Sara Wiertsema (Wageningen University), Kristina Thompson (Wageningen University), Ruben van Gaalen (Centraal Bureau voor de Statistiek)
    Abstract: Today, people with higher socio-economic positions live longer lives than people with lower positions. There is also evidence that these differences have widened over the twentieth century. The reasons why this has occurred are not precisely known. Previous theoretical research has proposed an untested explanation: that rising intergenerational mobility has led to greater homogeneity within lower socioeconomic groups, concentrating personal characteristics associated with poor health in these lower groups. Our study is the first to test this hypothesis empirically. Gaining a deeper understanding of how and why these inequalities between socioeconomic groups have developed over time can help academics and policymakers address the root causes of poor health across generations and better anticipate future trends in population health. This study utilises the unique linked Historical Sample Netherlands and the system of Social Statistical Datasets (HSN-SSD), an intergenerational sample of Dutch people living and dying between 1883 to 2023. The dataset includes over 250.000 individuals from up to four generations. We investigate how intergenerational trajectories of socioeconomic status and health interact, and how the compositions of the social strata change regarding personal characteristics. We do this by constructing intergenerational mobility typologies to examine whether the distribution of health-related characteristics varies across mobility groups. Subsequently, Cox proportional hazards models will be employed to assess how longevity differ by mobility trajectory, thereby testing the association between mobility-driven group composition and lifespan. This work enriches the evidence base regarding how health inequalities emerge and persist and may help to inform strategies to reduce them.
  28. Green vs. Oil: Regional differences in green and oil energy related ads in Mexico, Sofia Gil-Clavel (Vrije Universiteit)
    Abstract: Research done in Latin America shows that communities’ climate change concern depends on whether oil extraction is already happening in their communities and the economic dependence they have to oil. However, an alternative explanation could come from the type of ads communities are exposed to: pro-oil, pro-green energy, or green washing. This work aims to study the type of environmental and oil related ads that Mexicans are exposed to and how this relates with Mexican states contextual differences. It uses data from the Facebook Library for Ads on Social Issues, Elections, and Politics to study the differential ad exposure broken down by Mexican state. As a method, we use natural language processing and statistical models.Preliminary results show that those states where oil extraction is the main economic activity are the ones where ads tend to mostly be pro-oil. For the rest of the states, we found that the way “energy transition” is portrayed depends on other contextual factors. For example, in Queretaro and Campeche it is mentioned in the context of PEMEX and other oil companies, while in Guanajuato and Oaxaca it is mentioned in the context of climate change and clean energy. These last results are relevant, since in the case of Queretaro and Campeche it may be associated with green washing.
  29. Does staying or moving pay off? Internal migration and early career outcomes of international university graduates in The Netherlands, Steven Kema (Rijksuniversiteit Groningen), Viktor Venhorst (Rijksuniversiteit Groningen)
    Abstract: Our study investigates the relationship between internal mobility and labour market outcomes among international university graduates who remain in the Netherlands after completing their studies. We address the following research question: How are career outcomes (likelihood of obtaining a sufficient quality job) associated with internal migration (interregional residential moving post-graduation) for international university graduates in the host country? To answer our research question, we employ register data from Statistics Netherlands (CBS). International graduates are often viewed as an opportunity to address human capital deficits, with their post-graduation choices regarding residence and employment having substantial implications for both national and regional talent pools, particularly in places facing skill shortages, brain drain, or demographic aging. Understanding the conditions under which these graduates stay and integrate economically into the host country is therefore becoming increasingly important for policy. Yet, the labour market outcomes of international graduates in host countries and the mechanisms that drive these outcomes remain empirically largely unexplored, especially from a regional perspective. It is in particular this regional context and internal geographic mobility that may play a substantial role in shaping the labour market outcomes and long-term retention for this non-standard group of graduates. Our contribution is two-fold. We investigate labour outcomes of international graduates in the host country, which hasn’t been done before empirically, and it gives us insight into how internal migration interfaces with career entry for a group that has specific advantages and limitations (e.g., different personal networks, language barriers, visa restrictions).
  30. Assortativity of COVID-19 vaccination uptake: a population-scale social network analysis in the Netherlands , Taymara Abreu (Universiteit Utrecht), Javier Bernardo-Garcia (Universiteit Utrecht), Mitchell Matthijssen (Vrije Universiteit), Danielle Timmermans (Vrije Universiteit)
    Abstract: Background: Existing research shows that vaccination uptake can cluster within social networks, yet underlying mechanisms are not fully understood. To address this, we examined the tendency for individuals to have institutional or structural ties with others of the same COVID-19 vaccination status (i.e., assortativity) in the Netherlands. Methods: We used social networks data from the Dutch population registry (POPNET, Statistic Netherlands), including over 14 million adults in 2021. Assortativity was measured across five network layers: neighbours, workmates, schoolmates, household members, and family, and across specific family relations. Results: Assortativity was positive in all layers, with the highest value among cohabitants of the same household (0.51), suggesting a moderate tendency for clustering by vaccination status. Among neighbors, schoolmates, and workmates, assortativity was close to neutral (0.07 to 0.08). For the overall family layer, assortativity was 0.18. Subgroup analysis of specific family relations revealed higher assortativity within close family relations, notably partners (0.67) and coparents (0.56), while extended family ties showed lower values (all below 0.16 except siblings-in-law, which was 0.27). Conclusions: These findings show that vaccination status tends to cluster most within households and close family, highlighting the potential role of specific social networks in spreading vaccination behaviour. Public health strategies aiming at increase vaccination uptake in low-coverage groups could benefit from focusing on social networks with higher assortativity, such as specific family relations and cohabitating individuals, where peer influence may be strongest. Ongoing analyses investigate how vaccination uptake assortativity varies across population subgroups and distinct vaccination decision profiles. These are based on individuals’ beliefs, values, and behaviours collected through LISS panel and are extrapolated to the entire Dutch population. Additionally, we also employ survival analyses to examine whether the propensity to vaccinate increases with the proportion of vaccinated individuals in one’s network, and whether this differs by network layer.
  31. MICSM 3.0: An updated structural labour supply microsimulation model for tax-benefit reform analyses in the Netherlands, Thijs Busschots (Centraal Planbureau), Toon VanHeukelom (Centraal Planbureau)
    Abstract: This study presents an update of MICSIM, a discrete-choice microsimulation model for the response in labour supply and formal childcare use of Dutch households. In a random utility model approach, we estimate household preferences for consumption, leisure and formal childcare, differentiating between many household types and heterogeneity along household characteristics. The structural model allows us to simulate household responses to tax-benefit reforms, for which the previous versions of this model have been extensively used. In this update, we use more recent data and implement several improvements to the simulation of net incomes. Comparing simulation outcomes with the old and new method, we are able to quantify which type of reforms have become more/less beneficial over time due to changes in household preferences. One relevant example is the planned major reform in the childcare allowance, strongly reducing its cost for the middle and high income groups. How much is demand for formal childcare expected to increase? And how much additional labour supply does it generated, as per the predictions of the model? Currently, this project is work in progress, meaning that no preliminary conclusions can be included in this abstract.
  32. Interpersonal Influences in Negotiation Behaviours and the Effects of Individual Differences, Tom Nyhoff (Tilburg University), Bennett Kleinberg (Tilburg University)
    Abstract: Negotiations are key to resolving disagreements in daily life, yet little research has focused on their immediate interpersonal dynamics. This study examines the direct influences of verbal negotiation behaviours on the communication and overall negotiation outcomes. To achieve this, this study used the CaSiNo: Campfire Negotiation Corpus, including 1030 simulated virtual chat negotiations completed by 846 participants, together with demographic and personality variables. Using Natural Language Processing, the presence of the two important negotiation behaviours, verbal emotion expression and negotiation strategy, was assessed from the negotiation texts. Afterwards, linear mixed models (emotion expression) and multinomial models (negotiation strategy) were used to analyse if the presence of one negotiation behaviour in a previous utterance influences the extent to which the same negotiation behaviour is also present in the current utterance and if gender or personality moderates this effect. In addition, it was analysed whether changes in negotiation behaviours are associated with the negotiated agreement, opponent liking and satisfaction with the negotiation. Findings show that expressed emotion as well as negotiation strategy have a significant but weak influence on the presence of the same behaviour in the following utterance. While no moderating effect of gender or personality was found for emotion expression, some personality traits seemed to moderate the effect of prior negotiation strategies. No effect of the behavioural changes on the negotiation outcomes was found. Findings highlight the complexity of (digital) negotiation interactions and the need for nuanced models that account for contextual and relational variables beyond static predictors.
  33. How Biased Is the Voice of People? Introducing The Political Voice Inequality Index (The R-Index), Vardan Barsegyan (WODC Research and Data Centre)
    Abstract: The paper presents a new measure of political voice inequality: the R-index. Contemporary democracies operate on the fundamental premise that everyone’s voice counts and carries equal weight. However, research shows that political voice often skews toward groups such as more educated males without a migration background. The extent to which individual characteristics collectively influence political participation defines political inequality. The R-index captures the overall predictive power of individual characteristics within a society, serving as a diagnostic measure of political voice inequality. Among various ways to measure this predictive power, I use the adjusted coefficient of determination – the adjusted R-squared – as the basis for the R-index. I demonstrate that the R-index possesses desirable properties for an inequality measure and correlates with theoretically relevant country-level indicators, including democracy, freedom, and economic development indexes. I show, for example, that more democracy does not automatically lead to more equality in political voice. I also show that political voice inequality is a multidimensional concept that reflects variability across different forms of political participation. To validate the index, I conduct two studies. In Study 1, I calculate the R-index by regressing different forms of political participation on ten key individual characteristics (e.g., age, sex, income, parental SES). I analyze 272 country-years from the ESS survey (2002–2023). In Study 2, I expand the dataset to over 700 country-years by including additional sources – ISSP, ANES, EVS, and WVS. I aim for this index to become a central tool in comparative research on political voice inequality across time and societies.
  34. Migration aspirations and crises in redditors’ trajectories: a case of data extraction using local large language models, Vestin Hategekimana (University of Geneva)
    Abstract: In the contemporary landscape of global mobility, understanding how personal life trajectories are shaped by migration decisions has become increasingly crucial (Akhila Chakshu & Rituparna, 2024; Soto Nishimura & Czaika, 2024). This research focuses on exploring how life-course events impact the migration intention of individuals who express a desire to move to Switzerland, using mixed-method longitudinal analysis and leveraging data from Reddit as a primary source and large language models as data extraction strategy. Using the subreddit r/AskSwitzerland, I aim to capture nuanced narratives that might otherwise be overlooked in traditional survey-based studies. Reddit provides an invaluable platform for longitudinal analysis, enabling researchers to track individuals’ migration journeys over extended periods and offering insights into how personal circumstances evolve alongside shifting social, economic, and political contexts. A particularly significant focus is the impact of the COVID-19 pandemic as a major crisis event. The non-official API (PullPush) facilitates the collection of the study’s dataset spanning from January 2018 to March 2022. By collecting and analyzing user profiles and posts from two years before and after the declared relocation decision (until March 2020), I aim to map out the temporal progression of migration intentions. One significant innovation is the use of large language models for sophisticated data extraction which allows researchers to infer demographic information and life-course events from users’ posts. This mixed-method longitudinal approach is divided into two phases. First, a quantitative analysis using a multilevel model will be conducted to identify the factors influencing the decision to migrate initially, and subsequently, whether an event like the COVID-19 pandemic has an impact on the willingness to migrate. Second, a qualitative analysis will be performed on individuals whose intention to migrate was halted by COVID-19 to gain further insights into their reasons.
  35. Measuring in-betweenness: a method to identify interplaces in Belgium and The Netherlands, Wander Demuynck (Universiteit Utrecht), Michiel Van Meeteren (Universiteit Utrecht), Evert Meijers (Universiteit Utrecht), Ben Derudder (KU Leuven)
    Abstract: A significant conceptual blind spot in current models of metropolitan regions lies in the treatment of interplaces – secondary localities situated between the conventional boundaries of the main cities and their suburbs that nevertheless play a pivotal role in shaping metropolitan economic, political, and cultural dynamics. Despite their importance, there is little consensus on how to define interplaces conceptually or how to identify them empirically within a given urban system. This paper addresses these gaps by first unpacking the multifaceted nature of interplaces’ “”in-betweenness””: as places characteristically between city and suburb, as nodes within broader urban networks, as intermediaries across inter- and intra-regional scales, and as localities straddling jurisdictional territories. Second, we propose a novel empirical method grounded in network-analytical techniques – specifically community detection – to identify interplaces in Belgium and The Netherlands. Drawing on longitudinal commuting datasets, the results show an increase in the number and diversity of interplaces in both settings – especially in their respective metropolitan areas (the Flemish Diamond and Randstad). These findings are robust and partially align with prior empirical efforts in both contexts, but also reveal new insights, underscoring the utility of our more systematic and comprehensive approach to interplace-identification.
  36. Disadvantage Clustering: A novel measure of relational inequality using machine-learning, Yohei Yoshizawa (King’s College London)
    Abstract: This paper proposes an innovative measure of relational inequality, which quantifies the extent to which the most disadvantaged social group overlaps across domains of well-being (e.g., income, education, health). To this end, I identify the most disadvantaged group in each domain by taking advantage of the machine-learning (ML) methods of inequality of opportunity (IOp) estimation developed recently. In addition to equality of opportunity, Wolff (2015) advocates for the de-clustering of disadvantages from the perspective of relational equality, that is, the idea that a just society is one in which individuals relate to each other as equals. He argues that, while the most disadvantaged group will inevitably exist so long as we compare groups of people based on a single metric (e.g., employment, education, health), it is morally objectionable when that group is constant across domains. In such cases, multidimensional disadvantages are clustered to a single social group. This suggests that there is a recognizable hierarchy in society. The IOp literature has recently adopted ML algorithms to take a data-driven approach to model-specification (Brunori & Ferreira., 2024). The by-product of this is the identification of the most disadvantaged group based on the decision-tree generated by the ML process. I take advantage of this hitherto overlooked merit of ML-IO to operationalize one of the distributive demands of relational inequality proposed by Wolf. Specifically, I apply the transformation-tree version (Brunori et al., 2024) to the 1970 British Cohort Study regarding income, employment, educational attainment, physical and mental health, and loneliness. This study’s contribution is twofold. First, this will be among the first to measure relational inequality (Segal, 2022). Second, it will be an innovative method for measuring multidimensional deprivation, which, crucially, identifies the group of people that social policies should target, unlike existing ones (e.g., García-Gómez et al., 2024; Alkire & Foster, 2011).
  37. The Social Networks of Chronic Disease, Zixuan Wang (Universiteit Utrecht), Erika Banzato (University of Padua), Solveig Cunningham (Nederlands Interdisciplinair Demografisch Instituut)
    Abstract: Understanding how these diseases cluster is important for prevention and treatment; for example, which diseases cluster, and why; and whether these patterns differ by age, sex, or social groups. In this study, we apply social network analysis to bring new insights into the clustering and progression of multimorbidity. This method has been used to identify relationships among individuals, for example, identifying who is connected with many or few individuals, who are more closely connected, and who is the central. We propose that these methods can also tell us about how chronic diseases relate to each other. We use population-representative multi-country data from wave 9 of the Survey of Health, Ageing and Retirement in Europe (SHARE) to map the networks of chronic diseases in the European 50y+ population. We examine self-reported diagnosis of heart attack, hypertension, stroke, diabetes, chronic lung disease, arthritis, cancer, respiratory diseases, stomach ulcers, Parkinson’s disease, cataracts, Alzheimer’s disease, rheumatism, osteoporosis, chronic kidney disease, fractures, obesity, and affective disorders. We measure number of disease communities; closeness centrality, reflecting distance to all other diseases; betweenness centrality, indicating how a disease connects different parts of the network. Then we conducted stratified analyses by age and sex to examine whether clustering patterns are consistent across demographic groups. Furthermore, we applied a graphical model to identify conditionally independent diseases. We identified three disease clusters using the Louvain method, with a modest modularity level of 0.114: a cardiometabolic cluster including five diseases, such as heart conditions, hypertension, and diabetes; a neurocognitive cluster including three diseases, such as Parkinson’s disease and stroke; and a systemic cluster including nine diseases, such as cancer and conditions related to the digestive, respiratory, and bone health. The diseases with the highest levels of both closeness and betweenness centrality were Parkinson’s disease (1st), obesity (2nd), and cancer (3rd). Across age groups and sexes, the patterns remain the same, except in women and adults aged 50–74 the Neurocognitive and Systemic clusters tend to merge into one cluster.
  38. You Don’t Have to Live Next to Me: Towards Demobilizing Individualistic Bias in Computational Approaches to Urban Segregation, Anastassia Vybornova (University of Copenhagen), Trivik Verma (University of Bristol)
    Abstract: The global surge in social inequalities is one of the pressing issues of our times. The spatial expression of social inequalities at city scale gives rise to urban segregation, a common phenomenon across different local and cultural contexts. The increasing popularity of Big Data and computational models has inspired a growing number of computational social science studies that analyze, evaluate, and issue policy recommendations for urban segregation. Today’s wealth in information and computational power could inform urban planning for equity. However, as we show here, segregation research is epistemologically interdependent with prevalent economic theories which overfocus on individual responsibility while neglecting systemic processes. This individualistic bias is also engrained in computational models of urban segregation. Through several contemporary examples of how Big Data — and the assumptions underlying its usage — influence (de)segregation patterns and policies, our essay tells a cautionary tale. We highlight how a lack of consideration for data ethics can lead to the creation of computational models that have a real-life, further marginalizing impact on disadvantaged groups. With this essay, our aim is to develop a better discernment of the pitfalls and potentials of computational approaches to urban segregation, thereby fostering a conscious focus on systemic thinking about urban inequalities. We suggest setting an agenda for research and collective action that is directed at demobilizing individualistic bias, informing our thinking about urban segregation, but also more broadly our efforts to create sustainable cities and communities.
  39. Control features in app-based data collection: The effects and best implementation to give participants agency over the data collection process, Thijs Carrière (Universiteit Utrecht), Bella Struminskaya (Universiteit Utrecht), Laura Boeschoten (Universiteit Utrecht)
    Abstract: In social and behavioral sciences, researchers increasingly collect their data with mobile apps and sensors, as such data collection methods can obtain more granular and precise measurements. However, app-based data collection generally deals with low study participation rates, for example caused by privacy concerns. Studies show that providing participants with control over the data collection process can mitigate privacy concerns and increase participation rates. However, the mechanisms behind control increasing participation rates and what ways to best implement control features in app-based data collection are yet to be understood. Our study addresses the question what app-features provide participants with control over the data collection process.To understand what features of apps influence willingness to participate, we conducted a vignette study in December 2024 in the Centerpanel, a former-probability, online panel in the Netherlands. Participants were presented with a hypothetical study that uses app-based data collection. The study was introduced with both a text and a mockup of the study app. As an experiment, we varied the presentation of the study randomly over two conditions. In their description and mock-up, the experimental group received control features, such as the option for data deletion, data overview, and pausing data collection. The control group received a description and mockup without any of these features. We investigate whether control leads to higher willingness to share data, and through which mechanisms. We will test the leverage-saliency theory to see whether effects of control are larger for people that find control important. Furthermore, we investigate what control features app users rate as more important in their decision to participate in app-based data collection. The results of this study advance the understanding of control features in data collection and might inform researchers when designing study apps.
  40. Bayesian Penalisation and Variable Selection in Relational Event Model, Aliasghar Rostami Charati (Universiteit Utrecht), Mirjam Moerbeek (Universiteit Utrecht), Sara van Erp (Universiteit Utrecht), Mahdi Shafiee Kamalabad (Universiteit Utrecht)
    Abstract: Addressing the limitations of classical MLE-based REMs—such as overfitting, computational burden, and instability—in high-dimensional settings highlights the need for more scalable, interpretable, and statistically robust modeling approaches. In response to this need, Bayesian penalization and variable selection in Relational Event Models (REMs) are explored, beginning with simulated data and later extended to real-world datasets. Maximum likelihood estimates obtained via the remstimate and glm functions are compared with Bayesian estimates derived from the brm package under both noninformative and informative normal priors. Variations in prior specifications—such as location, scale, and degrees of freedom in priors like the Student-t—are examined to assess their influence on model inference. Furthermore, Bayesian methods are systematically compared to classical regularization techniques, including LASSO and elastic net, in terms of variable selection accuracy, predictive performance, and computational efficiency. Practical guidance is also provided on the implementation of REMs using the remstats, remulate, and brm packages.
  41. Visualizing political debates – extracting causal pathways with LLMs, Anna Machens (Universiteit Twente), Karel Kroeze (Universiteit Twente)
    Abstract: This project analyzes parliamentary debates to identify how political actors construct problems, explain their causes, and propose solutions. Rather than focusing on rhetorical interaction, it maps causal arguments—showing how issues are framed and which political parties support specific causal claims. Understanding these causal narratives is crucial, as they influence how policies are justified and which solutions are seen as legitimate. The analysis offers a value-neutral representation of problem framings, allowing for comparison across ideological lines without judging the validity of the claims. The project uses large language models (LLMs) to help extract causal relations from political text. This allows to compare several debates over time and between countries (Germany, Netherland). The extracted problem–cause–solution chains will be associated with their proponents’ party positions. They can then be interpreted in light of established frameworks such as issue framing or the Advocacy Coalition Framework. The result is a structured, interpretable map of political reasoning in debates. It contributes to research on political discourse and policy processes, and supports democratic transparency by clarifying how different parties construct public problems and argue for specific actions.
  42. A Formal Cognitive-Behavioral Model of Major Depressive Disorder, Jari Zegers (Tilburg University), Jelte Wicherts (Tilburg University)
    Abstract: The present study aims to develop a formal theory of Major Depressive Disorder (MDD) by integrating cognitive-behavioral and network theories of depression within the Productive Explanation Framework. Psychological theories are predominantly verbal, which introduces ambiguity in testing and refining theories. Formal theories, expressed through mathematical formulas or computer code, provide precise predictions, reducing this ambiguity. We develop a formal model from a dynamic systems perspective defining the change in four model components (mood, withdrawal behavior, future expectations and environmental influences on mood) as a set of differential equations. We used the proposed formal model of MDD to simulate data to assess whether key phenomena associated with depression can be replicated, including the prevalence of MDD, spontaneous remission, recurrence of depressive episodes, the influence of stressful life events, and the treatment effect of cognitive behavior therapy. We find the formal theory to adequately reproduce the statistical patterns representing five key depression-related phenomena. As such, we consider the model a first step towards formalization of cognitive-behavioral theories of depression. Future work may investigate measurement of model components, parameter selection, and the possibility to account for other depression related phenomena.
  43. Real-Time Insights into AI Adoption via Corporate Web Presence, Ana Pastor-Merino (Universitat Politècnica de València), Josep Domenech (Universitat Politècnica de València), Xavier Martínez-Barbero (Universitat Politècnica de València)
    Abstract: Artificial intelligence (AI) is rapidly reshaping Spanish industry, yet traditional data sources—surveys, patents, annual reports—struggle to keep pace with firms’ day-to-day technology deployment. This study introduces a real-time monitoring framework that mines corporate websites with Large Language Models (LLMs) to trace the diffusion of AI across more than 118 000 Spanish companies in 2023 and 2025. Public webpages are automatically harvested, segmented, and vectorised with contextual embeddings to surface AI-related text. Candidate fragments are evaluated by a GPT-based classifier that flags whether a firm employs AI and, if so, distinguishes internal process use from external product offerings, while tagging six functional domains such as marketing, production and logistics. A manually labelled, stratified sample of 228 firms benchmarks the model and calibrates confidence thresholds. The resulting dataset reveals a steady expansion of AI discourse: overall prevalence rises from 3.5 % to 4.8 % of firms, with clear differences by size, sector and region. Large enterprises and information-and-communication businesses lead adoption, whereas micro-firms and traditional sectors lag behind. Internally, marketing, administration and production tasks dominate AI usage; externally, two-thirds of adopters embed AI in products or services, often describing these innovations as radical. By transforming web narratives into structured indicators, the framework delivers timely, fine-grained evidence for policymakers, investors and scholars, underscoring the value of corporate websites as a scalable, low-cost data source for tracking emerging technologies.
  44. Scaffolding Self-Regulated Learning with a Conversational Agent: A Framework and LLM-Based Pipeline for Scalable, Adaptive Feedback in Higher Education, Gabrielle Martins van Jaarsveld (Erasmus Universiteit Rotterdam), Qixiang Fang (Universiteit Utrecht), Erik-Jan van Kesteren (Universiteit Utrecht)
    Abstract: As higher education increasingly adopts technology-enhanced learning environments, students are expected to become self-directed learners. However, many struggle with the self-regulated learning (SRL) skills needed to succeed in these autonomous contexts. To address this, we developed a conversational agent (CA) that supports students in goal setting, monitoring, and reflection—core components of SRL. The CA intervention provides scalable, personalized support by guiding students through weekly goal-related activities. Our project makes three primary contributions. First, we proposed a theoretical framework for quantifying 10 different indicators of SRL processes exhibited in students (e.g., goal specificity, goal measurability, and realistic multisource planning). Based on this framework, we developed a codebook that can be used to score these 10 indicators based on conversations between students and the CA. We also validated the validity and reliability of this codebook manually. Second, we implemented a Python-based pipeline to achieve automatic high-quality annotations of these indicators of SRL processes using large language models and prompt engineering techniques. We applied this pipeline to annotating rich conversation data between 100 students and the CA over five weeks, and demonstrated its potential for providing real-time identification of SRL processes and adaptive support strategies tailored to individual needs. Lastly, we analyzed the relationships between the CA interventions, different indicators of SRL processes, and student achievement outcomes. We hope our work not only enhances understanding of student engagement with SRL tools but also lays the groundwork for implementing personalized, scalable feedback mechanisms in digital learning environments.
  45. Comparing FAIR Assessment Tools and their Alignment with FAIR Implementation Profiles using Digital Humanities Datasets, Shuai Wang (Vrije Universiteit), Andre Valdestilhas (Vrije Universiteit), Menzo Windhouwer (KNAW), Ronald Siebes (Vrije Universiteit)
    Abstract: FAIR principles serve as guidelines for implementing data and metadata to improve Findability, Accessibility, Interoperability, and Reusability. In recent years, numerous tools have been developed to assess how well datasets adhere to each FAIR principle. However, due to their diverse designs, these tools interpret the FAIR principles differently and provide varying assessment results, which can be confusing. Many communities publish datasets that follow similar data management practices, and some of these common practices have recently been compiled into community standards known as FAIR Implementation Profiles (FIPs). This paper compares the metrics of FAIR assessment tools with FIP. We illustrate these differences by analyzing the assessment results of two datasets in the Digital Humanities domain and further explore how these results compare with their corresponding FIPs. More specifically, they correspond to the ODISSEI portal community and the CLARIN community, respectively.
  46. Enhancing machine learning-assisted systematic reviews with state-of-the-art data balancing strategies, Qixiang Fang (Universiteit Utrecht), Timo van der Kuil (Universiteit Utrecht)
    Abstract: Machine learning-assisted systematic reviews (ML-SRs) have shown great potential to reduce workload during evidence screening. However, the inherent class imbalance—where relevant studies form a small minority—poses a major challenge, often leading to poor recall of crucial evidence. This study investigates the impact of state-of-the-art data balancing strategies on the performance of ML classifiers in systematic reviews. We evaluate a range of techniques, including undersampling, oversampling (e.g., SMOTE), and ensemble-based methods, across multiple real-world systematic review datasets. Classifiers are assessed using metrics sensitive to imbalance, such as recall, precision-recall AUC, and work saved over sampling. Results show that while traditional methods often struggle with the extreme imbalance ratios common in SRs, tailored strategies such as dynamic resampling and cost-sensitive learning significantly improve recall without excessively compromising precision. We further demonstrate that integrating these strategies into active learning pipelines enhances prioritization of relevant studies in early screening rounds. Our findings underscore the importance of addressing imbalance not only for classifier performance but also for the practical utility of ML-SRs. This work offers actionable insights for researchers seeking to make ML-driven evidence synthesis more effective, reliable, and suitable for high-stakes decision-making in health, policy, and other domains.
  47. Mapping the Complete Network of Sweden, Károly Takács et al. (Linköping University)
    Abstract: We map population-wide networks that integrate multiple social contexts (i.e. multiplex networks) by utilizing information from the Swedish population register data. We study the properties of the population-wide exposure networks, how they change over time, and the consequences they have for various social outcomes. Methodologically, we develop novel approaches for constructing, measuring, and analyzing full-population networks. Substantively, we provide novel insights into how interwoven and segregated multiplex exposures are, how influence between domains affect individual outcomes, and how individuals’ exposure networks develop from birth to death through the life course. 

New data infrastructures

Room: Mission 1

Chair: Marcel Das

Building a research app for data collection using smartphone sensing for the social and behavioral sciences, Bella Struminskaya (Universiteit Utrecht), Thijs Carriere (Universiteit Utrecht), Niek De Schipper (Universiteit Utrecht), Laura Boeschoten (Universiteit Utrecht)

Abstract: The data from smartphone sensors and wearables have the potential to transform social and behavioral science research by ensuring in-the-moment, longitudinal, rich, precise, and scalable data collection. Researchers are increasingly taking advantage of such technologies using smartphone research apps with a wide range of sensing functionalities, including ecological momentary assessment, geolocation sensing, physical activity sensing, app usage, and others. Numerous apps and platforms exist that researchers can use for data collection. They vary in functionalities, needed IT know-how, and possibilities for re-use. In this presentation, we provide a comprehensive review of existing apps and platforms used for app-based data collection in social and behavioral sciences and official statistics based on a scoping review of existing platforms for app and sensor-based data collection. Based on this, we discuss methodological, operational and software-development considerations for mobile sensing research infrastructures. Secondly, we introduce an open-source smartphone app research infrastructure under development that will be available to social science researchers in the Netherlands and being developed by the authors with the support of the Dutch Research Council (NWO). We present the conceptualization of data collection modules (geolocation, EMA, physical activity, smartphone app use, etc.), features of providing participant feedback and engaging participants, back-end organization for users (i.e., researchers) and features that allow conducting randomized experiments for users. We share our experiences of software development and ensuring how the infrastructure can be developed in a sustainable manner and be integrated into national and international social science research landscapes. We focus on integrating research-based methodological aspects of privacy protection and participant engagement in software (i.e., research app) development, and discuss more general aspects of maintenance, integration with other parts of the research landscape, sustainable funding, as well as legal and ethical considerations when creating socio-technological systems for data collection.

Laura Boeschoten (Universiteit Utrecht), Niek De Schipper (Universiteit Utrecht), Theo Araujo (Universiteit van Amsterdam)

Abstract: The Digital Data Donation Infrastructure (D3I) is a collaborative initiative among six Dutch universities. The primary aim of D3I is to facilitate data donation for research in the social sciences and humanities in the Netherlands. D3I enables individuals to securely and transparently donate their digital trace data for academic research purposes by making use of their rights granted under the General Data Protection Regulation (GDPR). D3I empowers researchers to work directly with individuals to collect data, thereby bypassing the need for platform-mediated data access. This opens new avenues for studying human behaviour and platform dynamics.
To facilitate that researchers collect digital trace data in a safe, ethical and efficient way, we have developed a comprehensive research infrastructure that enables researchers to design a complete data donation study, considering consent, local processing of the digital traces, integration with other data collection approaches, participants invitations and data storage, while also providing recommendations regarding methodological and ethical aspects.
This infrastructure has been in place since 2023 and has been used by more than 20 research projects in and outside the Netherlands. The infrastructure continues to be expanded, improved and further embedded within the wider Dutch research infrastructure landscape for social sciences. In this presentation, I will first explain the sustainable development model that we use for this infrastructure, that facilitates the continuous development and improvement through multiple funders. Second, I will showcase how researchers both inside and outside the Netherlands interested in collecting digital trace data for their research purposes can use this infrastructure, where I cover topics such as participant recruitment, informed consent, data storage and re-use of data donation datasets.

A historical disease database of the Netherlands and based on newspapers using natural language processing techniques, Kristina Thompson (Wageningen University), Qixiang Fang (Universiteit Utrecht), Erik-Jan van Kesteren (Universiteit Utrecht)

The nineteenth and early twentieth century Netherlands had a very high burden of infectious diseases. Outbreaks of diseases such as tuberculosis and smallpox were common. To study health and living standards in this period, capturing the disease burden is critical. While today we are able to measure the disease burden with infection rates, this is not possible for the past. In the Netherlands, there are no available measures of the disease burden prior to the mid-twentieth century, aside from mortality rates. Mortality rates are also not perfect proxies of disease: they fail to capture diseases that negatively impacted an individual throughout their life, but did not kill them. This prevents researchers from, for instance, studying how exposure to infectious diseases in early life might impact someone’s health in later life. An indicator of the infectious disease burden that is not based on mortality information would enable researchers to do so. In the nineteenth and twentieth centuries, Dutch newspapers reported on the number of new cases and deaths from specific infectious diseases. A large share of these newspapers are digitized and publicly available via the Royal Library’s Delpher.nl. This project aims to convert these newspaper mentions into a quantitative indicator of the disease burden. To do so, we leverage natural language processing methods to: process newspaper mentions of a disease, connect them to a geographical region, map them onto numeric scales, and assess their performance as a proxy of the disease burden in a given municipality and year. We have thus released an open-source disease database and several historical maps. Further, we show how this disease database can be used by researchers, and how to account for measurement uncertainties in the disease database.


Workshop – Working with sensitive data in SANE

Room: Mission 2

Workshop leader: Tom Emery

Abstract: During this session, datasets that are available via SANE, such as LISS and FIRMBACKBONE, will be demonstrated, including what type of linkage and enrichment can be done with this data in a secure environment. You will also be taken through the data application process and how to launch a SANE environment on your machine, request additional memory or compute, and ensure that your SANE project is well documented and reproducible.  You will need to bring your own laptop to this session.

Networks, Contexts, and Outcomes

Room: Progress

Chair: Yuliia Kazmina

Using Prediction Gaps and Registry Data to Understand the Unequal Influence of Social Contexts on Educational Outcomes, Javier Garcia-Bernardo (Universiteit Utrecht), Eva Jaspers (Universiteit Utrecht), Weverthon Machado (Universiteit Utrecht), Samuel Plach (Boconni University)

Abstract: Social contexts—such as families, schools, and neighborhoods—shape life outcomes. The key question is not simply whether they matter, but rather for whom and under what conditions. Here, we argue that prediction gaps—differences in predictive performance between statistical models of varying complexity—offer a lens for identifying surprising empirical patterns (i.e., not captured by simpler models), highlighting where theories succeed or fall short (interactions or complex). Using population-scale administrative data from the Netherlands, we compare logistic regression, gradient boosting, and graph neural networks to predict university completion from early-life contexts. Overall, prediction gaps are small, suggesting that previously identified indicators, particularly parental socioeconomic status, capture most of the measurable variation in educational attainment. However, gaps are larger for disadvantaged students, indicating that the effects of social context for these groups go beyond simple models in line with sociological theory. Our paper shows the potential of prediction methods to support sociological explanation.

Simulating the influence of social networks and food environments on the formation of health-taste beliefs, Jiri Kaan (Wageningen University), Yara Khaluf (Wageningen University), Sonja Kunz (University of Vienna), Kristina Thompson (Wageningen University)

Abstract: Health-taste beliefs – the belief whether healthy eating requires compromising on taste – can significantly influence food choices and is found to vary by socioeconomic status (SES). Lower SES individuals tend to believe more frequently that healthy foods are less tasty, potentially contributing to health inequalities. Yet, the mechanisms driving these SES differences in health-taste beliefs remain poorly understood. Based on previous work, we propose that two factors shape health-taste beliefs differently across SES groups: social network influences (health-taste beliefs of close ties) and food environment exposures (skewed availability of unhealthy and tasty foods). To test this, we developed an agent-based model using data from 1046 participants in the LISS panel from the Netherlands. Agents preferentially connected in a scale-free network based on their socioeconomic status and were initialized with observed sociodemographic and dietary behaviors. Over time, agents updated their health-taste beliefs through simulated social network interactions and food environment exposures. The simulated health-taste beliefs were validated against the health-taste beliefs observed in the LISS panel dataset. Our simulations suggest that lower SES agents were more susceptible to social network influences compared to higher SES agents. This differential susceptibility suggests that social network structure and dynamics may perpetuate differences in health-taste beliefs across SES. These findings have significant implications for how health interventions are designed. Traditional interventions may attempt to directly change lower SES groups’ health-taste beliefs via education; however, our model suggests that targeting social networks may be more effective.

Community-level networks on a societal scale, Rense Corten (Universiteit Utrecht)

Abstract: The emergence of online social networks like Facebook in the early 2000’s promised a breakthrough in social networks research social network research by enabling analysis of societal-scale interactions due to abundant data. However, despite many groundbreaking studies, progress has been limited by a lack of freely available data. In rare cases where such data have been made available by platforms to researchers, individual-level data can typically not be shared with the wider research community. However, data that are aggregated to higher social entities, such as municipalities, can be shared more easily. This paper presents one such data set based on the (now defunct) Dutch social network platform Hyves. From an underlying individual-level network covering a significant fraction of the population, we create a data set consisting of topological features of within-municipality networks, covering all municipalities in the Netherlands. This provides a unique insight into features of social connectivity within municipalities that is not readily available from other resources. We present descriptives of topological features of municipality networks, explore associations with other properties of municipalities, and demonstrate the usefulness of municipality-level network measures for social research.

Talking Across Divides: Online and Offline Perspectives

Room: Quest

Chair: Tim Reeskens

Understanding Controversiality in Online Discourse: A Multi-Method Analysis of the DerStandard, Elena Candellone (Universiteit Utrecht), Emma Fraxanet (Universitat Pompeu Fabra), Darja Cvetkovic (Institute of Physics Belgrade), Simon D. Lindner (Complexity Science Hub)

Abstract: Controversial discussions drive online discourse by attracting user engagement and shaping public opinion. Depending on how these conversations are framed and moderated, they can either support healthy democratic debate or increase toxicity and polarization. Understanding and quantifying controversiality is crucial for identifying sources of disagreement, influential narratives, and antagonistic dynamics. Previous studies have examined controversy on online platforms from various angles, such as partitioning user interaction networks. However, to our knowledge, no approach integrates ideological positioning, network structures, and thread dynamics. To address this, we analyze a decade-long dataset from the Austrian newspaper platform DerStandard, comprising over 75 million user comments and 412 million up- and down-votes. This German-language corpus offers a rare opportunity to study controversy at scale beyond English-language platforms, with structured user interactions, topic context, voting behavior, and threaded discussions. Our work has two main components. First, we examine whether latent ideologies can be extracted from user interactions with selected influential users—identified based on positive, negative, and controversial vote patterns. We define controversiality using a metric that favors a balanced volume of positive and negative votes. Correspondence analysis (CA) provides user embeddings, and we validate influencer classifications using large language models (LLMs). The network structure reveals tightly connected clusters with internal alignment and polarized cross-cluster interactions. Second, we assess thread-level controversy through characteristics like thread depth, number of posts, negative votes, and a controversy index inspired by the h-index. We also examine response times, the original author’s controversiality score, and polarization based on ideological stance. Preliminary results show that structural thread measures strongly correlate, while vote-based metrics capture distinct facets of controversy. Our work highlights how controversy manifests in online communities, contributing insights to debates on polarization and content moderation.

The construction of territorial stigma: mapping representations of the Bijlmer, Gijs Custers (Erasmus University Rotterdam)

Abstract: Newspaper media play a dominant role in (re)producing and circulating negative representations of neighborhoods. This study will focus on the Bijlmer, one of the most well-known and ‘notorious’ areas in the Netherlands. The aim of this study is to show how the Bijlmer is portrayed in newspaper media, how these representations change across time and what might explain these changes. A large corpus is build including articles on the Bijlmer from the five largest newspapers between 1990-2024. Computational methods like topic modelling and sentiment analysis are used to analyze this corpus. This presentation will focus on the work in progress, outlining some of the potential benefits and challenges relating to corpus building and textual analysis. The study aims to unravel how and why areas become (de)stigmatized.

Analyzing Country-Level Heterogeneity in Public Attitudes Toward Migration Topics on Bluesky Jisu Kim (Universiteit Utrecht), Jiho Kwak (Seoul National University), Egor Kotov (Max Planck Institute for Demographic Research), Tom Theile (Max Planck Institute for Demographic Research)

Abstract: Migration is a globally significant political issue, with discourse varying across countries. This study examines the regionally varying patterns of migration discourse on Bluesky, an emerging platform that saw an increase in number of users following Elon Musk’s acquisition of Twitter (now X) and its subsequent policy changes. We explore how Bluesky users discuss migration issues in the United States, the United Kingdom, and Germany, focusing on key topics, sentiment, and variations in word usage. To achieve this, we identify country-specific themes, sentiment dynamics, and construct dialectograms using embeddings trained on posts from each country to explore nuanced linguistic differences. During this process, since Bluesky lacks geolocation metadata, we employ a Large Language Model (LLM) to annotate posts based on the country they reference. Our findings reveal distinct themes shaping Bluesky users’ migration discourse in the United States, the United Kingdom, and Germany. In all countries, negative sentiment posts outnumber positive sentiment posts approximately five times, and the predominant themes of negative sentiment posts were more partisan compared to those of positive sentiment posts. A bilateral comparison using dialectograms highlights words co-occur distinctively with words related to recently relevant subtopics in each corpus. This study provides new insights into regional variations in migration discourse on Bluesky and introduces an innovative approach for analyzing social data sources that lack traditional metadata.”

Longitudinal and spatial perspectives on educational inequalities

Room: Royal Lobby

Chair: Christoph Janietz

How Neighborhood and School Contexts Jointly Shape Primary School Outcomes in Dutch Cities, Dieuwke Zwier (European University Institute), Herman van de Werfhorst (European University Institute), Carla Haelermans (Maastricht University)

Abstract: Neighborhoods and schools are both seen as important extrafamilial social contexts shaping children’s education: advantaged environments offer greater access to various resources, whereas exposure to concentrated disadvantage can limit life chances. A pressing question is how neighborhood and school contexts are interconnected, and jointly shape educational outcomes. Residential and school choices are intertwined: some (prospective) parents choose where to live to access desirable schools, or choose schools depending on where they live (e.g., attending a local school or “opting out”). Subsequently, (the degree of) these interconnections may shape contextual effects. Yet, despite theoretical recognition of these links, empirical research on school and neighborhood effects has largely developed separately. This paper examines how neighborhood and schools jointly shape (inequality in) educational outcomes in the Netherlands. Using geo-coded longitudinal register data from the four largest Dutch cities (Amsterdam, Rotterdam, The Hague, and Utrecht), we study how families sort into neighborhood-school structures and how this shapes children’s track recommendations in the transition to secondary school – a consequential moment in the Dutch system. We use a recently developed two-step approach that explicitly accounts for contextual selection processes, and compare multiple urban regions to assess variation across socio-spatial contexts. Results will be presented at the conference.

Trends in inequality of opportunity: studying heritability of education over time, Marjolijn Das (Centraal Bureau voor de Statistiek), Mayke Nollet (Vrije Universiteit)

Abstract: Inequality of educational opportunity and its trend over time is one of the most prominent policy concerns in Western countries. Inequality of educational opportunity can be defined as a situation where children with similar innate capabilities have different educational outcomes as a result of their different (socio-economic) circumstances, which contributes to social stratification. Intergenerational transmission of education is often used as an indicator of inequality of opportunity. However, shared genes are a major source of resemblance between parents and children, and studies on intergenerational transmission are unable to distinguish between genetic effects and the influence of the family environment. Twin studies are designed to parse out the effects of genes versus the shared environment by making use of the difference in genetic relatedness of monozygotic versus dizygotic twins. Twin studies are usually based on survey data. The current study uses population-wide data on twins and siblings of Dutch origin from integral registers of Statistics Netherlands, birth cohorts 1964-1997. These integral data allow us to assess detailed trends over in genetic influence (‘heritability’) and the influence of the shared environment on educational attainment in the Netherlands. We adapt the classical twin design to estimate heritability in the absence of information on zygosity. Over these thirty-four years, heritability of educational attainment increased and the influence of the shared environment decreased. Hence, educational attainment has become increasingly determined by innate capabilities, which strongly suggests a decrease in inequality of opportunity in these birth cohorts of Dutch origin. This study contributes to insight in contemporary patterns of stratification. Future research should focus on younger cohorts and non-Dutch origin groups.

School Competitiveness And The Development of Motivation: A Longitudinal Cohort Study On Social Inequalities In Secondary Education, Nathalie Aerts (Universiteit van Amsterdam)

Abstract: The role of competition in educational contexts is usually viewed critically, especially among socio-economically vulnerable groups. Yet, more recent insights show that extrinsic motivators, such as performance goals, can also have a positive effect on motivation and achievement, particularly for students who start of less intrinsically motivated. Students with lower SES tend to respond more strongly to extrinsic motivators, such as competitiveness, because they start off less intrinsically motivated on average and consider school less personally meaningful. For this group, a competitive learning environment may actually function as an incentive to activate engagement and interest. In this study, we investigate whether school competitiveness (peer climate, school type and level) contribute to the motivation development in the early years of secondary education, focusing on differences by SES. We use data from the Onderwijsmonitor Limburg (OLM), a large-scale longitudinal study among almost all school in South Limburg. Students are followed throughout their entire educational trajectory, with regular measurements of motivation, achievement, SES, class and school characteristics. In a four-wave design, whether performance goals (T1), caused by school competitiveness, lead to increased interest (T2) and ultimately better performance (T3), controlling for baseline motivation and other relevant factors. We test whether these pathways are stronger for lower SES students, for whom external stimuli are hypothesized to be an entry point for engagement. Linking OML to data from the National Cohort Study on Education (NCO) allows us to continue follow pupils beyond their primary education. In doing so, this contribution not only provides theoretical insights into motivation development, but also policy-relevant starting points for equity and effect learning environment in secondary education. 

The Power of Prediction: Models for Society

Room: Expedition

Chair: Ana Macanovic

Foundation Models for Life-course Modeling, Flavio Hafner (eScience Center), Tanzir Pial (Stony Brook University), Ana Macanovic (European University Institute), Lucas Sage (Sorbonne University)

Abstract: By conceptualizing life events—such as education, employment, and familial milestones—as sequences of temporal markers, we build machine learning models inspired by sequence models from natural language processing. Using Dutch registry data and a vocabulary where each token represents a life event attribute, we pre-train this model with language modeling objectives akin to Masked Language Modelling, aiming to create life-course embeddings for diverse sociological predictive tasks. Preliminary results demonstrate the embeddings’ utility in predicting outcomes such as income levels and fertility decisions. Future work will move in two directions. First, we will explore improvements in model architecture to include a person’s wider life context, such as information about their neighbors, co-workers, and firm identifiers, during training. Second, we will explore the extent to which these models are foundational, i.e., whether their embeddings are generally predictive of outcomes in social science datasets.

Using Interpretable Network Embeddings to Understand Populist Voting Behavior in Population-Scale Registry Networks, Malte Lüken (eScience Center), Javier Garcia-Bernardo (Universiteit Utrecht), Sreeparna Deb (TU Delft), Flavio Hafner (eScience Center)

Abstract: Applying machine learning to administrative registry data has challenged the limits of predicting social outcomes. In deep learning, prediction models often rely on embeddings that project complex input data into a latent numerical space. In the social sciences, embeddings can be used to compress social networks, which reduces their dimensionality while preserving information about the network structure. Registry data can be used to construct social networks at population-scale by creating ties between persons who belong to the same administrative entity (e.g., households, neighborhoods). Such ties are understood to represent social opportunities, reflecting potential interactions and shared environments. Using registry data from Statistics Netherlands, we created person-level embeddings for the entire Dutch population solely based on social opportunity ties in five domains . We demonstrate the usefulness of these embeddings by predicting right-wing populist voting in the Dutch parliament election 2023, linking embeddings with data from a representative survey. While embeddings alone showed a predictive signal, they performed worse than established individual covariates of populist voting. Moreover, they did not improve predictions when combined with said covariates. To make the embeddings more interpretable, we applied regularizing auto-encoders that disentangle the embedding dimensions, enforcing sparsity and orthogonality. We found that one of the regularized embedding dimensions was highly predictive of right-wing populist voting. From this embedding dimension, we created a weighted version of the population network that revealed differences in network structure between higher- and lower-educated persons. These structural differences can help explain right-wing populist voting decisions. We see this study as a starting point to create interpretable population-scale social representations that researchers can use to relate social structure and social outcomes. The embeddings are available to the Dutch research community through the Data Storage Facility by Statistics Netherlands and ODISSEI.

Are births predictable with register data? Evidence from the Predicting Fertility data challenge, Elizaveta Sivak (Rijksuniversiteit Groningen), Gert Stulp (Rijksuniversiteit Groningen)

Abstract: Accurate predictions of life outcomes can inform social theory and policy. Yet, in contrast to predictive success in fields like biology, social science predictions often perform poorly. Why are life outcomes so difficult to predict? One common explanation is small sample sizes of social datasets. We examine this in the context of fertility by studying the predictability of having a child within three years under conditions that are close to the best currently possible. We use full-population data from the Dutch registers and high-quality LISS survey data in the Predicting Fertility data challenge. LISS data includes a wider range of theoretically relevant variables, including subjective measures such as fertility intentions. Administrative data provide vast, high-resolution coverage of life-course trajectories but lack subjective measures. This unique framework enables us to leverage the strengths of these datasets to assess the current predictability and contrast them to learn more about what limits predictability and gain insights about fertility behaviour. Over 150 people participated in the data challenge and submitted over 70 models, ranging from traditional machine learning approaches to state-of-the-art foundational models. Despite the administrative data’s scale and detail, predictive performance remained modest: even the best models fell substantially below the theoretical upper limit of predictability caused by randomness inherent in conception and fetal survival. Survey-based predictions performed slightly better. Analysis of the most important predictors across survey- and register-based models suggests that this difference may reflect the absence of certain key variables in administrative data. Still, the improvement was small, indicating that register data likely captures most major determinants of fertility. These results suggest that accurately predicting individual fertility remains highly uncertain, even in the short term and with extensive data. Modest predictive accuracy is unlikely to stem primarily from limited sample size, but may reflect the inherent unpredictability of life outcomes.

Studying Social Dynamics via Social Media Data

Room: Mission 1

Chair: Jessica Piotrowsk

Public Responses to Jihadist Terrorism on Social Media, Christian Czymara (Nederlands Interdisciplinair Demografisch Instituut)

Abstract: In recent years, several major terror attacks linked to political Islam have shaken Europe. Based on social media data, this study examines how public reactions to Jihadist terrorism are expressed and shaped to explore the formation of public views of immigration and intergroup relations. The theoretical predictions derived from combining terror management theory, group threat theory, and the concept of social resilience, are empirically tested using validated Keyword-Assisted Topic Modeling (keyATM). Unlike traditional topic modeling approaches, keyATM incorporates pre-specified keywords, enhancing both the precision and interpretability of results while mitigating the impact of researcher subjectivity. I analyze over 100,000 time-stamped and geo-coded Tweets on immigration and related issues across four languages in the week following eleven major Islamist terrorist attacks in nine European cities posted by users located in the city of the attack. Consistent with the theoretical predictions, the findings reveal that both threat-related and tolerance-related topics emerge prominently across all four languages. However, deeper analyses uncover temporal shifts in the discourse: While initial reactions often reflect nationalist and exclusionary views, later stages show a rise in inclusionary and tolerance-oriented topics. These results highlight the dynamic nature of social media debates on immigration issues after dramatic events, where initial threat-driven responses often give way to more inclusive and tolerance-oriented discussions as time progresses, underscoring its role in shaping and reflecting the evolving public response to terrorism.

Detecting Asymmetrically Polarized Group Dynamics on Reddit, Laima Baldina (Rijksuniversiteit Groningen), Russell Spears (Rijksuniversiteit Groningen), Hedy Greijdanus (Rijksuniversiteit Groningen), Leah Henderson (Rijksuniversiteit Groningen)

Abstract: Political polarization research increasingly recognizes group identification as a key driver of partisan divides. Social identity approach shows that group identification creates comprehensive psychological alignment over and above shared opinions. This alignment should produce observable group-level dynamics like boundary maintenance, norm development, and coordinated behavior patterns. Yet we have limited evidence of how these collective processes actually unfold over time in real political contexts, shaping both the groups themselves and their members’ experiences. Due to the high level of detail, resolution, and precision, social media data enables investigation of these group processes simultaneously with tracking individual behaviors. Reddit’s architecture provides a particularly instrumental context for studying partisan group behavior due to distinct community boundaries, anonymous interaction, and collective rating system, which amplify group dynamics and make them easier to detect. We analyzed ten years of Reddit data (2012-2021) from five major U.S. political communities, developing computational measures to track group-level behavioral dynamics. For each partisan community, we calculated monthly cross-participation rates in the shared forum r/politics and analyzed how these communities’ contributions were rated by the broader public there. We find systematic asymmetries in how partisan communities evolved their engagement with shared political discourse. Left-wing communities maintained stable participation in r/politics throughout the decade, while right-wing communities showed progressive withdrawal interrupted by temporary mobilization during major political events. Analysis of comment ratings revealed corresponding patterns, where posts by right-wing community members received increasingly negative evaluations from the broader r/politics audience over time, while left-wing members’ posts maintained ratings similar to the community baseline. These findings reveal self-reinforcing dynamics between group participation and reception that shape polarization’s asymmetric evolution. By capturing group-level behavioral mechanisms as they naturally unfold, this computational approach reveals the complex collective dynamics underlying political polarization in naturalistic settings.

Digital visibility and spatial inequality: uneven representation of urban hotspots on TikTok, Shuyu Zhang (TU Delft), Claudiu Forgaci (TU Delft), Lei Qu (TU Delft), Maarten van Ham (TU Delft)

Abstract: In contemporary cities, social media has become an important intervention in the spaces people perceive, imagine and experience. The identification of urban ‘hotspots’ is no longer limited to tourist guides or official planning discourses, but it is increasingly driven by the interplay between platform algorithms and user-generated content. Studies have noted the impact of social media on spatial structure and urban inequality. This study proposes and quantifies an indicator of ‘social media place visibility’ for Amsterdam hotspots on the TikTok platform, revealing how social media reshapes the spatial inequality structure of urban places. By collecting and analysing a large dataset of TikTok posts hashtaged Amsterdam from 2020 to 2023, and applying geo-parsing, spatial mapping, visualisation and place-hashtag network analysis, this study examines the uneven representation of urban hotspots on TikTok. The findings show that the platform hotspots weaken the traditional centre-periphery spatial structure to a certain extent, making peripheral areas such as De Pijp, West and Zuid ‘unexpected hotspots’ due to their high visibility. At the same time, social media platforms have created new inequalities in reconfiguring visibility. The study found that places dominated by consumer activities, such as food and drink, were reinforced on TikTok, while a large number of other urban areas were digitally marginalised, creating a new uneven representation of urban activity spaces and spatial inequalities. In addition, the spatial discourse constructed by the platforms poses a challenge to urban governance, posing the risk of unpredictable consumption behaviours and spatial appropriation. Overall, this study demonstrates how digital platforms have become new actors in the production of urban space, emphasising that in the age of social media, understanding urban spatial inequalities must go beyond the physical to a deeper focus on algorithmic power and the politics of visibility.

Software Showcase

Room: Mission 2

Chair: Kasia Karpinska

Generate synthetic tabular data in a transparent, understandable, and privacy-friendly way. Raoul Schram (Universiteit Utrecht), Erik-Jan van Kesteren (Universiteit Utrecht)

Abstract: Social scientists often handle highly privacy-sensitive information, such as mental health questionnaires, personal income data, or social media information. This makes it ethically and legally problematic for the researcher to share the real data publicly as part of the research process. Among other issues, this can create reproducibility problems: while the analysis code for the research might be published, other researchers cannot reproduce the results without the original data. One solution that can improve the situation is to create a synthetic copy of the original data and publish this synthetic copy. However, generating “safe” synthetic data is not trivial, as it does not automatically protect privacy. There are many methods to create synthetic data, but often it is not transparent how much information is actually retained in the synthetic data. We will present the python package metasyn, which can generate synthetic data in a transparent, understandable and privacy-friendly way. In contrast to most other synthetic data software, we make the explicit choice to strictly limit the statistical information in our model in order to adhere to the highest privacy standards. In the presentation, we first explain the basic principles of metasyn in contrast to other synthetic data methods, after which we show how easy it is to use metasyn in a real-world reproducible social science data workflow. Finally, we present our disclosure control plug-in, which adds an additional layer of security to the synthetic data based on well-known official statistics guidelines for publishing data products. Metasyn is currently funded through a TDCC-SSH grant on synthetic data, which is a collaboration between many different organizations, including Utrecht University, Maastricht University, DANS and SURF. We will show how metasyn contributes to this project and how it will benefit national infrastructure organisations, and talk about the progress in the implementation for these organisations.

Running online behavioral experiments on the SURF Research Cloud, Rense Corten (Universiteit Utrecht), Sandor Spruit (Universiteit Utrecht), Rob Franken (Universiteit Utrecht)

Abstract: Behavioral experiments have become increasingly popular in the past decades across a variety of social science disciplines, including economics, political science and sociology, because such experiments allow to study causal social mechanisms in great detail. When implemented in web-based interfaces, behavioral experiments can be deployed in a flexible manner in a diversity of contexts (e.g., embedded in a survey or in lab-in-the-field settings) or be scaled up to study complex social dynamics in large groups. However, setting up such online experiments is challenging to the typical researcher, as it requires appropriate software tools, programming skills, access to subjects, and hardware to securely and robustly host online experiments. This project aims to address the latter barrier. We present a method to easily set up and up and run a virtual server to host experiments based on the popular open-source platform oTree, levering the computational resources available via the SURF Research Cloud. We demonstrate how a custom-built SRC component can be used to quickly and easily set up a dedicated virtual oTree server, reducing the need for researchers to rely on commercial services or custom-built local hardware solutions.

AmCAT (Amsterdam Content Analysis Toolkit), Sofia Gil-Clavel (Univeristy of Amsterdam), Wouter van Atteveldt (Univeristy of Amsterdam)

Abstract: A key challenge to open science practices, such as sharing and reusing data, is that many researchers consider it a burdensome afterthought rather than a core part of the research itself. The AmCAT  (Amsterdam Content Analysis Toolkit) addresses these needs by providing an open collaboration environment for discovering, accessing, analysing and (re)using textual data. It also facilitates non-consumptive research when full data access is impossible. Combining a user-friendly GUI with a powerful API, AmCAT ensures researchers with varying computational skills can participate. AmCAT aims to bring collective benefit by offering an open source and decentralized solution for storing and analysing potentially sensitive documents, decreasing dependence on commercial or foreign providers and their terms of use throughout the research life cycle. 

Shaping career trajectories

Room: Progress

Chair: Konrad Turek

From Full-Time to Part-Time: Long-Term Maternal Employment Trajectories After Childbirth in the Netherlands, Flora Zhou (Erasmus Universiteit Rotterdam), Tom Emery (Erasmus Universiteit Rotterdam), Jennifer Holland (Erasmus Universiteit Rotterdam)

Abstract: The Netherlands has one of the highest rates of part-time employment in the world, particularly among mothers. While prior research often attributes maternal part-time work to childbirth, a key puzzle remains: when does this shift to part-time employment begin? Is it temporary or permanent? And what factors influence whether mothers return to their pre-birth employment levels? Existing studies have largely focused on short-term or cross-sectional patterns, leaving long-term trajectories underexplored. Using administrative data from Statistics Netherlands, this study applies multilevel linear spline models to track maternal employment trajectories over the 10 years following first childbirth. Unlike traditional non-linear models, linear spline models use piecewise linear segments joined at knots to reflect discontinuities or shifts in slope at key time points. The analysis identifies four distinct stages. The first stage (Year-2 and Year -1) indicates that a high proportion of Dutch mothers worked full-time at the pre-birth stage. The second stage (Year-1 – Year1) indicates that maternal employment sharply decreased during childbirth. The third stage (Year1-Year5) presents that employment continues to decline as the dominant pattern becomes a 3-day (approximately 0.6 FTE) workweek. The last stage (Year5-Year10) presents a very small increase in maternal employment, but it remains well below pre-birth levels. The 3 days working patterns are still dominant. The results also suggest that the number of childbirths and early daycare usage affect the slope of employment at the last stage. Contrary to the common assumption of a V-shaped recovery in maternal employment, our findings suggest a permanent shift: Mothers in the Netherlands tend to adopt and maintain a long-term part-time work pattern following their first childbirth.

New skill adoption and cities, Harm-Jan Rouwendal (Rijksuniversiteit Groningen), Anet Weterings (Planbureau voor de Leefomgeving), Sierdjan Koster (Rijksuniversiteit Groningen)

Abstract: This paper studies the relationship between new technology adoption and agglomeration economies. We asses to what extent the adoption of new technologies differs across space by studying differences in the adoption of new digital skills in the Netherlands. The use of online job vacancy data allows us to study the implementation of new technologies on the skill level in the workplace and empirically analyse the spatial variation in new digital skill adoption within occupations. Results show that employers in urban areas more often demand new digital skills for open jobs compared to similar jobs elsewhere. This indicates new technology adoption is higher in cities, which helps explain the productivity premium of urban areas.

Should I work part-time? Evidence from a Dutch Survey Experiment on Individual Preferences and Societal Norms, Simona Cicognani (Universiteit Leiden), Lea Hauser (Universiteit Leiden), Natalia Montinari (University of Bologna)

Abstract: This study examines individual preferences and perceived social norms regarding part-time work in the Netherlands, a country characterized by one of the highest female part-time employment rates globally. We run a vignette-based survey experiment embedded within the LISS panel dataset, a nationally representative longitudinal survey. Respondents (1,769) were randomly assigned to one of four treatment groups that varied along two dimensions: the protagonist’s gender (female or male) and parental status (with or without a child), in a between-subject design fashion. The study measures first-order beliefs (FOB), representing respondents’ personal preferences about whether the vignette character should reduce working hours, and second-order beliefs (SOB), capturing respondents’ perceptions of societal norms regarding the same. Our results show that both personal and perceived social norms favor reduced working hours more strongly for female protagonists and when a child is present. Importantly, perceived social norms (SOB) reflect that gender and parenthood interact to influence expectations, meaning their combined effect is stronger than each alone. In contrast, personal beliefs (FOB) treat gender and parenthood as having separate, additive effects. Contrary to expectations, female respondents express more conservative personal norms than males, favoring greater reductions in working hours, especially for female protagonists with children. Furthermore, respondents whose mothers worked full-time are less likely to support reduced working hours, suggesting intergenerational influences on work preferences. The analysis of perceived gender pay and pension gaps reveals that women perceive larger disparities than men, but these perceptions do not significantly affect part-time work preferences. Awareness of the negative financial consequences of part-time work predicts stronger support for increased working hours, although this effect does not differ by respondent gender. Overall, these findings highlight how both personal beliefs and perceived social expectations, shaped by gender and parental status, influence attitudes toward part-time work in the Netherlands.

Parenting, Fertility, and the Stories We Tell

Room: Quest

Chair: Pearl Dykstra

Family networks in the transition to parenthood: A predictive machine learning approach, Nicolás Soler (Erasmus Universiteit Rotterdam), Tom Emery (Erasmus Universiteit Rotterdam), Agnieszka Kanas (Erasmus Universiteit Rotterdam)

Abstract: The transition into parenthood implies making choices about childcare that have critical implications for parental employment trajectories, child development, and gender equality. Previous work has studied how childcare choices are influenced by the availability of both formal and informal family care. The latter is typically conceptualized in terms of the availability of support from the paternal or maternal grandparents. However, such perspective disregards how formal childcare uptake is also influenced by the family obligations and resources afforded by a larger set of kin, including both sets of grandparents, as well as aunts, uncles, and cousins. To fully grasp the role that family plays in parents’ childcare strategies, there is a need to adopt a multidimensional perspective that moves beyond simple measures of grandparental presence or geographical proximity to family. We provide a more comprehensive test. We adopt a machine learning approach where we quantify the importance of care availability by predicting whether first-time heterosexual parents will use formal childcare services in the Netherlands. We use register data that allows to measure parental characteristics, the location and capacity of nearby formal childcare providers, and a comprehensive set of characteristics of the family network of children, disaggregating their kin by age, gender, and lineage. The use of machine learning allows us to model multidimensionality while maintaining interpretability through the use of Shapley values. We further interpret the results by studying whether predictive performance and the relative importance of predictors varies across parental educational status. Put together, we provide novel evidence of the role of family care availability on parenting strategies and formal childcare uptake, as well as on how these are stratified by educational status.

Nature-nurture beliefs and parenting, Tim Wienand (Erasmus Universiteit Rotterdam), Hans van Kippersluis (Erasmus Universiteit Rotterdam), Niels Rietveld (Erasmus Universiteit Rotterdam)

Abstract: A long-standing question in the study of child development is to what extent genetic endowments (“nature”) and family environment (“nurture”) influence human capital formation. While extensive research underscores the importance of both dimensions, less is known about individuals’ beliefs regarding their relative importance. At the same time, parents’ perceptions about the process of child development are known to play an important role for their parental investment decisions. Bridging these two strands of literature, we design a survey experiment to investigate whether parental beliefs about the role of genetic endowments for children’s educational achievement influence parental educational time investment. To identify the causal effect of beliefs about the importance of genetics, we incorporate an information experiment in our survey that randomly informs participants about different published heritability estimates of children’s performance at the end of primary school. We then measure a range of outcomes relevant to parental investment decisions. Specifically, we develop new hypothetical vignettes to assess respondents’ views on optimal levels of parental time investment and its and allocation between siblings, perceived returns to ability and parental investment, and attitudes toward equality of schooling outcomes and parental investment between siblings. To validate responses to the hypothetical scenarios, we also ask questions to the subsample of parents related to their own children. The survey was administered in June 2025 as part of the LISS panel, a nationally representative longitudinal household survey in the Netherlands. We aim to present initial findings of our experiment at the ODISSEI conference 2025.

Gender Differences in Fertility Narratives Based on Neural Topic Modelling: Perspectives from Dutch men and women, Xiao Xu (Nederlands Interdisciplinair Demografisch Instituut), Anne Gauthier (Nederlands Interdisciplinair Demografisch Instituut), Gert Stulp (Rijksuniversiteit Groningen), Antal van Den Bosch (Universiteit Utrecht)

Abstract: The impact of gender roles on fertility decisions with distinct perspectives shaped by societal and family roles and their perceived responsibilities. In this study, we examine gender-based differences in where and how individuals express their concerns about fertility intentions and uncertainty in the Netherlands. We use open-ended textual responses collected both in the LISS panel and in the Dutch round of the Generations and Gender Programme (GGP) in 2023. By applying the BERTopic, a state-of-the-art transformer-based topic modeling approach, to thousands of open-ended survey responses, we uncover narrative structures that reflect the differences in underlying beliefs, priorities, and concerns related to childbirth plans by men and women. Our results reveal that gender continues to shape not only fertility plans but also the discourses and attitudes surrounding the reproductive decision-making process.

Urban Inequalities and Social Transitions

Room: Royal Lobby

Chair: Fabio Tejedor

Quantifying social tipping potential: Empirical evidence from three European countries, Christine Hedde-von Westernhagen (Universiteit Utrecht), Francesco Pasimeni (Eindhoven University of Technology), Floor Alkemade (Eindhoven University of Technology)

Abstract: Social tipping dynamics offer a promising mechanism to accelerate net-zero transitions. These dynamics involve rapid, often nonlinear shifts in social-environmental systems, triggered by feedback mechanisms and small interventions, with lasting effects. When change in one domain catalyzes transformation in others, tipping cascades may emerge, reinforcing systemic change. While recent conceptual work has outlined potential tipping domains, empirical evidence remains limited, particularly regarding household energy behaviours. This paper presents empirical analyses of social tipping potential in household energy behaviours across three European countries: Germany, the Netherlands, and the United Kingdom. We developed a new cross-national dataset based on representative surveys conducted in early 2025, including around 3,300 respondents in total, of whom 300 are members of energy communities (ECs). The survey examines energy behaviours, preferences, and attitudes—such as adoption of rooftop solar, electric vehicles, and heat pumps—as well as broader lifestyle practices, views on climate change, and socio-demographic characteristics. Our three-step analytical approach begins by identifying clusters of co-occurring energy behaviours. We then estimate the causal effect of EC membership on these behavioural profiles using econometric matching techniques. Finally, we simulate changes in the population distribution of energy behaviour clusters under different policy and intervention scenarios.
A first descriptive analysis indicates notable differences between countries and groups. For example, the share of households in the general population having adopted or considering the adoption of heat pumps is substantially higher in the Netherlands (60%), as compared to Germany (45%) and the UK (40%). In contrast, EC members in the UK have a higher share of adoption or adoption intention (70%) than EC members in Germany (52%) or the Netherlands (63%). Our inferential analyses outlined above will provide a nuanced account of these differences to help identify the country-specific drivers and barriers to whole-system change in the household energy sector.

Reshaping Economic Segregation in the Netherlands: The Role of Social Housing Allocation (2007–2023), Diego Buitrago (TU Delft), Clementine Cottineau (TU Delft)

Abstract: This paper evaluates the impact of targeted social housing allocation on spatial income segregation in the Netherlands. In 2010, the Dutch government introduced the “90% rule,” requiring housing associations to allocate the vast majority of social rental units to households below an income threshold. While the policy was designed to improve targeting and reduce competition with private landlords, it also altered the spatial distribution of low-income households by constraining access to other neighborhoods. Using detailed administrative data covering the universe of households over a 16-year period, this study assesses how the policy reshaped the geographic allocation of poverty across Dutch cities. The analysis focuses on micro-level changes at the 500-meter grid scale, allowing us to identify localized effects beyond broader national or demographic trends. We disentangle policy effects from compositional and structural forces, and employ an instrumental variable strategy to estimate the causal relationship between changes in household composition and spatial concentration. The results suggest that even income-based policies not explicitly designed as spatial interventions can generate strong place-based effects. These findings contribute to ongoing debates in urban economics on the spatial consequences of non-price housing allocation mechanisms and the trade-offs between targeting efficiency and spatial equity.

The Economic Urban Divide: A Detailed Study of Income Inequality and Segregation in Dutch Urban Areas (2011–2022), Javier San Millán (TU Delft), Clémentine Cottineau-Mugadza (TU Delft) , Maarten Van Ham (TU Delft)

Abstract: Research on segregation and economic inequality is often limited to major capitals and conurbations, neglecting smaller cities. This oversight can lead to public policies based on insights that may not be universally applicable. Leveraging geo-coded register data, this study addresses this problem in the case of the Netherlands by computing income inequality and residential segregation annually in all urban areas from 2011 to 2022. Contrary to most literature, this paper shows that inequality and segregation have remained stable or decreased in most cases. In addition, when looking at how income is distributed among social segments, how segregated they are, and at which geographical scale segregation occurs, we find significant variation between urban areas. More unequal urban areas also tend to be more segregated, but patterns vary, and the same segregation levels can coexist with diverse inequality metrics. Four groups of urban areas are identified through a cluster analysis.

Reading Between the (Code) Lines: LLMs in Social Research

Room: Expedition

Chair: Kevin Wittenberg

Measuring Complexity and Reproducibility: A Comprehensive Benchmark of LLMs for Multilingual Policy Agenda Topic Annotation, Bastián González-Bustamante (Leiden University)

Abstract: Although Large Language Models (LLMs) have rapidly expanded the text-as-data toolkit available to social scientists, their performance remains highly sensitive to task complexity, prompting, language and model parameters. In order to assess these limits, we conduct a cross-lingual benchmark of 79 contemporary LLMs (e.g., OpenAI’s GPTs and o-series, Claude models, xAI’s Grok, Meta’s Llama, Alibaba’s Qwen, Mistral models, among others) on a demanding zero-shot classification task: labelling one of the 21 major topics in the Comparative Agendas Project codebook to bills and Acts drawn from Danish, Dutch, English, French, Hungarian, Italian, Portuguese and Spanish corpora. We then conducted a meta-analysis that shows reasoning models lift F1-scores by around nine percentage points on average. In comparison, reproducible models under deterministic deployment conditions incur a cost of about eight percentage points in F1-score. In addition, preliminary and ongoing fine-tuning experiments suggest that fine-tuned transformers may surpass the best zero-shot LLM classification performance. These findings quantify the trade-off between performance and reproducibility, highlighting chain-of-though reasoning as the most effective option for complex policy agenda annotation. They also set a baseline for future venues, where systematic fine-tuning and alignment strategies could be used to test whether some models can close the performance gap while retaining full reproducibility.

Care Work and Code Work: Can Language Models Read Between the Lines?, Elisabet Doodeman (Vrije Universiteit), Arjen de Wit (Vrije Universiteit), Dominik Meier (University of Basel), John Mohan (Third Sector Research Centre)

Abstract: Can LLMs distinguish intentions to care for family, friends, and unknown others? We explore this question with text answers to an open-ended survey question. In 2008, a national UK survey asked 9,760 participants, all born on the same day and year, to describe their imagined life at age 60. These open-ended responses offer a rich but underexplored opportunity for computational text analysis, particularly for identifying intentions to volunteer or care for others within broader life narratives. Notably, participants were not prompted to discuss volunteering or care, reducing the influence of social desirability bias and allowing these themes to emerge in more naturalistic and varied ways. We explore the use of large language models (LLMs) as research assistants in coding these responses, with a focus on detecting care-related themes. What began as a technical investigation quickly raised deeper conceptual questions: What does it mean to care? What qualifies as “care” in the context of future-oriented self-descriptions? What seems obvious to a human coder may be ambiguous to a model, and vice versa. How can we communicate socially constructed, value-laden concepts to an LLM in a way that preserves meaning and nuance? Our experiments with several models culminated in promising results using Llama 3.3 70B Instruct, with precision and recall metrics higher than 0.9. We reflect on both the model’s performance and the broader methodological implications of using LLMs for interpretive tasks in social science. This work contributes to emerging conversations on the role of AI in qualitative analysis, highlighting challenges in alignment, abstraction, and conceptual agreement between humans and machines.

Trying to be human: Linguistic traces of stochastic empathy in language models, Bennett Kleinberg (Tilburg University), Jari Zegers (Tilburg University), Jonas Festor (Tilburg University), Stefana Vida (Tilburg University) 

Abstract: Differentiating between generated and human-written content is important for navigating the modern world. Large language models (LLMs) are crucial drivers behind the increased quality of computer-generated content. Reportedly, humans find it increasingly difficult to identify whether an AI model generated a piece of text. Our work tests how two important factors contribute to the human vs AI race: empathy and an incentive to appear human. We address these aspects in three experiments: human participants and a state-of-the-art LLM wrote relationship advice (Study 1, n=530) or mere relationship descriptions (Study 2, n=610), either instructed to appear as human as possible or not. New samples of humans (n=428 and n=408) then judged the texts’ source. Our findings show that when empathy is required, humans excel. Contrary to expectations, instructions to appear human were only effective for the LLM, so the human advantage diminished. Study 3 showed that these effects persist when the instructions are targeted at avoiding the appearance of an LLM. Computational text analysis (Study 4) revealed that LLMs become more human because they may have an implicit representation of what makes a text human and apply these heuristics. The model resorts to a conversational, self-referential, informal tone with a simpler vocabulary to mimic stochastic empathy. We discuss these findings in light of recent claims on the on-par performance of LLMs. 

 

Green talk, brown walk: Computing the gap between pollution and sustainability discourse, Mave Powlick (VU Amsterdam), Rajat Ravi Rao (Universiteit Leiden), Koen Steenks (VU Amsterdam), and Ava Zwolinski (University of Amsterdam).

Abstract: This presentation examines the discrepancy between corporate sustainability rhetoric and actual environmental performance, commonly termed greenwashing. We combine existing sentiment analysis of company websites (Ernst, 2025) and FIRMBACKBONE data on firm characteristics and emissions to develop a greenwashing index. We then analyze the role of bullshitting, operationalized as the use of words indicating uncertainty or non-commitment (hedge words), to predict the presence of greenwashing.  Through dictionary analysis, we find that use of hedging language is significantly associated with higher greenwashing scores, with variation across industries and workforce demographics. These findings demonstrate how corporate communication strategies can obscure environmental behavior, while also showing the potential of computational linguistics to identify sustainability-related misinformation. Our approach can contribute to policy and stakeholder efforts to detect greenwashing, enhance transparency, and promote more reliable environmental reporting.

Room: Progress

Closing Keynote: AI Readiness for Computational Social Science and Humanities – Mercè Crosas (Barcelona Supercomputing Center)

Abstract: The Barcelona Supercomputing Center (BSC) has launched a new initiative, the Computational Social Sciences and Humanities (CSSH) Lab, which leverages the use of new data types, AI model systems, and computational workflows to study the past and present, to improve the future. A large part of what is required to conduct CSSH´s large-scale computational studies is to prepare the data so that it´s machine-actionable and appropriate for AI adoption. This talk will introduce some of the new AI-based research in the CSSH Lab at BSC and review what is needed to adopt AI in science and become AI-ready, while preserving responsible, rigorous, and open science.

 

The Closing Keynote will be followed by closing remarks from ODISSEI Scientific Director Daniel Oberski.