Media / Publication
March 30, 2026

How Much Data Should I Request? Balancing Richness and Compliance in Digital Trace Data Donations

© Image: Unsplash/Markus Spiske

Asking people to donate more of their social media data can lead to richer results, but may also mean losing more participants along the way. This study shows that Facebook and Instagram data requests spanning longer periods increase dropout largely due to the longer delivery time from the platform, while also fundamentally shaping what researchers can validly measure.

Abstract

Digital trace “data donation” studies offer researchers a unique opportunity to collect high-quality behavioral data, but decisions about the scope of requested data can impact both dataset richness and participant compliance. This paper examines the tradeoffs between requesting larger data packages, which include more extensive historical records, and participants’ willingness to donate. In a randomized experiment with Facebook and Instagram data donations, we compare a control condition where participants are asked to request the default 1-year data period to a treatment condition in which they are asked to request data for their entire account history. We analyze how different request sizes affect (1) participants’ compliance rates and (2) the characteristics of the data resulting from these different requests. We find that participants asked to request more data are less likely to complete the task. However, we propose that this is not primarily due to heightened privacy concerns, but rather because these data packages are significantly larger and therefore take longer for the platforms to deliver. This additional time to deliver data packages results in increased attrition. In terms of the effects on the data itself, we show that decisions about the time-span of the data impacts not only the volume of data requested, but also has implications for measurement validity, as the temporal window fundamentally redefines what key constructs represent, potentially transforming intended static indicators into narrow snapshots of recent behavior. We provide guidance for researchers navigating these decisions, considering both the benefits of richer longitudinal data and the risks of reduced participation.

More results /

/ algosoc
AI insiders warn it could end humanity. Does the public agree?

By Ernesto de León • Claes de Vreese • September 17, 2026

Hou DigiD uit Amerikaanse handen nu het nog kan

By José van Dijck • April 30, 2026

/ health
The expectation game

By Martijn Logtenberg • March 13, 2026

Labour politics and why AI hasn't 'fixed' healthcare

By Martijn Logtenberg • November 20, 2025

Burying the Lead: Adjusting Goals to Manage Functional Limitations of AI Tools in Healthcare

By Jacqueline Kernahan • Richard Bartels • Mark de Reuver • Daniel Oberski • Roel Dobbe • October 20, 2025

/ media
The Label Paradox: Can AI transparency create more distrust?

By Natali Helberger • September 21, 2026

How does political efficacy condition clicks on politics? Understanding information-selection behavior in algorithmic feeds over time

By Jin Wan • Theo Araujo • Natali Helberger • Claes de Vreese • September 17, 2026

A Jug of Settled-Down Juice: AI Guidelines on Transparency Obligations

By João Pedro Quintais • September 03, 2026

Subscribe to our newsletter and receive the latest research results, blogs and news directly in your mailbox.

Subscribe