Human Engagement Through Technology Way Beyond Counting Views, Clicks and Likes

By Lesandro Ponciano ORCID iD, on 24 May 2026.

Our research over the last 15 years has systematically investigated how people engage through technology-mediated systems. The primary pursuit is both scientific and technological. The scientific aspect focuses on understanding of human behaviour when engaging in activities that require some kind of action on their part. We are especially interested in contexts where these actions are not financially motivated, which includes contexts spanning citizen science platforms, human computation projects, and social media for climate action. The technological aim is to model how people participate in an activity via technology, so that the technology can be designed to support better participation.

Here, we put it all together summarizing what we have learned. It is organised into the following questions.

What is the operational definition of human engagement?

The theoretical foundation of the research draws primarily from the conceptual framework of user engagement proposed by O'Brien and Toms (2008), adapted for the study of technology-mediated participation. Conceptual framework is a theoretical structure of interconnected ideas, assumptions, and principles that guides a study or project. It outlines how the core variables are expected to relate to one another, providing a clear map for the entire research process. This framework conceptualises engagement as a process comprising four distinct stages: the point of engagement, which is the moment an individual performs the first action within a system (e.g., executing the first task or posting the first reply); a period of sustained engagement, which is a continuous, uninterrupted span of participation (e.g., a single working session on a citizen science platform or a sequence of replies to a social media thread); disengagement, which occurs when the sustained period ends; and re-engagement, which denotes a new cycle of engagement after a period of absence. This process-oriented view is deeper than counting clicks or 'likes' because it captures the temporal dynamics, recurrence, and intensity of human investment. It acknowledges that engagement is not merely an event but a psychological and behavioural process wherein individuals self-invest personal resources—including time, cognitive effort, and social skills—over potentially multiple cycles.

To operationalise the conceptual framework, the thread of research discussed here adopts a quantitative, data-driven approach that relies on behavioural logs generated by the systems under study to reconstruct the process of engagement. This approach requires fine-grained, time-stamped records of individual actions. Specifically, the necessary data for this type of analysis include: (1) a unique identifier for each participant (volunteer or citizen), (2) a unique identifier for each task performed, (3) a timestamp for each action (date and time, often to the second), and (4) contextual metadata, such as the project identifier, task type, or, in the case of social media, the publication to which a reply is addressed. From these logs, sequences of actions per individual are extracted, inter-action intervals are computed, and participation sessions are reconstructed.

This type of data allows us to derive quantitative metrics that capture both the degree and the duration of engagement. Importantly, this approach is non-intrusive, scalable, and capable of revealing natural behavioural patterns without relying on self-reported data.


Which probability distribution best describes engagement behaviour?

Published in 2014, the Volunteers' Engagement in Human Computation for Astronomy Projects study is the first to characterise statistics distribution of engagement of volunteers in citizen science and human computation fields. It is based on four primary metrics:

The study divides participants into two broad behavioural classes: transients and regulars. Transient volunteers are those who executed tasks on only one day and never returned. Regular volunteers, in turn, are those who returned on at least one additional day after their first task execution. From the volunteering literature, this distinction between transient and regular operationalises theoretical categories from social psychology: helping behaviour vs volunteering behaviour.

Citizen science involves members of the general public actively collaborating in scientific research, typically through large-scale data collection and analysis. Human computation is a computer science technique that crowdsources tasks requiring human intelligence, such as pattern recognition, to solve complex problems that algorithms cannot yet handle efficiently. The paper analyses real citizen science projects based on human computation tasks from the Zooniverse platform and shows that the majority of participants are transient (>60%). Despite being the minority (<40%), regular volunteers contributed the majority of tasks (>75%). This characterization presented in the paper proved quite important in revealing that most participants in online citizen science contribute for only one day and then never return. Thus, they are closer to helping behaviour than to volunteering behaviour. After the publication in 2014, this finding has been replicated by dozens of other studies in citizen science.

Using goodness-of-fit tests, the study shows that Frequency follows a Zipf distribution. Daily productivity, typical session duration, and devoted time, in turn, follow log-normal distributions. Goodness-of-fit test is a statistical hypothesis test used to determine how well observed sample data matches a specific theoretical distribution. It assesses whether the discrepancy between the actual data and the expected values is merely due to random chance or is statistically significant. These distributions reveal a heavy-tail pattern: a small minority of volunteers exhibit very high frequency, productivity, or session duration, while the majority exhibit low values. The log-normal distribution for session duration was similar across both projects, despite differences in task complexity, indicating that volunteers tend to devote a characteristic amount of continuous time regardless of the task's cognitive load. This analysis, presented in the paper, proved to be quite important because it was one of the first to show that the unequal participation in online citizen science projects is similar to the unequal contribution in other online systems.

These findings provide the first large-scale, empirically grounded statistical characterisation of volunteer engagement in human computation for citizen science. They challenge the assumption that all volunteers contribute similarly, demonstrating that engagement is highly unequal. The study drew attention to the fact that most participants are not people who maintain long-term engagement, but are merely responding to a call for contribution that they received via outreach through traditional or social media. For citizen science practice, the results imply that retention strategies should focus on converting transients into regulars, as regulars drive project effectiveness.


What are the engagement profiles that people typically exhibit when participating through technology?

Also in 2014 the Finding Volunteers' Engagement Profiles in Human Computation for Citizen Science Projects study was published as the first study to propose a cluster-based approach to characterise volunteer engagement profiles in citizen science. It suggested that in addition to investigating the general engagement patterns common to the entire volunteer population, it is worthwhile to analyse distinct groups of volunteers who share similar behavioural traits. Clustering is an unsupervised machine learning method used to group a set of data points into distinct categories based on their inherent similarities. It ensures that elements within the same cluster share common characteristics, while differing significantly from elements in other groups.

The study defined four metrics derived from the engagement conceptual framework in order to consider both the duration of engagement and the degree of engagement:

The profile analysis was restricted to volunteers active on at least two different days (regular volunteers, class identified in the other study described above). The process to find profiles involved:

  1. Normalisation of each metric to [0,1] using range normalisation.
  2. Hierarchical clustering (dendrogram analysis) to identify a suitable range for the number of clusters (k).
  3. K-means clustering with initial centroids from hierarchical clustering, varying k from 2 to 10.
  4. Quality assessment using within-group sum of squares and average silhouette width.
  5. Selection of k = 5, as this yielded the best trade-off.

The analysis presented in the study, also carried out on big Zooniverse projects, has consistently identified five distinct engagement profiles:

The opposition between hardworking (high activity ratio, short duration) and persistent (low activity ratio, long duration) profiles reveals a trade-off: volunteers who engage intensely tend to burn out quickly, while those who engage lightly can sustain participation over months. The strong negative correlation between activity ratio and relative activity duration in the moderate profile suggests that, for most volunteers, high frequency and long-term persistence are mutually exclusive.

Thus, the study back in 2014 was the first to empirically derive natural engagement profiles from behavioural data in citizen science. It moves beyond aggregate statistics to provide a typology that can inform personalised engagement strategies. For volunteering literature, the profiles map onto different motivational and behavioural pathways, suggesting that volunteerism is not a single category but a family of behaviours. Practically, project managers can use these profiles to target interventions: for example, converting hardworking volunteers into persistent ones by pacing tasks, or re-engaging spasmodic volunteers during new project phases.


On platforms that host multiple projects, people test various projects, get involved in all of them, just explore, or what?

Published in 2019, the Characterising Volunteers' Task Execution Patterns Across Projects on Multi-Project Citizen Science Platforms study was the first effort to characterise citizen engagement in platforms that host multiple projects, e.g., whether volunteers migrate from a project to another. Beyond investigating engagement behaviour within individual projects in isolation, the research has progressed to exploring engagement across the broader context of multiple projects in which a person may participate. Interestingly, in the 2014 studies discussed above, it made sense to investigate engagement isolated citizen projects because that was the main paradigm; projects were created independently. But after that, projects began to aggregate and come together on common platforms, which introduces a new dimension of engagement that is addressed in the 2019 study.

The study proposes three classes of volunteer:

The study employed the Goal, Question, Metric (GQM) approach to define metrics from three perspectives: Goal, Question, Metric (GQM) approach is a software engineering and project management framework designed to drive meaningful evaluation. It establishes a high-level Goal, derives specific Questions to determine if that goal is being met, and identifies precise Metrics to supply the necessary data.

Three multi-project platforms were analysed in the study: Crowdcrafting, Societize and GeoTag-X. The results show that volunteers who engaged in multiple projects (explorers or regulars) exhibited significantly longer relative activity duration than those who stayed with a single project. This suggests that cross-project exploration acts as a retention mechanism: exposing volunteers to diverse projects may re-engage them when their initial project no longer holds their interest. However, the very low proportion of multi-project regulars (≤6%) indicates that deep, sustained engagement across multiple projects is rare. Gini coefficient is a statistical measure of dispersion that quantifies the degree of inequality in the distribution of any quantifiable resource or attribute across a system's elements. It ranges from 0, representing perfect equality (where the resource is shared evenly), to 1, representing absolute inequality (where a single element holds the entire resource). Inequality metrics (Gini > 0.9 on Crowdcrafting) show that a tiny subset of projects captures most volunteer attention, creating a winner-take-all dynamics. Interestingly, new projects recruited more volunteers externally than they inherited from the platform, but inherited volunteers performed more tasks on average (positive balance in computing), indicating that platform-sourced volunteers are more valuable in terms of effort.

The study contributed to extending engagement research from single projects to the platform level, revealing that cross-project dynamics are critical for platform sustainability. For volunteering literature, it shows that volunteering behaviour can be context-dependent: the same individual may be a transient on one project and a regular on another. Practically, platform designers should implement personalised project recommendations and provide feedback for cross-project participation to encourage exploration and retention. The finding that inherited volunteers contribute more tasks suggests that building a platform-wide community is more effective than external recruitment campaigns for securing high-quality effort.


How does credibility assessment relate to engagement, and why does it matter?

Engagement alone does not guarantee quality. Volunteers may be highly engaged, contributing frequently and for long periods, yet still provide low-quality answers due to lack of expertise, task difficulty, or ambiguous instructions. Conversely, a transient volunteer might provide a single but highly accurate answer. Credibility means "believability" and it is mainly associated with the notions of trustworthiness and expertise, reputation and relevance. From this point of view, a highly credible participant in a specific domain is someone having high expertise in such domain, and being able to provide relevant information and be trusted, so others can rely on that information. Therefore, understanding engagement without understanding credibility is incomplete. Published in 2018, the Agreement-based Credibility Assessment and Task Replication in Human Computation Systems study addresses this gap. The research investigated how to automatically measure the credibility of participants based on their patterns of agreement with others, and how this credibility information can be used to optimise the operation of human computation systems.

The study proposes four alternative metrics for measuring participant credibility based on inter-rater agreement:

Task replication is the strategy of having the same task executed multiple times by different participants. It is the standard mechanism for quality assurance in human computation systems, including citizen science platforms. The study also proposes an adaptive credibility-based task replication algorithm. Instead of replicating every task a fixed number of times (e.g., always 10 or 38 replicas), the algorithm stops replication as soon as a group of answers reaches a required credibility threshold. If participants strongly agree on an answer, fewer replicas are needed. If they disagree, the algorithm continues replicating until either the credibility requirement is met or a maximum replication limit is reached.

The evaluation used data from two real projects: Sentiment Analysis and Fact Evaluation. The results show that the adaptive algorithm reduces the number of replicas without compromising accuracy. Experienced agreement and reputed agreement consistently outperform surface agreement, appearing exclusively in the Pareto-optimal configurations. Task difficulty, measured via Shannon entropy, strongly correlates negatively with both answer credibility and replication reduction. Difficult tasks yield lower credibility and require more replicas. The replication algorithm allows requesters to balance competing requirements through three parameters: required credibility (desired confidence in the final answer), urgency (parallelism vs. sequential replication), and maximum replicas (cost or resource limit). Shannon entropy, used in the study to measure task difficulty, quantifies the divergence among answers to the same task. Higher entropy means more disagreement, indicating either a genuinely difficult task or a poorly designed one. This metric was originally introduced in information theory and has been widely applied in crowdsourcing quality assurance.

How does this connect to engagement? Credibility assessment and engagement measurement are two complementary dimensions of participant behaviour. Engagement tells us how much and how persistently people participate. Credibility tells us how reliably they participate when they do. A volunteer with high engagement but low credibility (e.g., a hardworking profile who consistently disagrees with the majority) may indicate a misunderstanding of the task or a deliberate cheater. Conversely, a volunteer with low engagement but high credibility (e.g., a persistent profile who contributes rarely but accurately) is a valuable asset that should be retained. The five engagement profiles identified in the 2014 study (hardworking, spasmodic, persistent, lasting, moderate) can be further enriched by overlaying credibility scores per difficulty level, enabling personalised task routing, feedback, and retention strategies.

Credibility assessment and engagement measurement are not separate concerns. They are two lenses on the same phenomenon: how people contribute to technology-mediated systems. Measuring both allows us to design systems that are not only effective (accurate answers) but also efficient (fewer wasted replicas) and respectful of participants' time, a critical factor for sustaining long-term engagement in citizen science and other non-financially motivated contexts.


When participation means interacting publicly with an authority, how do citizens engage?

As individuals often prefer to engage on platforms they already use daily, it is also worthwhile to understand how the public interact with information posted by an authority on generic, commonly used social media platforms, rather than on systems maintained by the authority itself. The How Citizens Engage with the Social Media Presence of Climate Authorities: the Case of Five Brazilian Cities study, published in 2023, operationalised engagement using metrics from the human engagement framework in the social media context.

The study proposed the following metrics:

In addition to the quantitative aspect, the study also analysed qualitative dimensions of engagement in terms of what citizens contribute when engaged. Since people participate in this context by posting responses to publications from authorities, the content of these responses is analysed using natural language processing.

Survival analysis is a statistical approach for modelling time-to-event data. It accounts for censoring and estimates the probability of an event occurring over time, making it suitable for studying durations, retention, dropout, and failure in longitudinal studies. The study also establishes a parallel between transient and regular citizens (citizens who replied only within 24 hours and never again vs. those whose activity spanned beyond 24 hours). The vast majority of citizens exhibited transient engagement: replying only once or within a 24-hour window and never returning. Survival analysis shows that the probability of engagement lasting more than 200 days was below 0.25. Reaction times were typically minutes to hours, indicating that citizens engage immediately when a publication appears in their feed, but rarely revisit older publications.

Latent Dirichlet Allocation (LDA) is a generative statistical model used in natural language processing to discover abstract, hidden topics within a large collection of text documents. It operates on the assumption that every document is a mixture of various topics, and that every word can be traced back to one of those topics. Unsupervised Latent Dirichlet Allocation analysis applied to identify themes in citizens' replies shows that the nature of contribution is corrective, complementary, updating, thanking, or donation-related. Citizens engaged in corrective ("The information is wrong; no rain here"), complementary ("Here in Pampulha, it is raining heavily too"), and updating ("It was raining heavily, but now it has stopped") behaviours. This suggests that citizens treat authorities as sources of operational weather information and tend to engage with them when they realise that something does not match their reality.

For climate action literature, this study fills a practical-knowledge gap by providing a year-long, behavioural analysis of citizen-authority communication. It demonstrates that social media, despite its potential, predominantly facilitates short-term, reactive, and location-specific engagement rather than sustained climate dialogue. For volunteering literature, it extends the transient/regular distinction to the context of public service communication, showing that even for highly relevant, risk-related information, sustained engagement is rare. Practically, authorities should not assume that social media builds long-term climate awareness or trust; instead, they should design for rapid, localised corrections and complements, and consider that citizens' preferred platforms (social media) may not align with the most effective communication channels (e.g., SMS with geolocation).


What are the main discoveries and implications for engagement theory?

Over a decade of research, these studies established that engagement with technology is a multidimensional process characterised by distinct behavioural patterns that can be robustly measured using log data. The key findings reveal that engagement is not a uniform phenomenon; rather, participants can be categorised into distinct behavioural classes, and engagement profiles are guided by psychological drivers and design strategies. The main discoveries are:

  1. engagement follows highly unequal statistical distributions (Zipf, log-normal) across all studied contexts (citizen science, human computation, and social media) with a small minority of participants (regulars, persistent volunteers) contributing the vast majority of effort;
  2. beyond simple transient/regular dichotomies, volunteers naturally cluster into five reproducible profiles (hardworking, spasmodic, persistent, lasting, moderate) that reveal fundamental trade-offs between the degree and duration of engagement;
  3. engagement profiles persist across different projects and platforms, suggesting they reflect deep-seated human behavioural tendencies;
  4. in some contexts, measuring engagement without measuring credibility is incomplete and may waste the effort of those who do engage;
  5. multi-project engagement acts as a retention mechanism, but deep cross-project regularity is rare; inherited volunteers from a platform are more engaged than externally recruited ones;
  6. when assessed in preparedness campaigns conducted by climate authorities through social media, citizen engagement is overwhelmingly transient, motivated by corrective messages.

Collectively, these discoveries are based on a diversity of types of participation and roles that the computer system plays in mediating these participation types. The discoveries underscore that effective engagement strategies must be profile-oriented, context-aware, and realistic about the limits of technology-mediated intrinsically motivated participation. They must recognise that sustained, deep engagement is the exception, not the rule, and that system design should accommodate this behavioural reality rather than assume or project uniform participation.

Further Reading

The following links to the source texts provide full details about the works discussed here.

  1. Volunteers' Engagement in Human Computation for Astronomy Projects
  2. Finding Volunteers' Engagement Profiles in Human Computation for Citizen Science Projects
  3. Characterising Volunteers' Task Execution Patterns Across Projects on Multi-Project Citizen Science Platforms
  4. Agreement-based Credibility Assessment and Task Replication in Human Computation Systems
  5. How Citizens Engage with the Social Media Presence of Climate Authorities: the Case of Five Brazilian Cities