AN EVENT-DRIVEN BEHAVIORAL DATA PIPELINE FOR DISCOVERING SOCIAL INTERACTION PATTERNS AND SUPPORTING EVIDENCE-BASED DECISION ANALYSIS - (SOCIAL PULSE)
DOI:
https://doi.org/10.62643/Abstract
Online communities, campus portals, and organisational collaboration platforms generate a continuous stream of small behavioural events such as posts, replies, reactions, mentions, group joins, and event check-ins. Each event on its own says very little, but taken together they describe how people actually interact, which groups are active, who connects otherwise separate groups, and how participation changes after a policy or programme is introduced. In most organisations this information is either ignored or examined through occasional manual reports built from spreadsheets. This paper presents Social Pulse, an event-driven behavioural data pipeline that captures interaction events as they happen, organises them into an analysable form, and turns them into evidence that decision makers can use. The pipeline receives events from the platform through a lightweight collector and publishes them to an Apache Kafka topic with a common event schema. A schema registry validates every event, and identifiers of users are pseudonymised with a keyed hash before the data leaves the ingestion tier. Spark Structured Streaming jobs clean, deduplicate, and sessionise the events and write them into a lakehouse organised as bronze, silver, and gold layers on Delta tables. The gold layer holds daily interaction graphs, per-user activity features, group-level engagement summaries, and a curated table of interventions such as new features, announcements, or policy changes along with their start dates. On top of these tables the system runs several analytical models. A weighted interaction graph is built for each week, and the Louvain algorithm detects communities while degree, betweenness, and PageRank centrality identify highly connected and bridging members. Sequential pattern mining with PrefixSpan discovers common chains of actions within sessions, and a K-Means model on engagement features groups users into behavioural segments. An isolation forest flags sudden bursts of activity that may indicate spam or coordinated behaviour. For decision analysis, a difference-in-differences model with bootstrapped confidence intervals estimates whether an intervention changed engagement compared with a similar group that did not receive it.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













