AN ORGANIZATIONAL EVENT INTELLIGENCE SYSTEM FOR DETECTING RECURRING OPERATIONAL FAILURES, DISCOVERING ROOT CAUSES AND GENERATING PREDICTIVE RESOLUTION INSIGHTS - (PROBLEM TO PREVENTION)
DOI:
https://doi.org/10.62643/Abstract
Organisations of every size depend on IT services, applications, networks and business processes that fail from time to time. Each failure leaves a trail of events: monitoring alerts, application logs, incident tickets, change records and notes written by support engineers. In most organisations these records are used only to close the immediate incident, and the same problems return week after week because nobody has the time to study the pattern behind them. This paper presents Problem to Prevention, an organisational event intelligence system that collects operational events from several sources, detects recurring failures, discovers likely root causes and generates predictive insights that help teams resolve and prevent incidents. The system builds a data pipeline that ingests alerts from monitoring tools, log streams from servers and applications, incidents and service requests from the ticketing system, and change and deployment records from the release management process. Each event is normalised to a common schema containing time, affected configuration item, service, severity and text description. Log messages are converted to templates with a log parsing algorithm, ticket text is cleaned and embedded with a sentence transformer model, and all events are linked to a configuration map that shows how applications, servers, databases and network devices depend on one another. Three analytical stages operate on the unified event store. First, a clustering stage groups similar incidents using text embeddings and shared configuration items, which reveals recurring failure families that were previously logged as unrelated tickets. Second, a root cause stage mines association rules and temporal precedence between events, such as a configuration change followed by a rise in database errors, and ranks candidate causes using the dependency graph. Third, a predictive stage trains a gradient boosting classifier on event sequences to estimate the probability that a service will suffer a major incident in the next few hours, and a resolution recommender retrieves fixes that worked for similar past incidents.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













