Zebrium's Unsupervised ML Transforms Software Incident Detection Across Multiple Platforms
Zebrium employs unsupervised machine learning to identify software incidents across multiple platforms. This detailed examination explores the integration of Zebrium's system with Datadog for real-time monitoring, the automated process of root cause analysis, and the platform's capabilities in data collection and transmission. Additionally, the article examines how Zebrium enhances incident management through comprehensive alerts, notifications, and seamless integrations with_ops management tools like Opsgenie, PagerDuty, and VictorOps.
To integrate with Datadog, Zebrium requires users to create an API Key in Datadog, which is then used to configure the integration within Zebrium. The basic setup involves creating a new integration, selecting Datadog Events and Metrics, enabling Send Detections, and entering the API Key. This configuration sends log metrics to Datadog for visualization, including counts of all logs, anomaly logs, and error logs per service group and deployment.
Zebrium provides detailed metrics sent to Datadog, including zebrium.logs.all.count for all log events, zebrium.logs.anomalies.count for anomaly logs, and zebrium.logs.errors.count for error logs - all reported at one-minute durations for service groups and deployments. Datadog displays this data in various formats, with root cause report suggestions appearing as events each time a detection occurs. The severity of these alerts is categorized as low, medium, or high, providing clear visibility into the significance of detected issues.
For more advanced use cases, Zebrium integrates with Datadog's triggered monitors through a webhook integration. This process begins by creating a Webhook integration in Datadog, configuring it to include alert_transition data in the payload. Users then enable the Datadog integration in Zebrium, selecting Send Detections and entering the API Key. When triggered monitors fire, they initiate a process where Zebrium finds coinciding anomalous log patterns and generates Root Cause reports. These reports are automatically displayed in Datadog dashboards, complete with summary information and correlated logs, allowing for rapid root cause analysis.
When Zebrium's AI/ML engine detects abnormal log patterns, it generates suggestions that appear on the Alerts page, the system's home screen. These suggestions contain several key elements:
AI-generated titles, developed using the GPT-3 language model, help summarize the nature of the detected anomaly. While these titles provide useful context, users should verify their accuracy before acting on the suggestion.
A word cloud displays relevant keywords selected by the AI/ML engine from the associated log events, helping users quickly identify potential issues.
The significance icon indicates the AI's confidence level in its detection, with red representing high confidence and yellow indicating medium confidence. Hovering over this icon reveals the specific confidence percentage.
The Root Cause (RCA) report summary presents the actual cluster of anomalous log lines identified by the AI/ML engine. Typically, up to eight of these lines are shown in the summary view, allowing users to quickly assess the nature of the detected issue.
Each suggestion includes one or two log lines marked with an alert key icon, which form the basis of an alert rule. Users can accept, reject, or ignore these suggestions, with rejected suggestions prevented from triggering future alerts.
During the initial 24-hour learning period, the AI/ML engine builds an event catalog and identifies normal log patterns. While the system typically achieves reasonable accuracy within this timeframe, users may need to allow up to 24 hours for new environments to be fully analyzed.
To optimize results, users are advised to provide comprehensive logs containing unusual events or significant errors. The system requires complete log collection, including troubleshooting information, and benefits from a 24-hour time range before triggering analysis.
The Zebrium OnPrem solution supports multiple data ingestion methods to collect and transmit logs to the platform. The command-line interface allows users to submit logs using the ze up command, while Kubernetes environments can leverage the LogCollector, which operates as a Helm chart with specific configuration requirements. For Logstash-based deployments, users can integrate Zebrium through a dedicated collector.
The system requires several parameters for configuration, including the ZAPI endpoint URL, authentication token, deployment name (as a single lowercase word), and host timezone. For Kubernetes deployments, additional settings such as zebrium.collectorUrl, zebrium.authToken, zebrium.deployment, and zebrium.timezone must be defined. The OnPrem system collects logs from all namespaces within Kubernetes clusters and transmits them to Zebrium's remote monitoring infrastructure.
To ensure secure communications, Zebrium supports multiple notification channels including Slack and email. The preferred channel for company communications is a shared Slack channel created by Zebrium, with additional debug alerts sent to a separate dedicated channel. Users must configure these channels within their Slack instance and invite Zebrium to join the relevant conversations.
For detailed integration, Zebrium provides two main collector options. The Kubernetes collector uses the official ze-fluentd-plugin and supports various Linux distributions, requiring configuration through the command line after initial setup. The Linux collector employs a Fluentd output plugin, allowing integration with existing td-agent setups without requiring a full installer. Both collectors support additional configuration options for custom log paths and advanced features like overwrite configuration.
The platform supports multiple operational modes, including Auto-Detect and Augment. In Auto-Detect mode, Zebrium continuously monitors all application logs using unsupervised machine learning to detect anomalies with over 95% accuracy. When issues are identified, it generates Root Cause reports and integrates with Elastic Stack, automatically adding report metrics to Kibana Dashboards for visual analysis.
The Augment mode operates in conjunction with Dynatrace, receiving signals from triggered monitors to find coinciding anomalous log patterns. It creates comprehensive Root Cause Reports that summarize findings and provide detailed word clouds of relevant keywords. Each report includes a set of log events that demonstrate symptoms and root causes, with clickable links to full reports in the Zebrium user interface. This integration streamlines incident response by reducing Mean Time to Resolution (MTTR) and minimizing the need for manual root cause analysis.
Zebrium generates three types of alerts through its platform:
Accepted Alert: Indicated by a green circle, these alerts represent suggestions that have been accepted by users or other Zebrium users.
Custom Alert: Triggered by writing regular expressions to search for specific patterns, these alerts are represented by a blue triangle.
Rejected Alert: Identified by a red triangle, these alerts mark suggestions that were rejected as irrelevant to the user's environment.
The platform presents these alerts using a Root Cause Timeline Widget that provides several key features:
Displaying date/time information, titles, and word clouds for specific suggestions/alerts
Highlighting spikes in suggestions/alerts with gray vertical lines, which users can click and drag to zoom in
Showing a Log Line timeline that displays ingested log lines within specified time intervals
Presenting Rare Event timelines that show the number of rare events (potential issues/problems) within time intervals
Displaying visual indicators for spikes (gray lines) and rare events (red lines)
To facilitate incident response, Zebrium integrates with multiple incident management tools through both pre-built integrations and flexible configuration options. Supported platforms include:
Opsgenie
OpsRamp
PagerDuty
VictorOps
These integrations enable several key features:
Automatically adding Root Cause (RCA) reports to incidents in Opsgenie
Providing comprehensive RCA report summaries that include word clouds and detailed log event timelines
Offering direct integration with existing workflow through dedicated webhooks
Supporting both Augment and Auto-Detect operational modes
Augment Mode integrates with third-party tools by:
Triggering Root Cause Analysis when incidents occur
Finding matching anomalous log patterns
Creating detailed Root Cause reports
Transmitting findings via webhooks or existing rules
Auto-Detect Mode continuously monitors logs across all applications using unsupervised machine learning, achieving over 95% accuracy in detecting anomalies. When problems are identified, the system generates suggested alerts containing metadata, root cause reports, and alert rules. Users can accept or reject these suggestions, customize notifications, and manage their environment through various settings controls.
Zebrium offers automated incident response through integrations with Opsgenie, supporting both augment and auto-detect modes. In augment mode, the platform uses its AI/ML engine to analyze application logs and trigger webhooks for root cause analysis when Opsgenie incidents occur. The system identifies anomalous log patterns coinciding with the incidents, creates detailed root cause reports, and summarizes these findings using Opsgenie's notes API. These reports display comprehensive root cause details directly in Opsgenie incidents, allowing users to quickly drill down into the Zebrium user interface for more context.
The auto-detect mode continuously monitors all application logs using unsupervised machine learning, maintaining over 95% accuracy in detecting anomalies. When problems are identified, Zebrium generates root cause reports and automatically sends them to Opsgenie via the webhook interface. These reports display as incidents in Opsgenie, maintaining detailed root cause information and enabling users to quickly correlate logs across the entire application environment.
To enable these integration modes, users must configure several key elements:
API Access Configuration:
Navigate to Opsgenie's Settings > API key management to create a new API key with Read, Create, and Update access. Save the API key for later use.
Opsgenie Integration in Zebrium:
Access Zebrium's User menu > Settings (hamburger) > Integrations > Incident Management > Opsgenie.
Create a new integration, setting the Integration Name, selecting Deployment and Service Group(s), enabling Receive Signals, selecting the Opsgenie region, and entering the saved API key.
Webhook URL Configuration:
In the Opsgenie Integration page, select the Zebrium integration.
Enable Sending alert details to Zebrium for Opsgenie Alerts and paste the saved Zebrium webhook URL.
Save the integration configuration to complete setup.
This configuration process enables seamless integration between Zebrium's root cause analysis capabilities and Opsgenie's incident management workflows.