Overview
An incident represents an anomaly in the Whatfix software system that causes service degradation or an outage. The goal of incident management is to resolve incidents quickly and restore services promptly. Whatfix implements a dedicated process to manage these incidents, ensuring swift resolution and restoration of normal service.
What Does This Apply To?
This process applies to all services provided by Whatfix.
Key Terms
SRE (Site Reliability Engineering): The first person to be alerted when an incident occurs.
Success Engineer: The person responsible for coordinating and resolving the incident.
Customer Success Engineering Manager: A leader who coordinates with multiple teams to resolve the issue quickly. They are also responsible for communicating updates about the incident.
Alert: A potential anomaly identified and notified through monitoring systems.
Incident: An alert that could potentially disrupt Whatfix services.
SLA (Service Level Agreement): Contractual obligations for a specific level of service.
Post Mortem Report: A detailed internal analysis of the incident.
RCA (Root Cause Analysis): An analysis of the cause and actions for a specific incident.
On-call engineer: Service owners who are available on a rotation basis to help with faster issue resolution.
The Incident Response Team
When an incident is identified, a team comes together to address it.
This team includes:
Customer Success Engineering Manager
Success Engineers
SRE
Engineering Leads
Customer facing teams
Roles and Responsibilities
Different teams play different roles in managing incidents:
Whatfix Management Team: Reviews and approves procedures, ensures all staff members receive training on these procedures, and regularly reviews critical Incident Reports and actions.
SRE: Responds to all critical alerts, identifies and addresses incidents, assesses and classifies incident severity, and escalates impacting incidents as Problems when necessary.
Customer Success Engineering Manager: Manages all customer-impacting incidents and oversees internal communications during incidents.
Engineering Teams: Support and participate in the quick resolution of incidents.
Incident Management Process
Identification
Incidents can be detected in two ways:
Identified through system monitoring
Reported by customers
In case of a system detected incident, the status page is directly updated. If customers report it, the status page is updated after validating if the impact is widespread.
Classification
Incidents receive internal prioritization and classification based on their impact, as outlined in our SLA. Widespread issues receive labels for proper tracking and communication.
Priority Levels and SLA
Once the issue has been prioritized, each of the priority levels has a different workflow within JIRA for managing the incident. For more information, see Whatfix Service Level Agreement (SLA).
Actions
The SRE attempts to manage the incident using a runbook. If the runbook fails, the SRE brings in the respective response team to address the incident. Whatfix includes engineers from each responsibility area in incident resolution. War rooms, chat channels, and virtual meetings facilitate quick incident resolution. The incident's progress updates periodically to all customer-facing teams and the status page.
Communication During an Incident
Keeping stakeholders and customers informed about the nature, status, and progress of incidents is critical. Whatfix has clearly defined SLAs for updates through these channels based on the priority. Modes of communication include status page updates, support tickets, emails, and RCAs.
After the Incident
RCA
Once an incident is resolved, an RCA is available. It captures the incident summary, timeline, corrective and preventive actions for an incident. Incident owners use templates to outline their analysis. These are reviewed by the Whatfix management team for correctness and a clear representation of problems and solutions. Once reviewed, the RCA is available to the customer.
Post Mortem Report
The owner of the incident prepares a Post Mortem report to capture in great detail the lead up to the incident with detailed RCA, references, and lessons learned from the incident. These are recorded and acted upon by the respective teams.
Incident Review Meetings
Incident review meetings are held where post mortem reports are thoroughly reviewed with key stakeholders. Action items are identified to ensure such incidents do not occur in the future.