One of the most common assumptions in engineering leadership is that a rising number of reported incidents signals declining system reliability. However, a recent article from Great Circle argues that the opposite is often true: an increase in incident counts may actually indicate that an organization's incident management culture is improving. As teams invest in better processes, tooling, training, and operational discipline, they become more willing to formally declare incidents that previously would have been handled informally or hidden altogether. The result is greater visibility into operational issues, not necessarily a deterioration in system health.
The article challenges the widespread practice of using incident volume as a key performance indicator. Incident count, it argues, measures an organization's willingness to surface and manage operational problems rather than the underlying reliability of its systems. Mature incident management cultures encourage engineers to declare incidents early, involve the appropriate stakeholders, and conduct structured post-incident reviews. While this often causes reported incidents to increase initially, it also creates more opportunities to learn, improve processes, and prevent larger outages in the future.
The phenomenon is familiar across other engineering disciplines. Organizations that strengthen vulnerability management frequently report more security findings, not because their systems have suddenly become less secure, but because they have improved their ability to detect weaknesses. Likewise, introducing better observability often leads to more alerts, while enhanced testing uncovers more defects before software reaches production. In each case, improved measurement exposes existing problems that were previously invisible rather than creating new ones.
Incident management follows the same pattern. Engineers who previously resolved "spicy bugs" or degraded services quietly may instead choose to declare formal incidents, triggering coordinated response processes, documentation, and postmortems. Although dashboards show more incidents, the organization has actually become more resilient by making operational knowledge visible and repeatable.
This perspective aligns with broader thinking across the reliability engineering community. Modern SRE practices have steadily shifted away from simplistic operational metrics toward measures that reflect customer impact, recovery effectiveness, and organizational learning. Shifting questions from "How many incidents occurred?" to "How quickly were users affected?" "How rapidly was service restored?" "Did we identify the root cause?" or "Are similar failures becoming less frequent?"
Recent guidance from Sygnia similarly argues that many traditional incident metrics, including raw incident counts, ticket volumes, and even mean time to respond when viewed in isolation, can create false confidence because they measure operational activity rather than organizational readiness. Instead, it recommends focusing on metrics such as containment effectiveness, escalation quality, post-incident improvements, and the maturity of incident response processes.
Likewise, modern reliability practices built around Service Level Objectives (SLOs) increasingly emphasize user-centric measures such as error budget burn, SLI degradation, and customer impact over infrastructure-centric statistics. Reliability platforms argue that understanding how users experience failures provides a much more accurate picture of service health than simply counting the number of incidents or measuring infrastructure uptime alone.
Perhaps the most important message from Great Circle is cultural rather than technical. Organizations should avoid discouraging incident declarations simply to improve dashboard metrics. If engineers believe they will be judged on keeping incident counts low, they may delay declaring incidents, attempt to resolve problems alone, or avoid escalating emerging issues until they become significantly worse. Such behaviors reduce organizational visibility precisely when rapid collaboration is most needed.
Instead, healthy engineering organizations reward transparency. They recognize that declaring an incident is not an admission of failure but the beginning of a structured learning process. By encouraging early reporting, blameless collaboration, and continuous improvement, organizations create an environment where operational knowledge accumulates over time rather than remaining trapped within individual responders.
The broader lesson is that metrics should reflect the outcomes organizations actually care about. A falling incident count may indicate improving reliability, but it could equally suggest under-reporting, inconsistent classification, or a weakening incident culture. Conversely, a temporary increase in reported incidents may represent healthier operational practices and better organizational awareness.