1700 DevOps Monitoring Experts Agree: Too Many Alerts from Too Many Tools Put Customers at Risk
May 18, 2016

Dan Turchin
BigPanda

We're all technology companies. Every second of downtime hurts. Monitoring at scale is hard. And that's just the beginning of what you shared in our recent survey.

We invited you to tell us about the state of monitoring. Tales of woe and glory from more than 1,700 ops experts provided the most articulate, profound, comprehensive summary of IT Ops life ever assembled.

We thought you'd all benefit from what you shared so we published the results. You represent five continents, large and small companies (modal reply: more than 10,000 employees), large and small teams (modal reply: less than 10 members), and both traditional IT and DevOps organizations.

Here's what fascinated me...

You rely on many tools to monitor your infrastructure.

■ Each team member is responsible for triaging between 10 and 50 alerts per day.

■ In an eight-hour shift, that means you're each working about 10 issues simultaneously assuming you don't inherit orphans from previous shifts (which you do!).

■ Translation: there are fire-swallowing, tightrope-walking, lion tamers working the e.coli route for Carnival Cruise Line with easier jobs than yours.

The more you've invested in agility and velocity, the more effective you are at reducing downtime.

■ Self-described "DevOps" organizations are more than twice as likely to deploy code and/or infrastructure changes at least a few times per day (31% for DevOps orgs vs. 15% overall).

■ They're also more than twice as likely to have cloud-based infrastructure (32% of DevOps orgs vs. 13% overall).

You're dissatisfied with the current reliability of your monitoring and incident management process.

■ Nearly 80% of you say the most challenging part of your job is suppressing alert noise.

■ The problem's not going away: more than 55% are dissatisfied with the current monitoring strategy. Your comments also indicate the problem won't improve in the next 12 months without a better way to manage the growing workload.

A bleak picture perhaps best summarized by Carlos from a midwestern credit union who says if he could change one thing about his organization's current monitoring strategy it would be "to focus on the only thing that matters: reducing noise." Carlos, you're right. Human beings alone can't fix a problem created by machines. We've been in this position before … before there was client-server, TCP, DNS, virtualization, cloud.

We've approached each challenge with the same tenacity, the same passion, the same commitment to solving problems with technology. We'll do it again. This time, with better automation and collaboration. Soon, machines and people will speak a common language. And when they do, we'll be the first to share how great technology plus your ingenuity makes life better for everyone.

Dan Turchin is VP Product at BigPanda.

The Latest

September 20, 2018

The latest Accelerate State of DevOps Report from DORA focuses on the importance of the database and shows that integrating it into DevOps avoids time-consuming, unprofitable delays that can derail the benefits DevOps otherwise brings. It highlights four key practices that are essential to successful database DevOps ...

September 18, 2018

To celebrate IT Professionals Day 2018 (this year on September 18), the SolarWinds IT Pro Day 2018: A World Powered by Tech Pros survey explores a "Tech PROactive" world where technology professionals have the time, resources, and ability to use their technology prowess to do absolutely anything ...

September 17, 2018

The role of DevOps in capitalizing on the benefits of hybrid cloud has become increasingly important, with developers and IT operations now working together closer than ever to continuously plan, develop, deliver, integrate, test, and deploy new applications and services in the hybrid cloud ...

September 13, 2018

"Our research provides compelling evidence that smart investments in technology, process, and culture drive profit, quality, and customer outcomes that are important for organizations to stay competitive and relevant -- both today and as we look to the future," said Dr. Nicole Forsgren, co-founder and CEO of DevOps Research and Assessment (DORA), referring to the organization's latest report Accelerate: State of DevOps 2018: Strategies for a New Economy ...

September 12, 2018

This next blog examines the security component of step four of the Twelve-Factor methodology — backing services. Here follows some actionable advice from the WhiteHat Security Addendum Checklist, which developers and ops engineers can follow during the SaaS build and operations stages ...

September 10, 2018

When thinking about security automation, a common concern from security teams is that they don't have the coding capabilities needed to create, implement, and maintain it. So, what are teams to do when internal resources are tight and there isn't budget to hire an outside consultant or "unicorn?" ...

September 06, 2018

In evaluating 316 million incidents, it is clear that attacks against the application are growing in volume and sophistication, and as such, continue to be a major threat to business, according to Security Report for Web Applications (Q2 2018) from tCell ...

September 04, 2018

There's a welcome insight in the 2018 Accelerate State of DevOps Report from DORA, because for the first time it calls out database development as a key technical practice which can drive high performance in DevOps ...

August 29, 2018

While everyone is convinced about the benefits of containers, to really know if you're making progress, you need to measure container performance using KPIs.These KPIs should shed light on how a DevOps team is faring in terms of important parameters like speed, quality, availability, and efficiency. Let's look at the specific KPIs to track for each of these broad categories ...

August 27, 2018

Protego Labs recently discovered that 98 percent of functions in serverless applications are at risk, with 16 percent considered "serious" ...

Share this