Cybersecurity is about small, targeted moves: 20% of the right actions prevent 80% of the risk. Find your 20%: assess your exposure →

04. Automation & AI

AIOps: fewer alerts, and the right ones

How you get from thousands of ignored daily alerts to usable detection, and what automation must never decide on its own.

Key points

  • The problem on critical infrastructure is not a shortage of alerts: it is their number.
  • An alert nobody acts on is worse than no alert, because it creates the illusion of monitoring.
  • Thresholds learned from real behaviour outperform thresholds set by hand.
  • Any automated remediation runs within an agreed scope, and every action is logged.

Why adding a monitoring tool solves nothing

Because alert volume is a design problem, not a tooling problem. An extra tool produces extra alerts, on the same events, seen from another angle. Teams end up filtering mentally, and incidents keep being reported by users before monitoring catches them.

The mechanism is well documented: past a certain volume, an alert stops being a signal and becomes background noise. That threshold is reached far sooner than people expect.

The work therefore consists of reducing, not adding. Correlate the events describing one cause, remove those that trigger no action, and surface what has a consequence for a delivered service.

The test is simple and unforgiving: if an alert never triggers an action, it should not exist.

How do you improve the signal-to-noise ratio?

Through correlation, learned thresholds, and a change in the unit of observation. You group events describing a single incident, replace hand-set thresholds with thresholds learned from real behaviour, and monitor delivered services rather than devices.

  • Correlation. A failed link produces dozens of events across as many devices. They describe one incident and must surface as one.
  • Learned thresholds. A threshold fixed at 80% utilization is arbitrary: on some links that is normal, on others it is already an anomaly. A threshold learned over several weeks accounts for business cycles.
  • Service-level monitoring. "The link to site B is saturated" calls for no decision from a management team. "Billing is unusable from site B" does.
  • Suppression. The least practised and most productive step. Any alert that has triggered no action in months is removed, or converted into an indicator.

What must automation never decide alone?

Anything not agreed with you in advance. Automated remediation runs within a written scope (which actions, on which devices, under which conditions)and every execution is logged with its trigger, its rule and its outcome. Anything outside that scope requires human validation.

This constraint is not excessive caution: it is an auditability requirement. In the sectors we serve, you must be able to explain after the fact why an action took place, and demonstrate that it fell within an agreed framework.

It also guards against the most expensive failure mode of unbounded automation: the corrective action that worsens the incident. Restarting a service, isolating a device or blocking a flow are useful actions, and destructive at the wrong moment.

We therefore always start with reversible, low-impact remediations, and widen the scope as confidence in the setup is built on evidence.

What data feeds these analyses?

Technical infrastructure data: telemetry, device logs, network flows, configuration state. No personal information enters without explicit contractual framing and a prior privacy impact assessment.

The distinction matters, because infrastructure logs sometimes contain user identifiers, addresses or file names. Treating that content as plain technical data would be a mischaracterization.

Before any deployment we therefore establish what is collected, for what purpose, and for how long. That is also what allows you to answer an auditor without having to reconstruct the information.

Where that processing physically runs raises the same question, and it is the subject of our sovereign AI approach.

Frequently asked questions

No. It reduces the volume to review by isolating, among a very large number of events, those that deserve attention. Qualification, decision and response remain human and documented, which is also an auditability requirement in regulated sectors. What changes is the time available for analysis.

Every automated action is logged with its trigger, the rule applied and its outcome. The scope of permitted remediations is written and agreed with you before activation, and anything outside that scope requires human validation. That framework is reviewed and amended: it is not frozen at deployment.

Rarely. Most tools in place already collect far more data than they exploit. The work targets correlation, removal of useless alerts, and the change in unit of observation: three efforts requiring no new tool. We recommend a change only where the gap is documented.

They must cover the organization's real business cycles, which varies by trade: a monthly close, a season, a production peak. We do not set a duration in advance. During that period manual thresholds stay in place: learning is observed before anything is replaced.

The principle yes, the tooling no. A smaller business does not have thousands of daily alerts; it has the opposite problem: nothing surfaces at all. The work then consists of instrumenting the essentials: backups, access, availability of critical services, rather than filtering an existing stream. That is what the ongoing engagement covers.

Sources

Is your infrastructure ready for the next threat?

An initial assessment, free and without commitment, to evaluate your security posture.

Home Expertise RISS 360 PME Assess