fix(linux): one logging-service alert per host - #2789
Merged
Merged
Conversation
"Audit or Logging Service Disabled" grouped by origin.host and origin.user, which an alert never carries (the alert holds the event's origin side as adversary). Grouping keys that do not resolve are skipped, so every matching record opened a new top-level alert. The rule now raises one alert per host (dataSource) and drops repeats for seven days. Its condition is unchanged. linux_alert_volume_test.go pins the key and checks that it resolves on fabricated journald records. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
"Audit or Logging Service Disabled" grouped by
origin.hostandorigin.user. An alert never carriesorigin.*(the alert plugin stores the event's origin side asadversary), and grouping keys that do not resolve are skipped, so every matching record opened a new top-level alert. On one v11 deployment the rule stored 335 alerts in 30 days, up to 88 in a day and 10 in one hour, and the rule flood guard can switch it off.Noise is reduced here only by de-duplication. The condition is unchanged.
What changes
groupByorigin.host, origin.user (never resolve on an alert)deduplicateBydataSource: one alert per host, repeats dropped for seven daysExpected volume
Every match was stored as its own alert, so the rule's alerts over the last 30 days on the v11 deployments are its matches (341 on three deployments). Replayed with the new key and seven-day de-duplication: 9 alerts in 30 days (6, 2 and 1), worst hour 1, worst day 1.
What the condition matched is worth a separate review: kernel out-of-memory messages that name the journald or rsyslog cgroup (
oom-kill: … cpuset=systemd-journald.service), rsyslog reloads after log rotation (systemctl kill -s HUP rsyslog.service), package scripts that unmask rsyslog (unmaskcontainsmask), journald watchdog restarts and a Filebeat status line that lists process capabilities.Tests
plugins/alerts/linux_alert_volume_test.go(new) runs fabricated journald records through the Linux raw model and the pinned go-sdk v1.1.36 CEL. It pins the de-duplication key, checks that it resolves, and checks that stop, disable and mask of a logging service match while a restart and an unrelated message do not.go test ./...inplugins/alertspasses on this branch, which is based on currentv11(89cd26c4).8a3ade7, go-sdk v1.1.36, the production events and alerts plugins, OpenSearch 2.19.1, this branch's Linux filter from currentv11) with fabricated journald records in two runs:systemctl stop auditd, a restart and an unrelated message on one host, thensystemctl disable rsyslog.serviceon that host andsystemctl stop auditdon a second host. Both runs processed all of their events.origin.hostandorigin.userresolved on none of them.🤖 Generated with Claude Code