DII OBSERVABILITY AUTOMATION Sep 22, 2026
Five disclosed production patterns
- Rogue-VM containment: sustained abnormal IOPS triggers an alert, owner notification, a remediation window, and then automatic shutdown if the condition persists.
- Hybrid inventory: DII consolidates consumption data from on-premises VMware and OpenShift plus AWS, Azure, and Google Cloud. NetApp explicitly says DII is not a billing or chargeback tool.
- Configuration and security: the team checks encryption, volume configuration, and growth policies, monitors file activity, and says suspicious patterns can trigger snapshots.
- Kubernetes automation: teams query inventory and performance data through APIs and feed it into existing internal workflows.
- CMDB population: VM metadata, relationships, and ownership flow into ServiceNow to improve change and incident correlation.
What an operator should copy—and challenge
The reusable design is the staged response: detect, verify persistence, identify the owner, allow a bounded remediation interval, then take an auditable action. A single instantaneous threshold should not power a destructive response. Use duration, reset thresholds, maintenance suppression, dependency context, and an explicit rollback path to prevent flapping or shutting down the wrong workload.
A snapshot trigger also needs a protection model. Define the protected scope, retention, immutability, replication, capacity reserve, and restore test. A new snapshot taken after suspicious activity begins may preserve evidence, but it does not prove the presence of a known-good recovery point.
A production control sheet
- Record the metric, threshold, minimum duration, baseline, scope, exclusions, and reset condition for every rule.
- Name the workload owner and platform approver; verify ownership data before enforcement.
- Run notification-only first and measure false positives, missed incidents, and time-to-acknowledge.
- Cap action rate and blast radius; provide a kill switch and idempotent rollback.
- Log input evidence, rule version, decision, action, actor, and outcome into the incident and CMDB record.
- Test collector/API/ServiceNow failure so stale or missing telemetry fails safely rather than becoming permission to act.
Bottom line: NetApp’s customer-zero story provides a credible workflow pattern for turning storage and infrastructure telemetry into action. Buyers still need their own scale test and safety case before enabling shutdown or protection automation.
ONTAP troubleshooting · Capacity operations · Ransomware protection