Home / News & Releases / DII automation

NetApp DII customer-zero automation: evidence and guardrails

NetApp Engineering IT says it runs Data Infrastructure Insights in production across tens of thousands of VMs, petabytes of storage, Kubernetes, and multiple clouds. The strongest evidence is not the scale claim; it is the description of closed-loop workflows and where humans remain in the decision path.

Evidence boundary This is NetApp’s account of its own environment. The post publishes no tenant sizing, alert accuracy, incident reduction, latency, or cost data, so it is a use-case reference—not an independent benchmark.

DII OBSERVABILITY AUTOMATION Sep 22, 2026

Five disclosed production patterns

What an operator should copy—and challenge

The reusable design is the staged response: detect, verify persistence, identify the owner, allow a bounded remediation interval, then take an auditable action. A single instantaneous threshold should not power a destructive response. Use duration, reset thresholds, maintenance suppression, dependency context, and an explicit rollback path to prevent flapping or shutting down the wrong workload.

A snapshot trigger also needs a protection model. Define the protected scope, retention, immutability, replication, capacity reserve, and restore test. A new snapshot taken after suspicious activity begins may preserve evidence, but it does not prove the presence of a known-good recovery point.

A production control sheet

  1. Record the metric, threshold, minimum duration, baseline, scope, exclusions, and reset condition for every rule.
  2. Name the workload owner and platform approver; verify ownership data before enforcement.
  3. Run notification-only first and measure false positives, missed incidents, and time-to-acknowledge.
  4. Cap action rate and blast radius; provide a kill switch and idempotent rollback.
  5. Log input evidence, rule version, decision, action, actor, and outcome into the incident and CMDB record.
  6. Test collector/API/ServiceNow failure so stale or missing telemetry fails safely rather than becoming permission to act.

Bottom line: NetApp’s customer-zero story provides a credible workflow pattern for turning storage and infrastructure telemetry into action. Buyers still need their own scale test and safety case before enabling shutdown or protection automation.

ONTAP troubleshooting · Capacity operations · Ransomware protection

Source

NetApp Engineering IT: DII for proactive observability, automation, and security

← Back to the news index