Home / Troubleshooting

Troubleshooting Guides

Recent issues worth reading

Incident-style coverage and release notes that feed these runbooks.

Auto-generated from the content index · 6 items · newest 2026-10-07

The problems that actually show up in the field, with the diagnosis path and the commands that resolve them. Each guide: symptoms → checks → fix → prevention.

Troubleshooting guides
NetApp FAS controller used in hardware troubleshooting
Start with the controller and shelvesJemimus · CC BY 2.0 · Wikimedia Commons

Start with a topic hub

secd Unresponsive or Authentication Timeouts

Users say SMB logons, LDAP lookups, or name mapping suddenly hang or fail; EMS shows secd messages.

wafl.vol.outOfInodes: Volume Cannot Create Files

Applications can read existing data but cannot create new files; ONTAP reports wafl.vol.outOfInodes.

vifmgr.reach.err: LIF Reachability Problems

A data or cluster LIF is up, but clients or peer nodes cannot reliably reach it; EMS reports vifmgr.reach.err.

NFS Latency: Slow Reads, Writes, or Metadata

NFS mounts work, but users report slow opens, reads, writes, directory listings, or intermittent stalls.

SnapMirror Latency or Lag Is Growing

The destination is hours behind, transfers take too long, or the protection dashboard shows growing SnapMirror lag.

Cluster LIF Problems: Node or Ring Connectivity

A node is unhealthy or out of quorum, cluster commands are slow, or cluster LIFs/ports appear down.

Low-Memory EMS Warnings

ONTAP raises low-memory or allocation warnings, and management or data services may feel slow or unstable.

Aggregate Offline: Volumes Are Unavailable

An aggregate/local tier is offline and its volumes are unavailable, often after power, disk, shelf, or HA disruption.

LUN Not Visible to the Host

The LUN exists in ONTAP, but the host cannot discover it or has lost all usable paths.

Volume space exhaustion

"Volume is full": snapshots holding deleted data, overcommitment, inode pressure, and how to reclaim space safely.

Performance analysis

Latency is up. Find the bottleneck: node CPU, disk utilization, protocol ops, and per-volume latency via sysstat and QoS statistics.

Network & connectivity problems

Troubleshoot LIFs down, no route to host, MTU and jumbo frames, ifgrps, iSCSI paths, and packet-tracer workflows.

LIF failover troubleshooting

Diagnose wrong or missing failover targets, non-home NAS LIFs, unsafe reverts, broadcast-domain mistakes, and client reachability without confusing SAN multipathing.

SnapMirror problems

Growing lag, failed transfers, stuck relationships — and how to break, resync, and restore without losing data.

Backup restore troubleshooting

Recovery-first triage for local Snapshots, SnapMirror vaults, NDMP and configuration backups—with safe restore paths and escalation boundaries.

NFS mount failures

Permission denied, RPC errors, and hangs: export policies, LIFs, routing, sec flavors, and the NFSv4 domain.

SMB/CIFS deep-dive

Shares, share ACLs vs NTFS ACLs, AD/Kerberos integration, SMB 3.x features (multichannel, encryption, CA), and the access-denied / auth-failure paths — with real commands.

ONTAP S3 object storage troubleshooting

AWS SigV4 signature mismatches, 403 clock skew, bucket policies, TLS cert handshakes, incomplete multipart capacity drains, and SnapMirror S3 replication lag.

MetroCluster DR failover

When a site dies: switchover and switchback workflows, mediator status checks, and what to verify after the dust settles.

MetroCluster troubleshooting

metrocluster check errors, mediator quorum, blocked switchover/switchback, IP-vs-FC link checks — the full triage runbook with the commands that matter.

NFSv4 ACL migration (POSIX → NFSv4)

Why ONTAP stores NFSv4 ACLs as NTFS ACLs, the NIS → AD/Kerberos identity move, and a phased runbook with nfs4_setfacl and vserver security file-directory commands.

ONTAP upgrade troubleshooting & recovery

ANDU pre-validation errors, LIF migration vetoes, paused rolling upgrades, mixed-version cluster recovery, and image rollback.

Aggregate offline after power loss

Dirty shutdowns and DC power cycles: safe read-only triage, "UUID is not valid" scares, what NOT to run, and when to call support — before touching anything.

Decommissioning & repurposing ONTAP

Retiring nodes, shelves, or whole clusters: config and license backup, clean removal, stuck audit volumes, disk sanitization, and used-shelf repurposing rules.

ONTAP with Proxmox & non-VMware hypervisors

Moving off VMware: what you lose (VASA, SRM, SnapCenter, Veeam snapshots), NFS nconnect tuning, iSCSI multipath with ALUA, and the backup/DR rebuild without VMware glue.

Before you start Every guide assumes you can reach the cluster: SSH to a node or the cluster management LIF. Use set -privilege advanced where noted. Run man <command> on the system for flag details — commands below target ONTAP 9.x.