Source: NetApp Knowledge Base — “VMware ESXi enters persistent APD during controller maintenance despite healthy FC service” (Reference · Validated · Public).
Troubleshooting SAN VMware ONTAP 9 2026-10-11The symptom: every path gone, storage still up
NetApp’s Knowledge Base documents a scenario where, during maintenance activities such as an ONTAP upgrade or a headswap, a VMware ESXi host loses all of its Fibre Channel paths and its LUNs enter a persistent All Paths Down (APD) state.
The detail that makes it worth reading: the storage side looks completely healthy at the same time. Aggregates, volumes and FC LIFs remain online; the node stays up and is still serving FC; and there is no matching FC logout or SCSI target error in the logs. So the two halves of the fabric disagree — the array reports the target as fine, the host reports that it cannot reach anything.
What the KB scope covers
The article is tagged as applying to:
- ONTAP 9
- FAS / AFF / ASA systems
- Fibre Channel (FC)
- VMware ESXi
NetApp classifies it as Article type: Reference, with the confidence tag Validated and governance Experience — a field-observed, operationally-sourced item rather than a theoretical one.
Why this is a maintenance-window problem, not a steady-state one
The trigger is the maintenance activity itself, which is exactly when teams can least tolerate a host dropping every path. A persistent APD on an ESXi host is not a cosmetic alert: it is the state in which the host has given up waiting for storage to respond and will not automatically resume normal I/O until the condition is cleared. The KB explicitly rules out the two explanations that would normally account for it — initiators were logged in before the APD, and no SCSI target faults were recorded — which is precisely why the article exists.
If you run FC-attached ONTAP behind VMware, this is worth checking against your own upgrade and headswap procedures before the next window.
What the public page does and does not contain
The KB’s opening sections — Applies to and Issue — are public and are what this article summarises. The cause, verification steps and resolution are gated behind a NetApp support sign-in, so this write-up deliberately stops at the symptoms; we neither reproduce nor guess the fix. The page was originally published 2026-09-23 and shows a last-updated date of 2026-10-11, and it remains KCS-enabled and Public.
What to watch next
Because the KB is public and continually updated, the practical next step is to re-read it as the fix sections are fleshed out, and to compare its symptom list against any APD you have seen after an ONTAP upgrade or controller replacement. Related reading on this site: the troubleshooting index, the ONTAP reference shelf, and the news index for other field advisories.