Based on the NetApp Knowledge Base article Database outage during NetApp takeover/giveback due to FC port queue depth congestion. Only claims present in that KB are asserted here.
SAN Fibre Channel ONTAP 9 2026-10-08What happened
A NetApp Knowledge Base article documents database servers being unavailable for over six hours during a planned NetApp cluster takeover/giveback (TO/GB) maintenance activity. Multiple volumes and LUNs logged I/O errors, and at least one volume was reported offline. Host symptoms were classic SAN-path failures: SCSI I/O errors and multipath path losses.
The environment in the KB is specific: NetApp AFF-A800 systems running FC/FCoE, Red Hat Linux multipath hosts, and Oracle database servers — the kind of stack where a dropped path is felt immediately.
The trigger: an FC target port past its queue-depth threshold
The cause was not a failover bug. It was FC target port queue-depth congestion. The KB records an ONTAP EMS event naming the exact condition:
[CLUSTER01-02:fct_tpd_work_thread_0:scsitarget.fct.port.thresh:notice]:
FC target port 2a has 1946 outstanding commands,
which exceed the threshold 1945 for this port.
An FC target port can hold only a finite number of commands in flight — here 1945. Once the count tips past it, ONTAP stops accepting new commands on that port, and attached hosts see queued I/O fail and paths drop. Our FC/FCP SAN reference covers how ONTAP presents Fibre Channel targets and LUNs.
Why a takeover/giveback makes it worse
During a TO/GB, LUNs and their paths move between HA partners — a routine operation that changes which target port serves the host. If the surviving node's port is already near its queue limit when the traffic lands, the re-routed burst can push it over the threshold. That is how routine maintenance becomes a multi-hour outage. The HA pairs reference explains the failover mechanics; the operational lesson here is that queue headroom on the receiving port matters as much as the failover itself.
What admins will see
The KB lists the host-side signatures to expect, and a companion NetApp article (“Performance Impact due to FCP Queue Depth Threshold Reached”) documents the same failure family. Look for:
- Linux —
blk_update_request: I/O erroranddevice-mapper: multipath: ... Failing path, plusqla2xxx ... QUEUE FULL detected. - ONTAP EMS — repeated
scsitarget.fct.port.threshnotices andSTIO TPD cmd alloc threshold reached ... Active commands:1945 threshold:1945. - VMware — NMP path failures and adapter resets; Windows —
SCSI Queue FULL reported on LUN.
On the fabric side, the KB points to congestion on host-facing switch ports (tim_txcrd_z, “Time TX Credit Zero”). Track host-side symptoms with our guide to SAN multipathing, and watch cluster alerts through the AutoSupport & EMS reference.
What to check before your next TO/GB
Before scheduling a takeover/giveback on a busy FC environment: confirm host and target queue depths are sized to what the workload really pushes, keep paths balanced across target ports so no single port absorbs the post-failover burst, and monitor the scsitarget.fct.port.thresh EMS message plus switch-port credit-zero counters. The KB's complete remediation steps sit behind a NetApp support login, so treat the numbers above as the documented symptom — not a substitute for the vendor's fix.
Sources: NetApp Knowledge Base — Database outage during NetApp takeover/giveback due to FC port queue depth congestion; NetApp Knowledge Base — Performance Impact due to FCP Queue Depth Threshold Reached