Home / News & Releases / Database outage during NetApp takeover/giveback due to FC po

Database outage during NetApp takeover/giveback due to FC port queue depth congestion

A public NetApp KB documents Oracle database servers down more than six hours during a planned cluster takeover/giveback after an FC target port on an AFF-A800 exceeded its 1945-command queue-depth threshold.

Field note · October 8, 2026. A public NetApp Knowledge Base article documenting a real ONTAP FC SAN outage. Facts below come from that KB and its companion article; where the KB's remediation steps sit behind a NetApp support login, we say so.

Based on the NetApp Knowledge Base article Database outage during NetApp takeover/giveback due to FC port queue depth congestion. Only claims present in that KB are asserted here.

SAN Fibre Channel ONTAP 9 2026-10-08

What happened

A NetApp Knowledge Base article documents database servers being unavailable for over six hours during a planned NetApp cluster takeover/giveback (TO/GB) maintenance activity. Multiple volumes and LUNs logged I/O errors, and at least one volume was reported offline. Host symptoms were classic SAN-path failures: SCSI I/O errors and multipath path losses.

The environment in the KB is specific: NetApp AFF-A800 systems running FC/FCoE, Red Hat Linux multipath hosts, and Oracle database servers — the kind of stack where a dropped path is felt immediately.

The trigger: an FC target port past its queue-depth threshold

The cause was not a failover bug. It was FC target port queue-depth congestion. The KB records an ONTAP EMS event naming the exact condition:

[CLUSTER01-02:fct_tpd_work_thread_0:scsitarget.fct.port.thresh:notice]:
FC target port 2a has 1946 outstanding commands,
which exceed the threshold 1945 for this port.

An FC target port can hold only a finite number of commands in flight — here 1945. Once the count tips past it, ONTAP stops accepting new commands on that port, and attached hosts see queued I/O fail and paths drop. Our FC/FCP SAN reference covers how ONTAP presents Fibre Channel targets and LUNs.

Why a takeover/giveback makes it worse

During a TO/GB, LUNs and their paths move between HA partners — a routine operation that changes which target port serves the host. If the surviving node's port is already near its queue limit when the traffic lands, the re-routed burst can push it over the threshold. That is how routine maintenance becomes a multi-hour outage. The HA pairs reference explains the failover mechanics; the operational lesson here is that queue headroom on the receiving port matters as much as the failover itself.

What admins will see

The KB lists the host-side signatures to expect, and a companion NetApp article (“Performance Impact due to FCP Queue Depth Threshold Reached”) documents the same failure family. Look for:

On the fabric side, the KB points to congestion on host-facing switch ports (tim_txcrd_z, “Time TX Credit Zero”). Track host-side symptoms with our guide to SAN multipathing, and watch cluster alerts through the AutoSupport & EMS reference.

What to check before your next TO/GB

Before scheduling a takeover/giveback on a busy FC environment: confirm host and target queue depths are sized to what the workload really pushes, keep paths balanced across target ports so no single port absorbs the post-failover burst, and monitor the scsitarget.fct.port.thresh EMS message plus switch-port credit-zero counters. The KB's complete remediation steps sit behind a NetApp support login, so treat the numbers above as the documented symptom — not a substitute for the vendor's fix.

Sources: NetApp Knowledge Base — Database outage during NetApp takeover/giveback due to FC port queue depth congestion; NetApp Knowledge Base — Performance Impact due to FCP Queue Depth Threshold Reached

Bottom line

An FC target port that reaches its outstanding-command threshold during a takeover/giveback can take database hosts down for hours — not because the cluster failed to fail over, but because the target queue saturated. Watch scsitarget.fct.port.thresh and host QUEUE FULL events before every maintenance window.

Outcome metric: indexed news page; first impressions for the story query are UNKNOWN until GSC data is available.

← Back to the news index