AI ONTAP NFS pNFS Oct 02, 2026
What Novus actually is
Novus is a scale-out parallel file system built on top of ONTAP. The design keeps a single, standards-based NFS namespace and moves the metadata service into a dedicated tier instead of running it on the same controllers that serve data. In the first supported configuration:
- Novus Data Director — NetApp’s metadata software — runs on qualified Supermicro x86 servers. It answers namespace and layout requests (the product diagram lists
OPEN,LAYOUTGETandLAYOUTCOMMITon the Data Director path). - AFF A90 HA pairs running ONTAP act as Novus data nodes, carrying capacity and bandwidth. NetApp says the first data nodes are a specific AFF A90 variant tuned for sequential bandwidth with additional NVRAM.
- Clients are stock Linux pNFS clients on the GPU servers’ front-end NICs, using NFSv4.2 with pNFS Flex Files,
nconnectand GPUDirect Storage, per NetApp’s product page — there is no proprietary client to install.
The data path rides a lossless Ethernet client fabric that is deliberately kept separate from the GPU backend fabric used for collectives; NetApp’s diagram makes the point that Novus is not in that backend path. Hosts mount once and data nodes can be added without remounting, so performance, capacity and metadata concurrency are meant to scale independently — the property that matters for multi-tenant AI factories run by neocloud and GPU-as-a-Service providers.
Why the metadata tier is the interesting part
In a briefing ahead of INSIGHT, NetApp VP of Technical Marketing Jeff Baxter described Data Director as answering lookups and layout requests out of band and in memory, so every GPU client can address every data node directly without being proxied through a cluster interconnect. That is the behaviour of a parallel file system — without a Lustre-style agent running on the client. For ONTAP shops, the implication is that the thing historically hard to scale for AI training — metadata throughput and small-file create/lookup rates — is no longer bound to the data controllers’ CPU. The broader mechanics of ONTAP metadata behaviour are covered in the performance hub.
NetApp frames the problem in GPU-utilisation terms: “Under traditional architectures, GPU utilization can drop below 30 percent when AI factories cannot feed enough data to their GPUs,” the release says. Baxter’s summary was blunter: “Probably the most expensive asset in any data center in the world is an unused GPU in an AI factory.”
Orderable versus projected
Two claims need to be kept apart.
Orderable. NetApp says the first hardware-bound configuration — Data Director on qualified Supermicro servers plus AFF A90 data nodes — is orderable today, and that the initial design “provides a path to software-defined deployments over time.” Baxter said the intent is to get Novus into the hands of lighthouse customers within a month or two of INSIGHT, with wider rollout to AI-factory customers in the months after. NetApp estimates the target population at roughly a hundred companies running tens of thousands to hundreds of thousands of GPUs.
It is also worth noting the AFF A90 is the current high-end platform, so the data-node tier is not a new array — it is new software plus a tuned build.
Projected. The 100 TB/s aggregate throughput and “dozens of exabytes” of effective capacity are not audited production results. NetApp says tests audited by Omdia showed near-linear scaling as ONTAP clusters were added to a single global namespace, and that Omdia’s modelling projects the architecture to 100 TB/s of sequential read. Omdia chief analyst Tony Palmer said: “We are confident in our projections that the solution can scale to the needs of zettascale deployments.” NetApp’s own sizing works out to more than 50,000 GPUs at about 2 GB/s each for the 100 TB/s figure. Treat “the fastest storage on the planet” — NetApp’s phrasing — as a marketing claim resting on that model, not as a measured result.
The design is also said to preserve ONTAP data services, security, QoS and multi-tenancy while supporting billions of files across thousands of storage nodes. Fine as a claim; it is exactly where the verification work lands.
Operational consequence for ONTAP admins
Novus is not aimed at a general-purpose enterprise ONTAP estate, and nothing in the announcement implies an upgrade to existing clusters. It is a platform decision for large AI factories. If you are evaluating it, the disaggregation changes the failure and sizing model in concrete ways:
- You now have two failure domains. The metadata tier (Data Director on third-party x86) and the data tier (AFF A90 pairs) fail and recover independently. Ask for the Data Director’s HA model, its recovery time, and what a client sees when the metadata tier is unavailable but data nodes are up.
- The client fabric is a new network to design. A dedicated lossless Ethernet fabric separate from the GPU backend fabric means separate switching, MTU, PFC/ECN and monitoring. Budget for it and stage it — the classic AI-factory networking pitfalls are mapped in the network architecture reference.
- Where do data services enforce? NetApp says QoS, multi-tenancy, security and SnapMirror/SnapVault/SnapLock-class protection are preserved. Confirm whether each is enforced on the data nodes, the metadata tier, or both — especially for multi-tenant QoS isolation and immutable snapshots.
- Verify the client support matrix. pNFS Flex Files with
nconnectand GPUDirect Storage depends on specific Linux kernels, NIC firmware and driver versions on the GPU hosts. Get the validated matrix before you commit host images. - Get behind the headline number. Ask for the audited test’s scale, file-size mix and client count, and the model’s assumptions. Sequential bandwidth projections say little about the metadata-heavy small-file patterns that dominate real training data pipelines.
- Track the software-defined path. The x86-agnostic ONTAP build “to follow” decides whether Novus is a closed appliance or deployable on your own qualified hardware. That timeline is uncommitted.
For context on the technology NetApp is folding in, see our earlier piece on the PEAK:AIO parallel-NFS-with-ONTAP results; NetApp agreed to acquire PEAK:AIO for its scale-out pNFS metadata work a week before INSIGHT. Unfamiliar terms such as pNFS, Flex Files and nconnect are defined in the ONTAP glossary.
Bottom line
Novus is the most architecturally interesting thing NetApp announced at INSIGHT 2026: a real disaggregation of the metadata path, built on standard NFSv4.2 pNFS rather than a proprietary client, delivered on ONTAP arrays that existing teams already run. The first configuration is orderable; the throughput headline is a projection. Put Novus on the architecture watchlist for AI-factory builds, not the change calendar for existing clusters, and make every decision against the validated support matrix and independent benchmark detail rather than the 100 TB/s claim.
Read the source (StorageReview) · AFF A90 reference · Network architecture · Performance hub