Home / News & Releases / AI Data Engine discovery

NetApp AI Data Engine now discovers storage beyond ONTAP

At NetApp Insight 2026, NetApp said its AI Data Engine will now discover metadata across ONTAP, StorageGRID and storage from other vendors, over NFS, SMB and S3. The earlier version worked only with data on NetApp systems. The change matters less for what it indexes than for what it removes: the pipeline no longer has to move data before it can reason about it.

Original reporting NetApp Black Box analysis, 2026-10-01. Claims are attributed to NetApp or to the reporting outlet, and roadmap items are labelled as such. Back to the news index.

AI CLOUD DATA MANAGEMENT Oct 01, 2026

What changed

NetApp's Vice President of Technical Marketing, Jeff Baxter, briefed press at Insight 2026 on a widened scope for the AI Data Engine. Where the first version discovered only data hosted on NetApp systems, the company now says the engine will discover metadata across ONTAP, StorageGRID and storage from other vendors, using NFS, SMB and S3. NetApp says customers get the capability at no charge for at least the first six months.

NetApp's Senior Vice President and General Manager for AI, Asad Khan, framed the rationale: “For too long, AI initiatives have depended on complex pipelines that move data away from where it is created, governed, and protected.”

Baxter also cited a Gartner projection that 40% of enterprise applications will embed task-specific AI agents this year, and that agents will handle 15% of daily business decisions by 2028 — after which, he argued, cost per token becomes the binding constraint. Discovery over heterogeneous storage is NetApp's answer to that constraint: index the metadata where the data already sits instead of copying it into a separate AI store.

Why discovery without movement is the point

Most enterprise AI data problems are governance problems wearing an engineering costume. The moment data is copied to a new location, its snapshot policy, retention class, quota, access controls and audit trail have to be re-created — and usually are not, faithfully. By discovering metadata in place over NFS, SMB and S3 rather than ingesting the data, the existing protection and capacity regime stays attached to the original copy.

That is why the storage-layer controls already configured on the source system keep doing the work. Protection policies built on ONTAP Snapshot copies remain the recovery baseline, and quota management remains the mechanism that stops one tenant's AI scratch space from consuming a shared aggregate. Discovery that respects those boundaries is only useful if the boundaries are actually defined first — which is the part most organisations skip.

The rest of the Insight 2026 wave

The discovery change did not arrive alone. The same briefing covered four other threads that administrators should read together:

Announced versus shipped

The honest read is that this was a platform and positioning event, not a release-notes event. No ONTAP version, minimum source release, upgrade path, CVE or known-issues list was attached to the AI Data Engine discovery change, and the multi-vendor scope comes from a press briefing rather than from published administration documentation. Treat the availability language — “now”, “orderable”, “this fall”, “within 12 months” — as four different levels of commitment, because they are.

No cluster upgrade is implied by any of it. Existing ONTAP systems should not be upgraded, reconfigured or licensed solely because of this announcement.

What hybrid multicloud admins should validate

  1. Confirm the third-party support matrix. “Other vendors' storage” is not a list. Before designing around discovery, obtain the specific arrays, gateways and object stores supported, the protocol versions, and the limits on namespace size and metadata volume — in writing.
  2. Decide where the metadata lives. A catalogue that spans on-premises ONTAP and two clouds raises residency and tenancy questions. Map which metadata may leave which jurisdiction before enabling discovery broadly; the same question is already live for Keystone Sovereign customers. The placement trade-offs are mapped in the hybrid cloud hub.
  3. Fix governance before you scale discovery. Snapshot policies, quota classes, retention and access controls should be consistent across every source the engine touches, or the catalogue will faithfully describe a mess at machine speed. NetApp's own framing of sovereignty — isolation through multi-tenancy, separating departments by access, routing and data management — only holds if those separations exist on every source.
  4. Measure the claim you are being sold. The pitch is that leaving data in place cuts cost per token and avoids pipeline sprawl. Baseline the current ingest cost, pipeline complexity and time-to-first-query on one representative workload, then compare against discovery-in-place before committing.
  5. Test the exit and the audit path. Confirm what the catalogue records, how it is exported or deleted, and how it behaves when a source array is decommissioned or a cloud account is closed.

Operational consequence

If the delivered capability matches the briefing, the practical shift is that storage administrators gain a second audience — AI platform teams — without surrendering their data. The catalogue becomes the contract between them: what exists, where it is governed, and what it is allowed to feed. That is a more durable role than competing with a vector database.

Bottom line: the interesting part of Insight 2026 was not raw throughput but the decision to index metadata across heterogeneous storage instead of copying data into an AI silo. Put multi-vendor AI Data Engine discovery on the evaluation list, not the change calendar, and require the support matrix and metadata-residency answers before you enable it on anything production.

Read the Insight 2026 briefing coverage (CIOL) · Hybrid cloud hub · Quota management reference · Snapshot management reference

Sources and corroboration

CIOL: NetApp Unveils AI Storage Architecture, Oracle Cloud Service And Agent Controls (Insight 2026 press briefing) · NetApp blog: unified data management for hybrid multicloud AI · NetApp blog: AI data services discovery · NetApp newsroom

Figures for throughput, scaling and detection accuracy are NetApp's or its cited analyst's own claims and are reported as such. The multi-vendor discovery scope, the no-charge window and the four availability statements come from the Insight 2026 press briefing; NetApp administration documentation for the widened scope was not available at the time of writing.

← Back to the news index