Enterprise AI agents need current, governed business data. That requirement has pushed 'zero copy' and 'zero ETL' into product announcements, architecture diagrams, and board-level AI discussions.

The labels sound decisive. The underlying designs are not.

One platform may use zero copy to mean a live query against data that stays in its source. Another may expose shared files through a common catalog. A service marketed as zero ETL may continuously replicate tables into a warehouse while removing the need to build the pipeline yourself. These patterns solve different problems.

The decision for a CIO or data leader is therefore not whether zero copy is better than replication in the abstract. It is which pattern gives each agent the required freshness, control, performance, resilience, and evidence.

That question became more urgent in September 2026. Salesforce and AWS expanded Data 360 zero-copy access across several AWS data services, while Salesforce and Google Cloud announced broader regional sharing and support for open catalog standards. The direction is clear: enterprise agents will increasingly reach across platforms. Architecture teams still have to decide how that access works for each workload.

Core Architectural Finding: Zero copy is an access pattern, not an all-encompassing data strategy. Enterprise AI agents require an intentional balance between in-place query federation, managed CDC replication, and curated analytical datasets.

Start With the Mechanism: Four Core Patterns

Use four working definitions during vendor evaluation to strip away marketing ambiguity:

  • In-place federation: The consumer queries data through a connector while the source remains authoritative. The query may be pushed to the source, and the consumer receives the result.
  • Shared-storage access: Different engines read the same underlying files or tables through a common format (like Apache Iceberg) and catalog. The data is stored once, although engines may still create caches, indexes, logs, or derived results.
  • Managed replication or change data capture: Copies an initial dataset and then applies changes to a target continuously or in micro-batches. AWS defines its RDS zero-ETL integration as a managed pipeline that replicates transactional data and schemas.
  • Batch ingestion: Copies data on a schedule and usually transforms it into a target model for reporting, analytics, or machine learning.

Ask vendors to place every relevant data flow in one of these categories. If the answer is 'zero ETL,' ask again. You need the mechanism, the storage location, the refresh behavior, and the control boundary.

PatternWhat HappensStrong FitMain Tradeoff
In-place federation or sharingThe consumer queries a governed source or shared storage through a referenceFresh reference data, exploratory analysis, cross-organization sharing, reduced duplicationSource dependency, query performance, read-only limits, cross-platform policy alignment
Managed replication or CDCA service copies changes into a target continuously or on a scheduleOperational analytics, performance isolation, local transformations, replay, downstream joinsAdditional copy, lag, schema drift, reconciliation, duplicated controls
Batch ingestionData is copied at defined intervals and transformed for a target modelStable reporting, historical analysis, reproducible snapshots, lower urgencyStaleness and pipeline maintenance overhead
Hybrid architectureEach data domain uses the pattern that fits its operational requirementEnterprise agents spanning reference, transactional, historical, and sensitive dataRequires unified cataloging, strict governance ownership, and pipeline observability

Apply Six Architecture Tests

Test 1: Determine Whether the Agent Reads or Acts

A read-only agent answering a policy or account-status question may work well with live, governed access. An agent that creates an order, changes a beneficiary, or updates a case needs an authorized write path into the system of record.

Do not hide write behavior inside a broad 'data access' design. Document the read source, the action API, the identity used for each call, and the approval or confirmation required. Microsoft's guidance for enterprise agents makes the same distinction: retrieval may use search, while real-time queries and actions can require authenticated tools or Model Context Protocol servers.

Test 2: Set a Measurable Freshness Requirement

'Real time' is not a service level. Define the maximum acceptable age of the data at the moment an agent reasons or acts.

Inventory availability may require seconds. A grantmaking foundation's annual eligibility reference data may tolerate a daily refresh. A fraud or payment decision may require both a current operational check and an immutable record of the data used.

Measure freshness from the source event to agent availability. For replicated data, include change-capture lag, transformation time, target apply time, and index refresh. For federated data, include source response time, cache behavior, and catalog propagation.

Test 3: Identify Where Meaning Is Created

Data access is not the same as usable business context. Agents need consistent definitions for customer, household, active policy, available inventory, restricted fund, or overdue case.

If definitions and quality rules already live in a governed semantic layer, in-place access may preserve them. If sources use incompatible codes, identifiers, or schemas, the organization may need mapping and transformation before the agent can use them safely.

This is where selective data movement remains valuable. A curated target can resolve identities, standardize units, preserve history, and apply quality rules once rather than forcing every agent to interpret raw operational tables.

Test 4: Protect Operational Performance

Federation makes the source part of the agent's runtime dependency. That can be appropriate when the source supports the expected query pattern and volume. It can be risky when complex joins or unpredictable agent queries compete with transaction processing.

Test concurrency, tail latency, query cancellation, rate limits, and failure behavior with representative traffic. A design that performs well for ten demonstration queries may not survive thousands of concurrent agent requests.

Replication can isolate analytical and agent workloads from the operational system. The tradeoff is lag and the burden of reconciling another copy. Choose the failure mode the business can manage.

Test 5: Trace the Full Governance Boundary

Keeping a source table in place does not mean the data never travels or persists elsewhere. Query results may enter model prompts, agent memory, logs, evaluation datasets, caches, analytics stores, or downstream applications.

Map the full path. Record which identity is evaluated, which row and column policies apply, which region processes the request, what is logged, how long derived data is retained, and how deletion or restriction propagates.

Google's BigQuery sharing documentation is instructive. Its sharing model provides linked datasets without replication and supports governance controls, but it also documents restrictions and interoperability limits. Zero copy can reduce duplication; it does not remove access design, cataloging, masking, audit, or lifecycle management.

Test 6: Design for Outages and Reproducibility

An agent that depends on live federation also depends on the source, network, catalog, and identity services. Decide what the agent should do when any dependency fails: stop, degrade to a read-only mode, use an approved cache, or route the case to a person.

Also decide whether the organization must reproduce the information that informed a past action. A live query shows what the source contains now, not necessarily what it contained when the agent acted. Regulated, financial, clinical-adjacent, or high-value decisions may require versioned snapshots, event logs, or an evidence record even when the primary access pattern is federated.

Choose by Workload: The Classification Matrix

WorkloadLikely Starting PatternReason to Reconsider
Current account or inventory lookupIn-place query or operational APISource cannot meet latency, volume, or availability requirements
Cross-platform customer or beneficiary viewHybrid architectureIdentity resolution and history require curated data while some attributes must stay current
Large historical analysis & trend forecastingReplicated or curated analytical copyData already shares storage and a governed semantic model
Agent action in a system of recordAuthenticated API with live validationOffline workflow is explicitly permitted and reconciled later
Model evaluation and audit replayVersioned or materialized datasetThe source provides durable time travel with the required retention
Cross-organization data collaborationGoverned sharing or clean-room patternJurisdiction, policy, or performance requires a controlled copy

The matrix is a starting point. Classify every data domain separately. An agent may use a replicated customer profile, a live inventory API, a federated policy table, and a versioned evidence log in the same workflow.

Building a Resilient Hybrid Control Plane

A hybrid architecture needs common controls even when the movement pattern varies across domains:

  • Unified metadata catalog: Record the source of truth, access method, owner, freshness target, permitted uses, sensitive fields, transformation rules, and fallback behavior for each data product.
  • Zero-trust authorization: Apply consistent identity and authorization at the point of access across all agents and tools.
  • Continuous observability: Monitor freshness, schema changes, query failures, replication lag, and policy denials in real time.
  • End-to-end lineage: Preserve lineage from the agent request to the underlying sources and tools it invoked.

This architecture also needs a clear separation between data preparation and agent reasoning. Deterministic controls should handle permissions, routing, masking, and transaction rules. The model should not decide whether it is allowed to see a field or execute a payment.

For organizations standardizing on Microsoft's data stack, YuniQ's Microsoft Fabric consulting and implementation services describe the same practical choice: workloads may be migrated, mirrored, accessed through OneLake shortcuts, or retained in source platforms based on latency, performance, governance, cost, and operational constraints.

Proof-of-Concept Scorecard: Eight Acceptance Criteria

Before approving an architecture pattern, run one realistic agent workflow through an objective, measurable proof of concept:

  1. Correctness: Compare agent-visible values with the source of truth across normal and edge cases.
  2. Freshness: Measure the 50th, 95th, and 99th percentile age of data when the agent receives it.
  3. Performance: Test latency and source impact at expected and peak concurrency.
  4. Schema change: Add, rename, widen, and remove representative fields; confirm detection, alerting, quarantine, and recovery.
  5. Access control: Test user-level permissions, service identities, denied fields, revoked access, and regional boundaries.
  6. Resilience: Interrupt the source, network, catalog, and target independently; verify the approved fallback.
  7. Reproducibility: Reconstruct the data, policy version, tool call, and outcome for a past agent action.
  8. Cost: Model query, compute, storage, egress, acceleration, observability, and operating labor at expected volume.

Require acceptance thresholds before the test begins. A proof of concept that only demonstrates connectivity cannot answer the architecture question.

Where SPEED AI Fits: Schema Intelligence & Hybrid Sync

Zero copy will reduce some data movement, but it will not eliminate the need to map schemas, synchronize selected data, validate values, mask sensitive fields, monitor pipelines, and route data to more than one target.

SPEED AI enterprise data architecture is designed for those operational requirements. YuniQ describes conversational pipeline configuration across more than 160 systems, schema matching and validation, multi-target flows, log-based change data capture, data spot checks, PII masking, metadata cataloging, and proactive pipeline alerts.

The right evaluation is not whether SPEED AI can make every data flow zero copy. It is whether the platform can help the organization implement and operate the movement, mapping, validation, and governance patterns that the workload decision requires.

Accelerate Your Agent Data Architecture

Start with one agent use case. List the systems it reads and updates, classify each access pattern, define the six test results, and measure them with real enterprise data.

Request a 14-Day SPEED AI Proof of Concept

Frequently Asked Questions

Is Zero Copy the Same as Zero ETL?

No. Zero copy usually means querying or sharing data without creating another full copy. Zero ETL may mean the same thing in one product, but another product may use the label for a managed replication pipeline. Always ask what data is copied, cached, indexed, transformed, and stored.

Does Zero Copy Eliminate Data Pipelines?

It can remove some ingestion pipelines. However, organizations still need access controls, semantic definitions, quality checks, lineage, monitoring, and often write-back or evidence flows. Hybrid architectures commonly retain selective replication or CDC.

When Is Replication Better Than Federation?

Replication is often stronger when the workload needs performance isolation, complex transformation, historical analysis, point-in-time reproducibility, local joins, or resilience when the source is unavailable. Its costs include another copy, synchronization lag, schema-drift handling, and reconciliation.

What Should Leaders Ask in a Zero Copy Demo?

Ask the vendor to draw the physical data flow; identify every persistent store, cache, index, log, and derived dataset; demonstrate row and column policies; show freshness and source-load measurements; break a dependency; change a schema; and reproduce a past agent result.