Amazon S3 Tables Unlocks Next-Generation Data Engineering with Full Apache Iceberg V3 Specification Support
SEATTLE — In a major development for petabyte-scale data lakes and modern enterprise architectures, Amazon Web Services (AWS) has announced full native support for the Apache Iceberg V3 specification across Amazon S3 Tables. The rollout, announced by AWS specialist Daniel Abib, equips data engineering, analytics, and machine learning teams with advanced capabilities designed to eliminate the bottlenecks that have long plagued large-scale data lakehouses.
By integrating Iceberg V3—including its advanced deletion vectors, built-in row lineage tracking, and a suite of rich new data types such as variant, nanosecond-precision timestamps, and geospatial geometries—AWS is positioning Amazon S3 Tables as the foundational cornerstone for high-performance, cost-effective open-table analytics.
Main Facts: What the Apache Iceberg V3 Integration Brings to S3 Tables
The adoption of Apache Iceberg V3 within Amazon S3 Tables introduces several architectural upgrades that directly target the performance degradation, storage bloat, and operational complexity historically associated with managing massive analytics datasets.
At its core, Amazon S3 Tables is a purpose-built storage tier designed to keep Apache Iceberg tables performant and economically scalable as they swell into the terabytes and petabytes. With the introduction of V3 support, users can now create brand-new V3 tables or execute seamless, in-place upgrades of legacy V2 tables. Key feature enhancements include:
- Deletion Vectors: V3 replaces traditional positional delete files—which often fragmented storage and degraded query efficiency—with compact, binary deletion vectors. Modifying or deleting tens of thousands of rows now generates a single lightweight vector file rather than thousands of localized delete fragments.
- Built-In Row Lineage: Every data record automatically receives unique identifiers (
_row_idand_last_updated_sequence_number), empowering downstream ETL pipelines to process incremental changes instantly without executing full-table scans. - Advanced Native Data Types: The specification adds native support for
variant(semi-structured JSON data stored in an optimized columnar format), nanosecond-precision timestamps, unknown data types, and geometry/geography spatial coordinates, ending the need to clumsily encode complex objects into basic strings or integers. - Seamless Managed Maintenance: S3 Tables continues to automate critical maintenance overhead, including background compaction, data replication, intelligent tiering, and cleanup operations for outdated delete logs.
Chronology: The Evolution Toward Open Table Formats and V3
To understand the magnitude of this release, it is helpful to trace the evolution of data lakehouse architectures and the role Amazon S3 has played in shaping modern analytics.
The Rise of Object Storage and Data Lakes
For over a decade, organizations stored massive volumes of unstructured and semi-structured data in Amazon S3, utilizing open-source file formats like Apache Parquet. While cost-effective and infinitely scalable, these raw data lakes lacked the transactional guarantees (ACID compliance) of traditional enterprise relational databases.
The Apache Iceberg Breakthrough
To solve the lack of ACID compliance, open-source table formats emerged, with Apache Iceberg quickly establishing itself as the gold standard for managing petabyte-scale datasets. Iceberg brought robust features like schema evolution, hidden partitioning, and time-travel queries to data lakes, allowing multiple compute engines to read and write data safely against shared Parquet files. However, as tables expanded past billions of rows, limitations in the V2 specification—particularly concerning rapid mutations, handling semi-structured data, and tracking incremental changes—began to surface.
The Arrival of S3 Tables and V3
AWS launched Amazon S3 Tables to directly address the operational friction of managing Iceberg metadata and performance at scale. Building upon that foundation, AWS has now bridged the gap to the Apache Iceberg V3 specification. Available immediately across all AWS Regions that support S3 Tables, this release provides the industry’s broadest native cloud support for V3, closing the loop across ingestion, storage, cataloging, and compute layers.
Supporting Data: Addressing Enterprise Analytics Pain Points
The architectural limitations of Apache Iceberg V2 created compounding hidden costs for data-driven enterprises. AWS’s release of V3 support directly addresses these friction points through measurable performance and efficiency gains.
1. Eliminating the Compaction Penalty of Deletion
In V2 architectures, executing a compliance mandate—such as deleting 50,000 specific user records from a massive 2-billion-row table—generated thousands of positional delete files. Every subsequent query had to read and evaluate these fragmented files until a heavy background compaction job finally ran.
Under V3 and S3 Tables, that same compliance delete writes a single, compact binary deletion vector. Query engines bypass the overhead of parsing thousands of tiny files, drastically reducing I/O latency, storage overhead, and compute costs.

2. Streamlining Semi-Structured Data with the Variant Type
Previously, semi-structured formats like JSON payloads had to be ingested as raw strings. Every analytical query required expensive parsing operations (PARSE_JSON) at runtime.
-- Creating a V3 table utilizing the new variant data type
CREATE TABLE my_catalog.namespace.clickstream (
event_id bigint,
event_time timestamp,
user_id string,
payload variant
)
USING iceberg
TBLPROPERTIES ('format-version' = '3')
With the V3 variant type, the storage engine shreds semi-structured data into hidden columnar formats during ingestion and gathers granular statistics. When queries run against the data, the engine uses these statistics to execute file pruning, slashing I/O workloads.
-- Querying variant columns directly without read-time parsing overhead
SELECT
event_id,
user_id,
variant_get(payload, '$.action', 'string') AS action,
variant_get(payload, '$.amount', 'double') AS amount
FROM my_catalog.namespace.clickstream
WHERE variant_get(payload, '$.action', 'string') = 'purchase'
AND variant_get(payload, '$.amount', 'double') > 50.00
3. Efficient Incremental Pipelines via Row Lineage
Data pipelines traditionally relied on full table scans or complex watermarking strategies to capture recently modified records. Iceberg V3 assigns unique identifiers to every row (_row_id) and tracks sequence updates (_last_updated_sequence_number). Downstream data pipelines can now pull only altered records via a lightweight predicate filter:
SELECT *, _row_id, _last_updated_sequence_number
FROM my_catalog.namespace.clickstream
WHERE _last_updated_sequence_number > 42
Official Responses and Ecosystem Integration
AWS has positioned its analytics portfolio to provide end-to-end compatibility for Iceberg V3, ensuring that organizations can immediately leverage the new specification across their entire data stack without friction.
Cross-Service AWS Compatibility
The introduction of Iceberg V3 support is deeply integrated across the broader AWS analytics ecosystem:
- Amazon EMR: Data engineers can write and transform V3 tables using Apache Spark on Amazon EMR.
- AWS Glue: AWS Glue 6.0 offers full Apache Iceberg V3 support alongside catalog interoperability via the Iceberg REST Catalog (IRC) API.
- Amazon Redshift: Enterprise data warehousing teams can seamlessly query and analyze V3-formatted datasets stored in S3 Tables.
Migration and Backwards Compatibility
To facilitate enterprise transitions, AWS has engineered graceful migration paths. Upgrading an existing V2 table to V3 can be performed atomically with a single command, requiring no immediate data rewriting:
ALTER TABLE my_catalog.namespace.existing_table
SET TBLPROPERTIES ('format-version' = '3')
AWS notes that this upgrade is designed with backward compatibility in mind; legacy V2 readers can continue accessing tables during phased enterprise rollouts. However, administrators should verify that all connected compute engines support the V3 specification, as downgrading from V3 back to V2 is not supported by the underlying Apache Iceberg specification.
Implications for the Future of Data Engineering
The broad availability of Apache Iceberg V3 on Amazon S3 Tables marks a significant maturity milestone for the data lakehouse paradigm.
For data engineering teams, the removal of cumbersome string workarounds for geospatial and semi-structured data translates directly into cleaner codebases, fewer pipeline errors, and lower infrastructure expenditures. By baking performance-enhancing features like deletion vectors and automated row lineage directly into managed storage, AWS is shifting the burden of low-level data optimization away from the developer and placing it firmly into the background storage layer.
Furthermore, as multi-engine architectures become the standard enterprise norm, the adherence to open specifications like Iceberg V3—coupled with universal REST catalog interoperability—ensures that organizations remain free from vendor lock-in. Companies can flexibly route workloads across Amazon EMR, AWS Glue, Amazon Redshift, and third-party analytics engines while maintaining a single, highly optimized, truth-bearing data lake.
Getting Started
Amazon S3 Tables support for Apache Iceberg V3 is available now at no additional charge across all AWS Regions where S3 Tables are supported; standard S3 Tables storage pricing applies. Developers and data architects can begin provisioning V3 table buckets via the Amazon S3 console, consult the official S3 Tables technical documentation, or leverage the AWS MCP Server and plugin toolkits for automated API interaction and troubleshooting.
