AI News Feed
Market watch
Products & Applications

AWS Embeds DuckDB Engine in Aurora PostgreSQL to Query Iceberg and S3 Data Alongside Live Records

Amazon Web Services is embedding the DuckDB analytics engine inside Aurora PostgreSQL, letting applications query Apache Iceberg and Parquet data in Amazon S3 next to live transactional records without reverse ETL pipelines or data copies.

The capability was described in an announcement by Esra Kayabali, a principal solutions architect at AWS, and reported by SiliconANGLE. According to that report, AWS is positioning the feature as a way to simplify application development and cut the engineering effort needed to maintain data pipelines. The company cited real-time dashboards, transactions enriched with historical information and artificial intelligence agents that need both current and archived records as potential uses.

Previously, combining recent transactions in Aurora with historical records stored in S3 generally required reverse ETL pipelines. AWS said that approach produced duplicated data, added infrastructure costs and required continuing work to keep records synchronized. “This challenge only grows as you increasingly embed AI agents into your applications, where it is impractical to predict and pre-replicate every dataset an agent might need,” Kayabali wrote in the announcement.

DuckDB performs analytical scans inside Aurora, avoiding additional network hops for query processing. AWS said a single query can reach data lake records alongside live operational data, including uncommitted writes, and described the integration as an example of how it is incorporating the DuckDB engine broadly across its services.

The feature also supports external catalogs that comply with the Iceberg Representational State Transfer Catalog specification, federating with the data catalog in the AWS Glue serverless data integration service. Customers register an external catalog with Glue and create foreign tables referencing its data, after which applications can join Aurora records with Iceberg tables registered across multiple catalogs.

To limit the volume of data read, Aurora filters records and selects relevant columns during query execution and caches frequently accessed data. Developers can inspect metrics including rows scanned, bytes read from S3 and cache hits. In a financial example, Kayabali demonstrated a query combining seven days of customer transactions in Aurora with five years of historical transactions stored in a Parquet file in S3; Aurora inferred the historical table’s schema from file metadata, eliminating manual column definitions.

For workloads requiring single-digit-millisecond latency, customers can copy selected data lake records into native Aurora tables using standard SQL commands. Read queries can run on the cluster’s writer or a read replica, offloading analytical scans from operational workloads, while commands that materialize data run on the writer.

Customers enable the capability through the aurora_analytics extension and an AWS Identity and Access Management role granting access to S3 and Glue. It supports Aurora PostgreSQL versions 17 and 18, beginning with versions 17.11 and 18.6 respectively. AWS said the feature is available in all commercial AWS regions with no additional feature charge; customers pay for the incremental Aurora computing resources their queries consume and the S3 requests used to read files.

AWS acquired DuckLabs B.V., the developer of the open-source DuckDB database, last month.

Editor's Summary

AWS has embedded the DuckDB analytics engine in Aurora PostgreSQL, allowing a single query to combine live transactional data with Iceberg and Parquet records stored in S3 and catalogs registered through AWS Glue, without reverse ETL pipelines. The feature is enabled through the aurora_analytics extension and an IAM role, supports Aurora PostgreSQL 17.11 and 18.6 onward, and carries no additional feature charge beyond the compute and S3 requests consumed. It follows AWS’s acquisition of DuckDB developer DuckLabs B.V. last month.