Skip to content

AWS embeds DuckDB in Aurora PostgreSQL to accelerate data lake queries

Amazon Web Services Inc. (AWS) has integrated the DuckDB analytical engine into its Aurora PostgreSQL database management system, enabling customers to query Apache Iceberg data lakes and Parquet files directly. The feature aims to accelerate queries of live and historical data without requiring ETL pipelines, reducing engineering complexity and infrastructure costs for developers.

Editor, Lazyfounder

Published 6 min read
AWS embeds DuckDB in Aurora PostgreSQL to accelerate data lake queries
Image: SiliconANGLE via source

Amazon Web Services Inc. (AWS) has integrated the DuckDB analytical engine into its Aurora PostgreSQL database management system, enabling customers to query Apache Iceberg data lakes and Parquet files directly. The feature aims to accelerate queries of live and historical data without requiring ETL pipelines, reducing engineering complexity and infrastructure costs for developers.

30 SEC SUMMARY

  • AWS has embedded DuckDB, an analytical engine, into Aurora PostgreSQL to enable direct querying of Apache Iceberg data lakes and Parquet files without ETL pipelines.
  • The integration reduces network hops by processing analytical scans within Aurora, improving query performance for operational and historical data.
  • AWS acquired DuckLabs B.V., the developer of DuckDB, last month, signaling a strategic move to enhance data lake capabilities.
  • The feature is available in all commercial AWS regions at no additional charge, though customers pay for Aurora computing and S3 requests.
  • Developers can use existing PostgreSQL tools to query data in Amazon S3 alongside live transactions, simplifying application development.

TABLE OF CONTENTS

  • AWS integrates DuckDB into Aurora PostgreSQL
  • How the feature works
  • Performance and cost considerations
  • Potential use cases
  • Background on DuckDB and AWS’s acquisition
  • What this means
  • Key takeaways
  • FAQ
  • Sources

KEY HIGHLIGHTS

  • AWS has embedded the DuckDB analytical engine into Aurora PostgreSQL, enabling direct queries of Apache Iceberg data lakes and Parquet files.
  • The integration processes analytical scans within Aurora, reducing network hops and improving query performance.
  • AWS acquired DuckLabs B.V., the developer of DuckDB, last month, marking a strategic expansion of its data lake capabilities.
  • The feature is available in all commercial AWS regions at no additional charge, though customers pay for Aurora computing and S3 requests.
  • Developers can query data in Amazon S3 alongside live transactions using existing PostgreSQL tools, simplifying application development.

AWS integrates DuckDB into Aurora PostgreSQL

Amazon Web Services Inc. (AWS) has embedded the DuckDB analytical engine into its Aurora PostgreSQL database management system, enabling customers to query Apache Iceberg data lakes and Parquet files directly. According to SiliconANGLE, this integration allows users to analyze live operational data alongside historical records stored in Amazon S3 without requiring data copying or ETL (extract, transform, load) pipelines.

The DuckDB engine processes analytical scans within Aurora itself, reducing the need for additional network hops that typically slow down query performance. This capability is designed to simplify application development and reduce the engineering effort required to maintain complex data pipelines.

How the feature works

The new feature leverages Aurora PostgreSQL’s existing tools and endpoints, allowing customers to query operational records in Aurora alongside data stored in Amazon S3. This includes support for uncommitted writes, meaning a single query can access both live transactions and historical data lake records.

Aurora PostgreSQL supports external Iceberg catalogs through federation with the AWS Glue data catalog, which aligns with the Iceberg REST Catalog specification. During query execution, Aurora filters records and selects relevant columns to minimize the amount of data read. It also caches frequently accessed data to optimize performance.

Developers can monitor key metrics such as rows scanned, bytes read from S3, and cache hits to assess query efficiency. Schema inference for historical tables is handled automatically from file metadata, eliminating the need for manual column definitions.

Performance and cost considerations

The feature is available in all commercial AWS regions at no additional charge for the functionality itself. However, customers incur costs for the incremental Aurora computing resources their queries consume, as well as for S3 requests used to read files.

For workloads requiring single-digit-millisecond latency, AWS recommends copying selected data lake records into native Aurora tables using standard SQL commands. This approach ensures faster access for time-sensitive applications while maintaining the flexibility of querying broader datasets when needed.

Potential use cases

SiliconANGLE reports that the integration could be valuable for applications requiring real-time dashboards, transactions enriched with historical context, or AI agents that need access to both current and archived records. Previously, combining operational data in Aurora with historical records in S3 required reverse ETL pipelines, which introduced data duplication and added infrastructure costs.

The feature is supported on Aurora PostgreSQL versions 17 and 18, specifically versions 17.11 and 18.6 and later. Customers can enable it through the aurora_analytics extension and an AWS Identity and Access Management (IAM) role granting access to S3 and AWS Glue.

Background on DuckDB and AWS’s acquisition

DuckDB is an open-source analytical engine optimized for fast queries on large datasets, particularly those stored in columnar formats like Parquet. Its lightweight architecture makes it well-suited for embedded use cases, such as the integration with Aurora PostgreSQL.

AWS acquired DuckLabs B.V., the developer of DuckDB, last month. While the terms of the deal were not disclosed, the acquisition aligns with AWS’s strategy of enhancing its database and data lake capabilities to support AI and analytics workloads.

What this means

Lazyfounder analysis — our interpretation, not reported fact.

For founders and operators, AWS’s embedding of DuckDB into Aurora PostgreSQL is a practical step toward simplifying data architecture. By eliminating the need for ETL pipelines and enabling direct queries of data lakes, AWS is reducing the engineering overhead required to maintain complex data workflows. This could be particularly useful for startups leveraging AI or real-time analytics, where access to both operational and historical data is critical.

The move also reflects a broader industry trend: cloud providers are increasingly integrating specialized tools—like DuckDB—into their managed services to reduce friction for developers. For teams already using Aurora PostgreSQL, this feature could lower costs and accelerate time-to-insight, without requiring a shift to new tools or workflows. That said, the real-world impact will depend on performance at scale, especially for workloads demanding single-digit-millisecond latency. Teams should test the feature with their specific use cases before committing to architectural changes.

Key takeaways

  • AWS has integrated DuckDB into Aurora PostgreSQL to enable direct querying of Apache Iceberg data lakes and Parquet files.
  • The feature eliminates the need for ETL pipelines, reducing engineering complexity and infrastructure costs.
  • DuckDB processes analytical scans within Aurora, minimizing network hops and improving query performance.
  • AWS acquired DuckLabs B.V., the developer of DuckDB, last month, reinforcing its focus on advanced data lake capabilities.
  • The capability is available in all commercial AWS regions at no additional feature charge, but customers pay for Aurora computing and S3 request costs.

FAQ

What is DuckDB, and why is AWS using it?

DuckDB is an open-source analytical engine designed for fast queries on large datasets, particularly those stored in columnar formats like Parquet. AWS has embedded it into Aurora PostgreSQL to enable direct querying of data lakes and historical records without ETL pipelines, improving performance and simplifying data workflows.

How does this integration improve query performance?

The DuckDB engine processes analytical scans within Aurora PostgreSQL itself, reducing the need for additional network hops that typically slow down query execution. Aurora also filters and caches data to optimize performance.

What are the cost implications of using this feature?

The feature itself is available at no additional charge in all commercial AWS regions. However, customers pay for the Aurora computing resources their queries consume, as well as for S3 requests used to read files from data lakes.

Can this feature replace ETL pipelines entirely?

For many use cases, yes. The integration eliminates the need for ETL pipelines when querying data lakes alongside operational data in Aurora PostgreSQL. However, workloads requiring extremely low latency may still benefit from copying selected data into native Aurora tables.

How do developers enable this feature?

Developers can enable the feature through the aurora_analytics extension in Aurora PostgreSQL. They must also configure an AWS Identity and Access Management (IAM) role to grant access to Amazon S3 and AWS Glue.

Related on Lazyfounder

Sources

  1. SiliconANGLE · 2026-09-30
    AWS embeds DuckDB in its PostgreSQL DBMS to speed queries of live and historical data

This story is an original summary drafted with AI by Lazyfounder from the reporting listed above and checked by automated validation. Facts are attributed to their original publishers; sections marked as analysis are Lazyfounder's. Where a source is in another language, facts were machine-translated and quotations are reported, not reproduced. Read the original coverage via the links, and see our AI policy and corrections policy.

About the author

Editor, Lazyfounder

Tarun Mottlia edits LazyFounders, covering Indian startups, funding rounds, AI and product launches. Every story on the site is AI-assisted and checked against its cited sources before publication.

More stories by Tarun Mottlia

Get the LazyFounder Brief

Startup, funding and AI news in a five-minute read. Join the early-access list.

Lazy Founder - Powered by Blogy.in

Contact us

Have a story tip, correction or partnership idea?

Write to us at tarun.kumar@blogy.in or talk to the founder directly. We read every message.

Contact us