Apache Iceberg C++ 0.4.0 Release
The Apache Iceberg community is pleased to announce the 0.4.0 release of Apache Iceberg C++. This release includes over 160 pull requests from 26 contributors, including 12 first-time contributors.
iceberg-cpp is a native C++ implementation of the Apache Iceberg table format, providing libraries for reading, writing, and managing Iceberg tables in C++ applications.
Release Highlights🔗
Iceberg v3 Support🔗
- v3 type definitions and v3 geometry and geography support
- Column default values end to end: representation, serialization, and validation, Parquet reads of missing fields, Avro reads of missing fields, and
UpdateSchemasupport - Row lineage: writing snapshot row lineage fields at the top level, reading row lineage metadata columns, and retrying stale row-lineage validation
- Deletion vectors:
deletion-vector-v1blob read and write, merging multiple DVs inMergingSnapshotUpdate, and Puffin and DV support moved into the core library
Table Update APIs🔗
- New snapshot-producing operations:
DeleteFiles,RowDelta,OverwriteFiles,RewriteFiles,ReplacePartitions(now exposed onTableandTransaction), and merge append - Retryable and cleanup-safe table updates,
FileCleanupStrategyhardened with retries and parallel deletes, and cleanup of delete files from expired manifests - Catalog operations including
RenameTableinInMemoryCatalogand purge support inDropTablefor the in-memory and SQL catalogs
Catalogs and FileIO🔗
- REST catalog gains SigV4 authentication, a session-aware catalog, OAuth2 token exchange sessions, and per-table FileIO bound from vended storage credentials
ResolvingFileIOselects a FileIO by location scheme and forwards vended credentials, backed by a registry-driven resolution, with S3-compatible schemes and theossscheme supported- Hive Metastore groundwork with vendored HMS IDL and generated bindings, an
iceberg_hivelibrary, andHmsClientconnection lifecycle and URI parsing - Closer alignment with the Java implementation for REST table update serialization and REST error handling
Scan Planning and Performance🔗
- Lazy scan planning streams and a fallible iterator utility
- Parallelism for reading manifests, writing manifests, and update and scan processing, plus an LRU cache
- Metadata tables with a base metadata table interface and streaming
SnapshotsTablescans - Arrow data access via
ArrowRowBuilderand reading list columns aslarge_list - Metrics with commit and scan reporting integration, Parquet NaN metrics collected during writes, and a unique-value optimization for
notEq/notIn
Logging🔗
- A pluggable logging stack built up across the release:
LogLevel, theLoggerinterface and default logger, astd::cerrbackend, logging macros, an optional spdlog backend behindICEBERG_SPDLOG, and aLoggersregistry - Transaction commit lifecycle logging as the first consumer
Build Changes🔗
- Meson build support has been removed; CMake is now the only supported build system
- Dependency upgrades to Arrow 25.0.0 and nanoarrow 0.9.0
Contributors🔗
$ git shortlog --perl-regexp --author='^((?!dependabot\[bot\]).*)$' -sn v0.3.0..v0.4.0
23 Junwang Zhao
14 Gang Wu
13 Manu Zhang
13 Zehua Zou
13 wzhuo
10 Abanoub Doss
10 Minh Vu
9 Xin Huang
8 kamcheungting-db
7 Jiajia Li
5 YangJie
3 Kevin Liu
3 Rahul Goel
3 Xinli Shang
3 lishuxu
3 liuxiaoyu
2 Anupam Yadav
2 Joey
2 Yuya Ebihara
1 Huangshi Tian
1 Innocent Djiofack
1 Rahul Shivu Mahadev
1 Sreesh Maheshwar
1 Timothy Wang
1 ZhaoXuan
1 kid
This release welcomes 12 first-time contributors to Apache Iceberg C++: @kamcheungting-db, @All-less, @goel-skd, @abnobdoss, @ebyhr, @huan233usc, @yadavay-amzn, @timothyw553, @LuciferYang, @Angelia-Wang, @rahulsmahadev, and @u70b3.
We thank all contributors for their efforts in making this release possible!
Roadmap for 0.5.0🔗
The community is tracking the next release in #959, which focuses on the remaining Iceberg v3 feature support.
Getting Involved🔗
We welcome questions and contributions from all interested. Issues can be filed on GitHub, and questions can be directed to GitHub or the Iceberg dev mailing list.