Add support for the OVERLAY string function. (#30026)
Add the ROW::fields function to return the field names of a row value.
(#29425)
Add support for invoking string functions as methods on varchar and char
values, such as value.length() and varchar::chr(65). (#29978)
Allow INSERT, UPDATE, and MERGE to assign values to array, map, and
row columns containing nested character types when the value is assignable
to the column type. (#29953)
Allow comparing and ordering row and array values that contain null
elements, for example in ORDER BY, DISTINCT, min(), max(), and
range comparisons. (#29891)
⚠️ Breaking change: Remove support for Alluxio-backed exchange storage for
fault-tolerant execution. (#29596)
⚠️ Breaking change: Reverse the implicit coercion between char and varchar: a
char value coerces to varchar with its trailing spaces removed, and
comparisons between char and varchar follow varchar semantics with no
blank padding, so trailing spaces are significant. The previous behavior can be
restored by setting the deprecated.legacy-varchar-to-char-coercion
configuration property to true. (#29939)
Disallow OMIT QUOTES in JSON_QUERY when the returned type is json.
(#29587)
Reduce memory usage that accumulates over time on long-running clusters that
execute many queries with aggregations. (#29523, #29524)
Improve performance of queries with a starts_with() filter by enabling
predicate pushdown. (#29712)
Fix incorrect results for queries on number columns containing NaN values.
(#29497)
Fix incorrect results when casting decimal values with precision smaller
than 19 to double. (#29831)
Fix DISTINCT and GROUP BY treating two NaN values in number columns as
distinct. (#29920)
Fix incorrect results for DISTINCT, GROUP BY, and = comparisons on
number values whose magnitude exceeds the supported precision.
(#29962)
Fix incorrect ordering, grouping, and aggregation of row values with more
than 64 fields. (#30013)
Fix incorrect column lineage for recursive common table expressions with
column aliases. (#29304)
Fix query failure when dereferencing a field of a null row value inside a
lambda expression. (#29504)
Fix failure when executing a SQL user-defined function that declares a
variable with a structural default value such as an array. (#29533)
Fix failure when a table function returns large pass-through columns.
(#29561)
Fix incorrect query progress reported through the /v1/query endpoint and the
Web UI. (#29793)
Fix query progress reported as NaN under fault-tolerant execution.
(#29791)
Fix failure when serializing custom data types provided by connectors over the
spooling protocol. (#29711)
Fix round() returning NaN for double and real values when rounding
to a number of decimal places that causes underflow. (#29808)
Fix DESCRIBE OUTPUT returning the number type to clients that do not
support it. (#29671)
Add the table location to the $properties metadata table. (#29521)
Add support for limiting the number of rows in Parquet writer row groups with
the parquet.writer.block-row-count configuration property or the
parquet_writer_block_row_count session property. (#29673)
Disallow creating tables partitioned by varbinary columns. (#24155)
⚠️ Breaking change: Remove support for the Alluxio file system. This does not affect
support for the Alluxio file system cache. (#29596)
Improve compression and read performance for high-cardinality string columns
in Parquet files. This can be disabled by setting the
parquet.writer.delta-length-byte-array-encoding-enabled configuration
property to false. (#29246)
Improve performance of queries that access nested fields inside lambda
expressions. (#29532)
Improve performance of reads from Google Cloud Storage when using static
credentials. (#29776)
Fix data loss caused by cleanup of a failed write deleting active data files
when deletion vectors are enabled. (#29361)
Fix failure when accessing Azure Storage with a hierarchical namespace using
vended credentials scoped to a subdirectory. (#29293)
Fix sporadic failures when reading data from S3. (#29419)
Add support for limiting the number of rows in Parquet writer row groups with
the parquet.writer.block-row-count configuration property or the
parquet_writer_block_row_count session property. (#29673)
Add the hive.parquet.max-split-size configuration property to control the
maximum size of splits for Parquet files, while hive.max-split-size applies
only to other file formats. (#29698)
⚠️ Breaking change: Remove the hive.max-initial-splits and
hive.max-initial-split-size configuration properties. (#29698)
⚠️ Breaking change: Remove support for the Alluxio file system. This does not affect
support for the Alluxio file system cache. (#29596)
Improve compression and read performance for high-cardinality string columns
in Parquet files. This can be disabled by setting the
parquet.writer.delta-length-byte-array-encoding-enabled configuration
property to false. (#29246)
Improve performance of queries that access nested fields inside lambda
expressions. (#29532)
Improve performance of reads from Google Cloud Storage when using static
credentials. (#29776)
Improve performance when a configured Hive metastore is unavailable by backing
off before retrying the connection. (#29653)
Fix sporadic failures when reading data from S3. (#29419)
Allow configuring the data_location table property without enabling
object_store_layout_enabled. (#28887)
Add support for limiting the number of rows in Parquet writer row groups with
the parquet.writer.block-row-count configuration property or the
parquet_writer_block_row_count session property. (#29673)
Add support for caching Parquet file footers, configurable with the
iceberg.parquet-footer-cache.type configuration property. (#29684)
Add support for the write.object-storage.partitioned-paths table property.
(#27633)
Add support for prefixed path storage credentials in the Iceberg REST
catalog. (#29641)
Add support for reading and writing geometry and geography values in
Iceberg v3 tables. (#27893)
Add support for the location property when creating views in the REST and
JDBC catalogs. (#29724)
Add the iceberg.max-split-size configuration property and max_split_size
session property to control the maximum size of splits. This replaces the
experimental experimental_split_size session property. (#29894)
Add support for the target_max_file_size and parquet_writer_row_group_size
table properties, persisted as the Iceberg-native write.target-file-size-bytes
and write.parquet.row-group-size-bytes properties. (#28250)
⚠️ Breaking change: Change the lower_bounds and upper_bounds columns of the
$files metadata table from map(integer, varchar) to typed row values.
Queries using lower_bounds[1] or casting these columns to json must be
updated. (#23147)
⚠️ Breaking change: Remove the target_max_file_size and parquet_writer_row_group_size
session properties, including the deprecated parquet_writer_block_size alias.
Use the equivalent table properties instead. (#28250)
⚠️ Breaking change: Remove support for the Alluxio file system. This does not affect
support for the Alluxio file system cache. (#29596)
Improve performance of queries on the $partitions metadata table.
(#23147)
Improve performance of queries on the iceberg_tables system table that
filter the table name with an IN predicate. (#29964)
Improve compression and read performance for high-cardinality string columns
in Parquet files. This can be disabled by setting the
parquet.writer.delta-length-byte-array-encoding-enabled configuration
property to false. (#29246, #29616)
Improve performance of queries that access nested fields inside lambda
expressions. (#29532)
Reduce memory usage when reading tables with many equality delete files.
(#29643)
Improve performance of reads from Google Cloud Storage when using static
credentials. (#29776)
Increase the maximum supported size of variant metadata from 16MB to 128MB.
(#29423)
Fix loss of materialized view data when running CREATE OR REPLACE MATERIALIZED VIEW for a view with a fixed storage location. (#29481)
Fix loss of the custom storage location of a view when running CREATE OR REPLACE VIEW in the REST and JDBC catalogs. (#25425, #29755)
Fix query failure when reading tables with delete files that reference nested
fields. (#29763)
Fix failure when reading statistics for a column whose type was changed.
(#29802)
Fix excessive token requests when listing nested namespaces in a REST catalog
using OAuth 2.0 with iceberg.rest-catalog.session set to NONE.
(#29722)
Fix creation of multiple delete files per data file when updating tables
partitioned only with identity transforms. (#28276)
Fix failure when listing schemas in an Iceberg REST catalog and a nested
namespace is dropped concurrently. (#29500)
Fix failure when listing columns or comments for tables in the Glue catalog
when access to a table is denied. (#29995)
Fix failure when accessing Azure Storage with a hierarchical namespace using
vended credentials scoped to a subdirectory. (#29293)
Fix sporadic failures when reading data from S3. (#29419)
Add support for limiting the number of rows in Parquet writer row groups with
the parquet.writer.block-row-count configuration property or the
parquet_writer_block_row_count session property. (#29673)
Improve performance of reads from Google Cloud Storage when using static
credentials. (#29776)
Fix sporadic failures when reading data from S3. (#29419)
Fix incorrect results when filtering a char or varchar column with an
equality or IN predicate where stored values differ only in trailing spaces.
Inequality and range predicates on character columns are no longer pushed down.
(#29952, #29939)
Fix incorrect results when a CAST from char to varchar is pushed down to
Oracle, which previously retained the blank padding instead of trimming it.
(#29939)
⚠️ Breaking change: Fix incorrect results when filtering a char or varchar column
with an equality or IN predicate where stored values differ only in trailing
spaces. Inequality and range predicates on character columns are no longer
pushed down, so a DELETE or UPDATE that relies on pushing down such a
predicate fails. (#29952, #29939)
Add the @StaticMethod annotation to register a scalar function as a static
method of a named type, invocable via T::method(args). (#29385)
Add the @InstanceMethod annotation to register a scalar function as a method
of its receiver type, invocable via expr.method(args). (#29399)
Add the @Name annotation to declare names for the arguments of functions
implemented in Java. (#29530)
Add support for lambda expressions in connector expression pushdown.
(#29532)
Add support for an optional description on connector table functions.
(#29830)
Add support for connector page sources to report memory usage through a
MemoryContext provided by the engine. (#29843)
⚠️ Breaking change: Replace the DynamicFilter parameter of
ConnectorSplitManager.getSplits with the set of dynamic filter columns, and
pass the current predicate to ConnectorSplitSource.getNextBatch as a
DynamicFilterSnapshot. (#29206)
⚠️ Breaking change: Remove the predicate from Constraint and add the
ConnectorExpressionEvaluator interface for evaluating expressions during
partition and split pruning. (#29289)
⚠️ Breaking change: Remove deprecated methods from ConnectorPageSourceProvider and
ConnectorPageSinkProvider. (#29562)