Spice v2.4.0-rc.1 (Oct 8, 2026)
Spice v2.4.0-rc.1 is now available! ๐ฅ
Spice v2.4.0-rc.1 is the first release candidate for v2.4.0. It adds performance improvements, S3 event-driven ingestion, SQL results-cache warmup, and adaptive HTTP rate controls. The release also upgrades to DataFusion v55, Ballista v55, Arrow v59, Vortex v0.86, Iceberg v0.11, and Turso v0.81.
Highlights in v2.4.0-rc.1 include:
- Performance & Query Engine: DataFusion v55, Vortex v0.86.1, Iceberg v0.11.0, and Ballista v55 update query planning, storage, and distributed execution.
- S3 Event-Driven Ingestion โ ingest newly created objects through SQS notifications.
- Cache Warmup and Shared Fetches โ warm the SQL results cache after the first refresh and coalesce concurrent source fetches.
- Adaptive HTTP Rate Controls โ bound request waits and coordinate request limits across instances.
- ORC Support โ query ORC files through listing connectors.
- Hugging Face Datasets: Query and accelerate files from the Hugging Face Hub.
What's New in v2.4.0-rc.1โ
Performance & Query Engineโ
This release upgrades Apache DataFusion to the v55.2.0 dependency line and Apache Arrow to v59.3.0. It also upgrades Vortex to v0.86.1, Apache Iceberg to the v0.11.0 fork, and Apache Ballista to v55.
Apache DataFusion v55โ
The DataFusion v55 release adds the following improvements:
- Sort pushdown and TopK pruning: Parquet scans reevaluate each unread row group as the threshold for
ORDER BY ... LIMITtightens. They skip groups that cannot contribute to the result. TopK pruning also supports multiple sort columns. - Join planning: The optimizer converts eligible inner joins to semi joins and removes redundant sides of outer joins. It also orders filter predicates by estimated cost.
- Aggregation and expressions: Multi-column
GROUP BYuses column-oriented storage for all supported key types, such as fixed-size binary UUIDs. More string functions preserve dictionary encoding, andINlists use specialized paths for small integer types. - Parquet reads: Scans skip nested fields that the declared schema does not contain. They also skip page-index reads when a file has no page index.
- Spill handling: Sorts bound the number of streams in a merge. If memory is insufficient, they spill the largest stream again in smaller batches.
- SQL diagnostics and functions:
EXPLAINaccepts PostgreSQL-style options andFORMAT pgjson. New array functions cover element-wise addition, subtraction, scaling, sums, and averages.
Spice carries these changes through its query plans and preserves statistics across plan wrappers. Cayenne keeps Vortex scans below 10 MiB unsplit to avoid repeated footer reads. See #14612.
Vortex v0.79.0 to v0.86.1โ
Spice v2.3.2 used the Vortex v0.79.0 fork. This upgrade covers the full upstream range from v0.79.0 through v0.86.1, not only the v0.86 changes:
- Types and arithmetic: Vortex adds native
Maparrays, Arrow map conversion, and operations and compression for maps. It also adds union arrays and decimal addition, subtraction, multiplication, and division. - Row selection: Piecewise-sequence indices represent contiguous selections without an expanded index for every row. Specialized paths handle chunked arrays, fixed-size lists, variable-length lists, and binary values.
- Scans and expressions: Layout scans gain a physical plan and expression optimization. Filters can pass through scalar functions with multiple arguments. Filters on wide lists restrict child elements to the selected range, and constant masks can resolve from metadata.
- Compression: Binary arrays support FSST compression with variable-length offsets. OnPair becomes a stable encoding for reads and gains a storage-backed dictionary. Vortex can convert run-end arrays of lists and decimals.
- File metadata: Files can store custom metadata. Readers cache decoded type descriptions, and file-format editions define the supported encodings and types.
- Memory and execution: Builders append nested values in batches, preserve list views, and propagate buffer allocators. Expression rewrites retain unchanged nodes, and row functions support batch execution.
- Correctness and validation: NULL handling changes cover dictionary predicates,
BETWEENbounds, and empty arrays. Readers add validation for footer offsets, compression metadata, and array indices. Other corrections cover nested scalar hashes, decimal operations, and variable-length binary selections above 4 GiB.
See the complete Vortex v0.79.0 to v0.86.1 changelog for every upstream change. Changes to standalone Vortex bindings and GPU execution do not imply new Spice features.
Apache Iceberg v0.11.0โ
Spice updates its Iceberg reader and catalog integrations from v0.10.1 to the v0.11.0 fork for DataFusion 55 and Arrow 59. The upstream v0.11 changes extend reads and catalog compatibility:
- Iceberg v3 reads: The reader applies deletion vectors from Puffin files and carries row identifiers and sequence numbers through scans.
- Metadata-only scans: Metadata-only projections do not require data-column reads. Manifest reads reuse partition types, and positional-delete processing buffers runs instead of allocating a key per row.
- REST catalogs: Clients negotiate server-advertised endpoints and use session-scoped OAuth2 authentication.
The upgrade also aligns Iceberg storage with OpenDAL v0.58. See #14771.
Apache Ballista v55โ
Spice.ai Enterprise feature. See the Enterprise documentation.
The Ballista v55 upgrade adds virtual-core resource accounting for distributed tasks. A protocol-version handshake detects incompatible schedulers and executors. Cancellation identifies tasks by their task IDs, and task-state records use an append-only model.
Schedulers and executors must run the same version. Upgrade all cluster components together.
S3 Event-Driven Ingestionโ
S3 listing datasets can use refresh_mode: changes with S3 event notifications delivered through SQS:
datasets:
- from: s3://my-bucket/events/
name: events
params:
file_format: parquet
s3_region: us-east-1
s3_auth: iam_role
s3_changes_queue_url: ${ secrets:events_queue_url }
acceleration:
enabled: true
engine: cayenne
mode: file
refresh_mode: changes
New object notifications append the object's rows. A periodic listing backfill covers missed or expired notifications. Each dataset needs its own queue and permissions to read and delete SQS messages, alongside its S3 read/list permissions.
This is object ingestion: removal notifications are ignored by default. Set s3_on_object_removed: rebuild to rebuild the entire prefix when an object is removed. An overwrite of an already applied object key is not ingested again; use new object keys for incoming data. See #14121.
SQL Results-Cache Warmup and Shared Fetchesโ
The SQL results cache can persist query plan shapes and replay them after a dataset's first full or append refresh:
runtime:
caching:
sql_results:
enabled: true
warmup: on_first_refresh
For example, run these queries against an accelerated orders dataset with warmup enabled:
SELECT id, status FROM orders WHERE id = 1;
SELECT id, status FROM orders WHERE id = 2;
Spice records one query shape because only the equality-filter value differs. After a restart and the dataset's first full or append refresh, warmup reruns that shape with distinct id values from the refreshed dataset, filling the cache before the dataset becomes ready. Keep the local .spice/data directory across restarts, or configure runtime.state.location, to retain recorded query shapes.
Warmup replays up to ten distinct recorded query shapes and tries up to 1,024 distinct filter-value combinations per shape, stopping when the cache is full. A dataset stays not ready until warmup completes. Later refreshes do not repeat warmup. The feature requires the default plan-based cache key; cache_key_type: sql is incompatible. See #14178.
Concurrent cache-miss fetches for the same request now share a source fetch. SQL results caching also includes changes for tables updated during query execution and for stale results served while revalidation runs. HTTP dataset caching supports RFC 5861 stale-if-error handling. See #14142, #14710, #14708, and #14134.
SQL, search, and embedding caches now use Spice's sharded cache backend. Existing engine: moka and engine: pingora values are accepted for configuration compatibility but no longer select a backend.
Adaptive HTTP Rate Controlsโ
HTTP rate controls now adapt admission to upstream failures while staying within configured request limits. rate_control_acquire_timeout bounds how long a request waits for capacity and defaults to the connector's client timeout. rate_control_failure_threshold and rate_control_window control the response to upstream failures.
In Spice.ai Enterprise, instances that share a runtime.state.location coordinate per-second and per-minute limits through shared state; OSS instances keep these limits in memory independently. Concurrency limits remain local to each instance. Components sharing an upstream origin must use matching rate-control settings. See HTTP rate-control documentation and #14143.
ORC Files and Object Metadata Queriesโ
Listing connectors support file_format: orc. Object-store listing also uses predicates on metadata columns, including _last_modified, to narrow eligible objects. Queries that select only partition or metadata columns can use those values without reading file contents. See #14075, #14265, #14303, and #14116.
Hugging Face Datasetsโ
The new Hugging Face data connector queries and accelerates datasets from the Hugging Face Hub. It supports Parquet, CSV, TSV, JSON, and ORC files. Public datasets need no credentials:
datasets:
- from: hf://datasets/stanfordnlp/imdb/plain_text/
name: imdb
acceleration:
enabled: true
The location format is hf://datasets/<owner>/<dataset>[@<revision>][/<path>]. A path can select a file, a folder, or a glob. A revision can name a branch, a tag, or a commit. Use @~parquet to select the Hub's automatic Parquet conversion.
Set hf_token for private or gated datasets. Set hf_endpoint for a Hub mirror or proxy. Each scan reads one commit. Refreshes follow the selected branch, but the dataset keeps its registered schema until reload. See #14877.
Automatic Primary-Key Handlingโ
Cayenne keeps one row per primary_key without requiring an on_conflict policy. When a dataset sets time_column, the row with the newest time wins; without it, the last arrival wins. This applies to full and append refreshes as well as writes. Existing explicit conflict policies remain accepted during the deprecation period.
datasets:
- from: s3://my-bucket/orders/
name: orders
time_column: updated_at
params:
file_format: parquet
acceleration:
enabled: true
engine: cayenne
mode: file
primary_key: id
For a read-write dataset whose writes should stay in its acceleration, set acceleration.write_mode: acceleration. Its source need not support writes. This mode cannot be combined with a dataset that refreshes by changes.
Cayenne secondary indexes also support dynamic join filters, and their write handling covers inserts, updates, deletes, and refreshes. See #14726, #14282, and #14593.
PostgreSQL and MySQL replication, and MongoDB change streams, rejected Cayenne datasets that omitted on_conflict. Their validation still required an explicit upsert policy. These sources now accept Cayenne datasets with primary_key alone. See #14880.
For file-mode Cayenne datasets, append refreshes failed with a configuration that combined primary_key, time_column, and retention_sql. The refresh selected a version-resolution path that did not support retention. The refresh now resolves each key's newest version before Cayenne applies retention. See #14878.
Cayenne maintained aggregates now share one compact index of per-key contributions across views. Rebuilds capture concurrent writes and apply them after the scan. The aggregate budget uses 10% of a bounded query pool without the former 512 MiB cap. An unbounded pool retains the 512 MiB budget. See #14762.
Read Datasets from Published Snapshotsโ
Spice.ai Enterprise feature. See the Enterprise documentation.
A dataset can now read published acceleration snapshots directly, without configuring the original source connector:
datasets:
- from: s3://my-bucket/spice/snapshots/orders/
name: orders
params:
file_format: snapshot
s3_region: us-east-1
Spice reads the snapshot metadata to select the engine, restores the published data, and checks for newer snapshots. The dataset is read-only. Snapshot-mode readers can also use S3 notifications delivered through SQS, with periodic checks retained for missed notifications. Each reader process needs its own queue.
This release includes changes to snapshot retries, slow-connection bootstrap, publication metadata, and coordination between snapshot archiving and Cayenne maintenance. Cayenne datasets with a datalake tier cannot create acceleration snapshots. See snapshot documentation, #14529, and #14335.
Connector and Protocol Updatesโ
- Connector status: ADBC, Databricks Spark Connect and SQL Warehouse, FlightSQL, Glue, HTTP/HTTPS, Iceberg, Localpod, and MongoDB are now Stable data connectors.
- MCP: support for specification version
2026-07-28, alongside the earlier protocol era. See #14043. - GitHub: nested GraphQL pagination and rate-limit pacing updates; the default concurrency limit is now four. See #14179 and #14431.
- GitHub nested pages: Scans could return incomplete reviews or comments because pagination accepted a short page as complete. The connector now rejects incomplete connections and repeated cursors. It retries a failed nested page without another fetch of the outer page. Datasets with the same token also share the REST quota. See #14862.
- Iceberg REST: Clients could read an empty dataset because the catalog synthesized metadata without snapshots. The catalog now returns the source metadata for Iceberg datasets that Spice reads unchanged. Other datasets return
400 BadRequestException. Clients need their own storage credentials. See #14588 and Breaking Changes.
Other Fixesโ
The release includes fixes in the following areas; the linked PRs provide details of the changes:
- Cayenne queries and writes: NULL-aware
NOT IN, maintained aggregates, dynamic filters, partition-filter forwarding, memory-mode DML and retention, and primary-key handling across CDC checkpoints. See #14429, #14761, #14370, #14047, and #14344. - Startup and reloads: retry datasets with unavailable sources, serve existing accelerations during source outages, and invalidate cached plans and results on catalog replacement or dataset unload. See #14623, #14624, #13914, and #14365.
- Federation: local evaluation of casts and functions whose source semantics differ, plus filter pushdown changes for DynamoDB, Cosmos DB, and MongoDB. See #14484, #14601, and #14419.
- HTTP and GraphQL: response-status handling during refresh, retry-budget handling, non-JSON gateway responses, and URL redaction in HTTP errors. See #13538, #14313, #14781, and #14490.
- Search: deletion of obsolete Elasticsearch chunks, non-finite embedding handling, and source-scan coordination during full-text refresh. See #13960, #13902, and #14663.
- Models and tools: tool-call-only assistant turns, required tool choices, streaming tool-use completion, model-load diagnostics, and propagation of the API-key principal into MCP tool calls. See #14232, #14460, #14548, and #14828.
- CDC shutdown and reconnects: source-position recording before accelerations close, and MySQL shared-stream reconnect handling. See #14702 and #14751.
- SQL weekdays:
date_part('dow')aligns withEXTRACT(dow), with Sunday represented as zero. See #14796. - Cayenne schema statistics: Decimal bounds could retain an old scale after schema evolution because maintenance published statistics from the previous schema. Cayenne now rejects statistics from an obsolete schema and keeps row counts conservative. See #14856.
- Vector search:
vector_searchplanning failed after the DataFusion 55 upgrade because a second optimization pass tried to reorder a join with a dynamic filter. The planner now preserves that join's input order. See #14857.
Default accelerator: Datasets and views that enable acceleration without engine now use Cayenne. Explicit engine settings keep their behavior. Storage still defaults to memory. Set mode: file for persistent acceleration. See Breaking Changes for migration guidance.
Dependency Updatesโ
| Dependency / Component | Version |
|---|---|
| DataFusion | v55.2.0 |
| Apache Arrow | v59.3.0 |
| Vortex | v0.86.1 |
| Apache Iceberg | v0.11.0 |
| Apache Ballista | v55.0.0 |
| Turso | v0.8.1 |
| ADBC | v0.24 |
| Rust toolchain | v1.98.1 |
Contributorsโ
Breaking Changesโ
Cayenne is the default accelerator on supported platforms. A dataset or view that omits acceleration.engine switches from Arrow to Cayenne. To retain Arrow, set engine: arrow explicitly before upgrading. Windows keeps Arrow as its default. Persistent datasets should continue to name their engine and use mode: file.
on_conflict is deprecated and scheduled for removal in v3.0. Cayenne automatically keeps one row per primary key, choosing the newest time_column value when configured, or the last arrival otherwise. Existing explicit policies remain supported during the deprecation period. Review those policies before removing them, especially drop or policies that reject conflicting rows.
on_conflict no longer routes writes to the acceleration. For read-write datasets whose writes should stay in the acceleration, use:
acceleration:
enabled: true
engine: cayenne
write_mode: acceleration
This setting cannot be used with refresh_mode: changes, including a connector's default change-stream mode. The default write_through and write_back modes require a writable source.
Cache engine selection is retired. engine: moka and engine: pingora remain accepted but are ignored. Remove the field and use caching_policy to select eviction behavior.
GitHub connector default concurrency is four. Review explicit concurrency settings if your deployment relied on the previous default.
HTTP rate-control waits are bounded by default. Requests waiting for rate-control capacity now time out after the connector's client timeout. Set rate_control_acquire_timeout to a suitable duration, or 0 to retain the previous unbounded wait behavior.
Cayenne acceleration snapshots are unavailable for datalake-tier datasets. Review snapshot settings on datasets using cayenne_datalake_location; this release disables snapshotting that configuration.
Iceberg REST no longer synthesizes metadata for unsupported datasets. GET /v1/namespaces/{namespace}/tables/{table} returns 400 BadRequestException for accelerated datasets, views, and other datasets that Spice does not read unchanged from Iceberg. If a client used this endpoint for schema discovery, use SQL DESCRIBE or information_schema.columns instead. Query these datasets through /v1/sql or Arrow Flight SQL. For eligible Iceberg datasets, clients read the source metadata and need their own storage access. See the Get a table API.
Cookbook Updatesโ
The Spice Cookbook provides recipes to help you get started with Spice.
Upgradingโ
To upgrade to v2.4.0-rc.1 once the release artifacts are available, use one of the following methods:
CLI:
spice upgrade v2.4.0-rc.1
Docker:
Pull the spiceai/spiceai:2.4.0-rc.1 image:
docker pull spiceai/spiceai:2.4.0-rc.1
For available tags, see DockerHub.
Helm:
helm repo update
helm upgrade spiceai spiceai/spiceai --version 2.4.0-rc.1
AWS Marketplace:
Spice is available in the AWS Marketplace. Marketplace availability follows its published versions.
What's Changedโ
Changelogโ
- fix(runtime): discard cached logical plans when a hot reload replaces a catalog (fixes #13910) by @claudespice in #13914
- fix(acceleration): let a schema repair correct a checkpoint without resetting the freshness clock (fixes #13817) by @claudespice in #13894
- fix(search): filter a chunked Elasticsearch delete on a field that can match the key (fixes #13714) by @claudespice in #13926
- fix(search): classify a partially non-finite embedding as unindexable on every backend (fixes #13872) by @claudespice in #13902
- fix: stabilize GitHub tests and bound GraphQL registration (fixes #13762) by @lukekim in #13939
- fix(postgres): decode versioned JSONB binary replication values by @phillipleblanc in #13962
- docs: require a reviewed Enhancement before any user-facing surface changes by @lukekim in #13970
- ci: upgrade spiceio setup action to v0.9.0 by @lukekim in #13971
- fix(postgres): preserve microseconds in timestamp writeback by @phillipleblanc in #13963
- feat(hash-index): verify the bloom filter's block index with Verus by @lukekim in #13777
- build(deps-dev): bump js-yaml by @dependabot in #13989
- docs: release notes for v2.3.0 by @bjchambers in #13999
- fix(ci): drop the dangling substrait-compliance submodule pointer by @bjchambers in #14002
- fix(cayenne): release the keyset bytes an abandoned PK checkout accounted (fixes #13668) by @grokspice in #13925
- fix(caching): keep a declared key from disabling eviction and stranding stale rows (fixes #13976) by @bjchambers in #13992
- Add Substrait compliance harness (IBM TPC-H Mode A + FlightSQL Mode B stub) by @lukekim in #13879
- ci: skip DynamoDB TPC-H benches in OSS testoperator dispatch by @phillipleblanc in #14016
- docs: update security support and roadmap after v2.3.0 by @phillipleblanc in #14024
- chore: post v2.3.0 release housekeeping by @bjchambers in #13969
- test(adbc): guard BigQuery corpus offline and in release gate by @phillipleblanc in #14017
- fix(duckdb): deny the regexp built-ins DuckDB cannot answer faithfully (fixes #13809) by @claudespice in #13871
- fix: Update tpch benchmark snapshots for federated/adbc[bigquery].yaml by @app/github-actions in #13984
- fix(search): drop the chunks a shortened row no longer produces from a chunked index (refs #13717) by @claudespice in #13960
- Reduce Cayenne allocations during primary-key validation and filtering by @lukekim in #14009
- test(forks): guard seven fork patches that had no repo-side test by @krinart in #13996
- fix(deps): bump arrow-rs to correctly-rounded DecimalโFloat cast (closes #13978) by @Jeadie in #14012
- perf(vortex): defer projection setup on filtered scans until the filter resolves by @bjchambers in #14035
- endgame: include spiceai/skills versioned release by @lukekim in #14031
- fix(caching): partition doomed entries at the survivor cutoff so eviction converges (closes #13994) by @Jeadie in #14021
- fix(deps): bump arrow-rs fork pin for Decimal->Float rounding fix by @Jeadie in #14049
- Fix subqueries with use_source acceleration by @phillipleblanc in #14022
- fix(cayenne): make DELETE, UPDATE and INSERT work on a
mode: memoryacceleration (fixes #12008) by @bjchambers in #14047 - fix: clarify OpenDAL S3 retry warnings by @lukekim in #14040
- test(s3): run the parquet-overwrite fixtures on RustFS by @bjchambers in #14067
- feat(cayenne): materialize multi-reference CTEs on the query path by @lukekim in #13918
- perf(vortex): answer a constant IN list by probing a set, and falsify it by interval by @bjchambers in #14061
- perf(vortex): skip a scan split whose zones cannot satisfy the filter by @peasee in #14064
- fix(arrow): report an exact row count from the indexed point-lookup scan by @krinart in #13972
- fix(cayenne): apply
sort_columnswithrefresh_mode: fullby @peasee in #14063 - fix(smb): pad an empty CREATE buffer so Samba lists the share root (fixes #13293) by @grokspice in #14050
- perf(cache): key the logical-plan cache on SQL text, not parameter values by @bjchambers in #14069
- fix: harden HuggingFace E2E chat against slow Metal generation by @lukekim in #14072
- fix(cayenne): move accelerator filesystem I/O off Tokio workers by @lukekim in #14073
- test(chbench): enable CTE materialization and IVM on mysql/postgres adaptive HTAP by @lukekim in #14070
- ci: run Substrait Mode A TPC-H on pull requests and the merge queue by @lukekim in #14071
- feat(mcp): support MCP specification 2026-07-28 (dual-era) by @lukekim in #14043
- fix(ci): call a linker that died of a signal an infrastructure failure, not a check failure (fixes #13614) by @grokspice in #14044
- fix: Provide temporary directory in docker images by @Jeadie in #14089
- fix: restore OSS installer, CLI and test workflow coverage by @phillipleblanc in #14025
- Delete v2.2.0.md by @Jeadie in #14095
- docs(release): add v2.3.1 release notes by @phillipleblanc in #14094
- fix(test): allow DELETE in the CORS allow-methods assertion by @claudespice in #14098
- fix(ci): stop install-protoc unzipping into a shared ~/.local by @lukekim in #14097
- feat(connectors): add ORC listing format via in-repo FileFormat by @lukekim in #14075
- fix(cayenne): run snapshot bootstrap check before opening the metastore by @Jeadie in #14093
- feat(cache): verify the results-cache namespace prefix with Verus by @lukekim in #14074
- feat(cloud-connect): add a GetDatasets command that answers the /v1/datasets document (refs #13369) by @grokspice in #14051
- Suppress Cayenne startup logs when no Cayenne dataset is configured by @Jeadie in #14042
- fix(cache): re-bind parameter values when revalidating a stale result (fixes #14099) by @bjchambers in #14100
- fix(cluster): support distributed HTTP scans by @phillipleblanc in #14108
- fix(bigquery): keep ILIKE evaluation local by @phillipleblanc in #14110
- docs: update security support for v2.3.1 by @phillipleblanc in #14117
- feat(cayenne): reuse ScanView until write, lag only for read-only CDC by @lukekim in #14055
- fix(deps): remediate open Dependabot alerts by @phillipleblanc in #14111
- fix(testoperator): validate results in every scale factor 1 TPC-H, TPC-DS and ClickBench benchmark by @lukekim in #14119
- fix(cayenne): reject ambiguous metastore paths by @phillipleblanc in #14130
- perf: serve results-cache hits where the request arrives and cut per-hit overhead by @lukekim in #14103
- test(runtime): record query previews in the management export test by @lukekim in #14155
- ci: upgrade spiceio setup action to v0.11.0 by @lukekim in #14152
- feat(caching): Make
caching_stale_if_errorRFC-5861 compliant (withstale-if-errorheader) by @Jeadie in #14134 - perf(runtime-table): defer cache-eviction key extraction to entries a delete actually names by @Jeadie in #14138
- fix(runtime): report the acceleration.ready_state deprecation once per component (fixes #13749) by @claudespice in #14006
- fix(runtime): write the inferred Arrow sort order under the prefixed key its validation accepts (fixes #14023) by @claudespice in #14032
- fix(connectors): Fix JSON/Orca files using metadata columns by @Jeadie in #14115
- feat(cayenne): build secondary indexes from
indexesin file and memory mode by @phillipleblanc in #14149 - fix(cayenne): round-trip decimal, binary, and time stats and drop them on scale change by @lukekim in #14139
- feat(cayenne): cluster warm and datalake tiers, and write full refreshes as key-range files by @lukekim in #14124
- fix(cayenne): compile the cold-tier pruning test and backtick a doc literal by @lukekim in #14175
- fix(cayenne): make the crates own targets lint and compile by @phillipleblanc in #14200
- test(forks): guard five more fork patches, and drop a row that is not fork state by @krinart in #14015
- perf(cache): promote encoded SQL results to raw after the second decode by @lukekim in #14199
- fix(turso): build a dictionary column directly so a dictionary over a list, map or boolean value reads back (fixes #13033) by @grokspice in #14181
- fix(ci): probe the macOS toolchain before reaching for brew in the release builds by @grokspice in #14203
- fix(ci): skip Metal kernel precompilation in the macOS release build by @grokspice in #14204
- fix(github): paginate nested GraphQL connections and pace to GitHub's rate limits by @lukekim in #14179
- fix(vortex): stop an IN list holding a NULL from panicking the scan by @krinart in #14163
- bench(cayenne): use std::hint::black_box in the clustering bench by @lukekim in #14129
- fix(runtime-table): stop rebuilding SessionContext on every cache fetch by @Jeadie in #14141
- test(chbench): cluster order_line, oorder and customer on the adaptive HTAP arms by @lukekim in #14192
- fix(runtime): count a first load as still loading in the Dataset load summary (fixes #13974) by @claudespice in #14020
- fix(duckdb): push regexp_count down again at a rendering that counts as the kernel does (fixes #13870) by @claudespice in #14153
- build: lint and test the sign-off under the same profile as the merge queue by @lukekim in #14180
- ci: require the Verus proofs in the merge queue as one check by @lukekim in #14189
- ci: run the longest macOS jobs on their own runner pool by @lukekim in #14229
- fix(cayenne): build the DELETE sink inside the execution-time write lock (fixes #13828) by @claudespice in #14218
- build(deps): bump the github-actions-dependencies group across 1 directory with 7 updates by @dependabot in #14231
- fix(cluster): decide what a Flight message carries by its IPC header, not its body length (refs #13737) by @claudespice in #14212
- ci: run Mode A TPC-H on merge queue and trunk/release push only by @lukekim in #14247
- chore(deps): bump spiceai/duckdb-rs to 76655d2f by @lukekim in #14246
- ci: stop exporting empty AWS and DuckLake endpoints to the schema test by @phillipleblanc in #14194
- fix(runtime): reload a localpod dataset when the dataset it reads through is reloaded (fixes #3288) by @claudespice in #14208
- build(deps): bump the aws-sdk group with 3 updates by @dependabot in #14255
- fix(ci): resolve Homebrew prefix when brew is the spice flock wrapper by @lukekim in #14210
- chore(deps): raise the datafusion-table-providers pin to include the NUMERIC result-column fix by @phillipleblanc in #14254
- ci: align the DuckLake bootstrap with the embedded DuckDB, wait for Databricks startup, and stop dispatching legs that cannot pass by @phillipleblanc in #14250
- fix(runtime): install the Spice function deny-list on the PostgreSQL catalog connector (refs #13664) by @claudespice in #14225
- perf(cayenne): share inline-cache view entries by Arc instead of cloning them per scan by @krinart in #14191
- build(deps): bump aws-actions/configure-aws-credentials by @dependabot in #14256
- ci: stop dispatching the indexed turso TPC-H SF1 tests by @phillipleblanc in #14252
- fix(catalog): keep the tables registered under an existing schema by @phillipleblanc in #14193
- feat(cache): Spice sharded cache as the sole LruCache engine by @lukekim in #14206
- ci: lint GitHub Actions definitions with actionlint, and fix the 73 findings it surfaced by @grokspice in #14223
- perf(cache): tighten the Raw SQL results-cache serve path by @lukekim in #14205
- ci: stop triggering the CUDA build on pull requests by @lukekim in #14267
- ci: run CodeQL on pull requests and the merge queue by @lukekim in #14269
- Release
2.3.2release notes by @krinart in #14271 - chore: make AGENTS.md the canonical agent instructions by @lukekim in #14237
- ci: run CodeQL Analyze on spiceai-dev-runners by @lukekim in #14281
- feat: TypeSafe Jev System One evaluation provider by @lukekim in #14215
- fix(cayenne): tag file statistics bounds as the column's Arrow type (fixes #14280) by @phillipleblanc in #14283
- ci: install spiceio when the runner has no gh (refs #14233) by @lukekim in #14288
- test(forks): 11 repo guards by @krinart in #14261
- docs: add Spice.ai in Action manuscript and companion labs by @lukekim in #13965
- test(duckdb): add an integration test for the index CTE materialization by @sgrebnov in #13885
- fix(cayenne): coalesce inline writes into one batch by @sgrebnov in #14279
- release: Update SECURITY.md and endgame template after 2.3.2 by @peasee in #14293
- fix(cayenne): count each Arrow allocation once in the inline-cache gauge by @krinart in #14272
- feat(s3): SQS event-driven changes for refresh_mode: changes by @lukekim in #14121
- feat(caching): single-flight coalesce concurrent cache-miss fetches by @Jeadie in #14142
- fix(caching): make caching_stale_if_error detect transient HTTP failures on real schemas by @krinart in #14161
- Use Cayenne secondary indexes for dynamic join filters by @phillipleblanc in #14282
- fix(ci): compare DynamoDB sets without relying on element order by @bjchambers in #14328
- docs(release): Remove QA analytics step from endgame by @peasee in #14329
- ci: run CodeQL on the merge queue and trunk, not on pull requests by @lukekim in #14290
- docs(cayenne): reposition the reference, add query serving, re-audit against trunk by @lukekim in #14338
- docs(endgame): update versioned docs release steps by @ewgenius in #14277
- test(forks): guard the ballista per-task file-scan restriction by @krinart in #14292
- fix(cli): surface the full error chain for spice chat connection failures by @krinart in #14289
- fix(ci): degrade the incomplete-sign-off handler when the runner has no gh (refs #14234) by @claudespice in #14304
- fix(runtime-table): serialize a direct write against acceleration snapshot creation (fixes #13548) by @claudespice in #14310
- fix(duckdb): screen regexp_like and regexp_replace as regexp_count is screened (fixes #14148) by @claudespice in #14321
- fix(http): honor retry budget without an extra origin request by @phillipleblanc in #14313
- Fix
SchemaCastScanExec's schema conversion infn partition_statisticsby @Jeadie in #14258 - fix(cluster): recognise a Flight keepalive by its empty envelope, not by what its header declares (fixes #13737) by @claudespice in #14327
- perf(http): defer zero-TTL acceleration lookup until origin failure by @phillipleblanc in #14302
- fix(cayenne): keep a key visible when it is re-inserted over a stale-insert tombstone by @sgrebnov in #14312
- perf(cayenne): read only key and filter columns in filtered key deletes by @sgrebnov in #14374
- fix(cayenne): let small protected-snapshot merges run during a long large-tier merge by @sgrebnov in #14296
- perf(cayenne): serve primary-key lookups on a freshly loaded table in ~1 ms by @lukekim in #14314
- fix(cayenne): skip min/max statistics for nested columns (fixes #14368) by @sgrebnov in #14392
- test(caching): cover SchemaCastScanExec statistics projection by name, retype, and SQL filter by @Jeadie in #14146
- fix(turso): keep a quantified comparison out of Turso SQL (fixes #14041) by @grokspice in #14393
- fix(udfs): declare the local_embed dev-dependency the embed tests need (fixes #13092) by @claudespice in #14377
- perf(cayenne): batch metastore manifest rewrites, upsert in place, and run every write on one writer connection by @lukekim in #14369
- fix(runtime-table): ignore zero-row batches in stale fallback by @phillipleblanc in #14331
- Prune object-store file listing by
_last_modifiedpredicates by @Jeadie in #14265 - Reading only partition or metadata columns needlessly scans all file contents by @Jeadie in #14116
- fix(duckdb): keep a concat over a binary operand out of the federated plan (fixes #13915) by @claudespice in #14333
- fix(ci): keep the sign-off attribution inside GitHub's 140-character status cap (fixes #14076) by @grokspice in #14390
- fix(ci): expose a present-but-unlinked cc tool on macOS runners instead of routing it through brew install (fixes #13479) by @grokspice in #14387
- ci: re-measure integration.yml's job bounds after the archive consolidation (fixes #13429) by @grokspice in #14388
- fix(ci): start DuckLake's local MinIO from an image that is still published by @grokspice in #14399
- test(runtime-table): make the metric-scraping refresh tests pass under cargo test by @claudespice in #14381
- fix(federation): keep a correlated subquery predicate above a join of two sources (refs #8220) by @claudespice in #14372
- fix(caching): snapshot staleness before the origin fetch for stale_if_error by @Jeadie in #14263
- perf(snapshots): skip unchanged snapshot metadata with a conditional GET by @sgrebnov in #14409
- fix(search): prune the rest of a key group from an Elasticsearch chunked index (refs #13717) by @claudespice in #14320
- fix(bench): derive MySQL's empty-field NULL handling from the column type (refs #13152) by @claudespice in #14345
- fix(graphql): debit a LIMIT by the rows a page returned, not the declared page size (fixes #14308) by @claudespice in #14353
- fix(ci): fit retention_oom's retry budget inside its workflow step, and guard the coupling (fixes #13512) by @grokspice in #14389
- docs: say plainly what Spice is, refresh the README for v2.3, and promote connector statuses by @lukekim in #14410
- fix(ci): pin, checksum and retry the oha download in the E2E graceful-shutdown jobs by @grokspice in #14418
- fix(snapshots): resolve snapshot entries relative to the metadata location (#14425) by @sgrebnov in #14426
- Lower the GitHub connector default concurrency limit to 4 by @lukekim in #14431
- build(deps): bump nvidia/cuda in the docker-dependencies group by @dependabot in #14439
- build(deps): bump the aws-sdk group with 3 updates by @dependabot in #14440
- build(deps): bump the github-actions-dependencies group across 1 directory with 5 updates by @dependabot in #14441
- test(cayenne): bound the refused-build guard by builds, not by the host's speed by @grokspice in #14424
- fix(cache): don't report an invalidation cancelled by runtime shutdown as a failure by @grokspice in #14417
- ci: run remote sign-off on the spiceai-macos pool by @lukekim in #14449
- ci: run CodeQL Analyze on spiceai-macos by @lukekim in #14442
- fix(runtime): keep the built-in date_part so both weekday spellings agree (fixes #13920) by @claudespice in #14154
- fix(cayenne): include the in-memory CDC tier when an overwrite or a retention pass covers the whole table by @lukekim in #14428
- fix: keep NOT IN null-aware through the join reorder and the Cayenne sort-merge rewrite by @lukekim in #14429
- fix(cayenne): discard a compaction whose snapshot an overwrite replaced mid-pass by @lukekim in #14432
- fix(cayenne): stop maintained views, Vortex IN lists and dynamic-filter sharing from returning wrong rows by @lukekim in #14427
- fix(cayenne): hide spilled rows on CDC upsert fallback by @bjchambers in #14416
- fix(ci): skip the integration, ADBC, chDB and E2E gate jobs on pull requests instead of passing them (fixes #13841) by @grokspice in #14454
- fix(cayenne): judge a filtered key delete's captured sources by the index captured with them (refs #13913) by @claudespice in #14455
- fix(ci): clear the macos-15 image's [email protected] symlink before installing MySQL by @lukekim in #14489
- ci: produce CodeQL SARIF in the Analyze job on spiceai-macos by @lukekim in #14500
- fix(cayenne): correctness fixes for the goal-driven adaptive controller, with a closed-loop simulation harness by @lukekim in #14443
- fix(cdc): keep the newest source commit timestamp on a coalesced change batch by @lukekim in #14463
- fix(cayenne): draw a sequence for a current-snapshot append so per-key OCC can order it (fixes #13685) by @claudespice in #14360
- fix(http): name a configured model's load failure instead of reporting it not found (fixes #13303) by @claudespice in #14395
- fix(duckdb): keep inferred source indexes off change-stream accelerations so upserts commit under concurrent reads (refs #13929) by @claudespice in #14396
- fix(duckdb): keep a text cast over a binary operand out of the federated plan (fixes #14355) by @claudespice in #14448
- fix(llms): honor tool_choice required and allowed_tools on mistral.rs-hosted models instead of panicking (fixes #14230) by @claudespice in #14460
- test(runtime): run the load-error counter test in its own process (fixes #13085) by @claudespice in #14462
- fix(runtime): stop a replaced dataset configuration's load from registering over the new one (fixes #1458) by @claudespice in #14367
- fix(install): stop asking for sudo on a first install into a fresh HOME (fixes #14445) by @claudespice in #14495
- perf(cayenne): serve primary-key point lookups in half the time by @lukekim in #14433
- fix(deps): bump DataFusion for upstream fixes to wrong results from filter pushdown, simplification and planning by @lukekim in #14430
- fix(runtime): size every internal DataFusion session from the CPU budget by @bjchambers in #14412
- fix(runtime): serve nested, zoned and half-float columns from the Iceberg catalog API (fixes #4815) by @claudespice in #14480
- fix(acceleration): keep cached results when a snapshot refresh finds no newer snapshot by @sgrebnov in #14497
- fix(runtime): sleep cron tests to the next boundary, not one that already fired (fixes #13759) by @claudespice in #14474
- fix(ci): derive every lint-rust guard make runs, whatever its separator or recipe layout (fixes #13783) by @claudespice in #14475
- fix(cayenne): stop the small-file compaction of a position-mode PK table from deadlocking on its own write lock (fixes #14420) by @claudespice in #14481
- fix(runtime-table): report refresh bytes for the rows each batch holds by @Jeadie in #14469
- fix(spark,databricks): keep Spice-only functions out of SQL sent to Spark Connect and Databricks SQL Warehouse (refs #13664) by @claudespice in #14498
- fix(vortex): size a cached footer by what it retains, not its serialized bytes (fixes #12917) by @claudespice in #14502
- fix(cayenne,telemetry): Register accelerated sink dataset immediately from existing acceleration by @peasee in #13955
- Update spicepod.yml by @Jeadie in #14533
- fix(cayenne): keep a rewrite's count inexact when it retains a late protected snapshot (fixes #14383) by @claudespice in #14385
- fix: Update tpch benchmark snapshots for accelerated/on_zero_results/file[parquet]-cayenne[file]-on_zero_results.yaml by @app/github-actions in #14408
- build(rust): upgrade toolchain to 1.98.1 by @lukekim in #14560
- fix(postgres): release a shared slot's hold on a table its publication cannot drop (fixes #13032) by @claudespice in #14527
- fix(cayenne): size the build side before sort-merging a non-Cayenne outer join by @krinart in #14520
- Update DF Upgrade template by @krinart in #14526
- fix(ci): wait for the refresh to invalidate the results cache, not a fixed 3s by @grokspice in #14598
- fix(deps): move the DataFusion pin past the four unparser fixes, and guard each of them (fixes #13022) by @grokspice in #14570
- fix(dynamodb, cosmosdb, mongodb): push filters down only where the source evaluates them as SQL does by @lukekim in #14419
- feat(snapshots): serve a dataset from published snapshots with
from: s3://โฆandfile_format: snapshotby @lukekim in #14529 - perf(cayenne): keep the per-shard PK index across checkpoint flushes and back off futile bakes; fix q17 and two wrong-results races (fixes #14235) by @lukekim in #14555
- test(chbench): lower the MySQL adaptive pods' CDC linger to 500 ms by @lukekim in #14615
- fix(graphql): show the parse-failure bytes in JSON decode previews by @Jeadie in #14530
- Defer source-first cache fallback planning by @phillipleblanc in #14556
- perf(cayenne): order whole-table rewrites per scan partition, then merge by @bjchambers in #14488
- fix(ci): tolerate the parquet-rename race's third DuckDB error in the E2E log scan by @grokspice in #14541
- fix(graphql): retry an inferred credential refusal instead of failing the refresh by @Jeadie in #14539
- fix(cayenne): serve a widened table's files from their persisted statistics (refs #13829) by @claudespice in #14316
- fix(acceleration): checkpoint DuckDB's write-ahead log before a snapshot copies its file (fixes #13912) by @claudespice in #14325
- fix(ci): let E2E cleanup run when the job failed before creating its working directory by @grokspice in #14553
- Replace unmaintained
backoffwith a workspace crate by @phillipleblanc in #14621 - test: read Expect stand-ins with the installed shell by @phillipleblanc in #14634
- fix(models): apply a tool_choice that forces a call to one round of the tool-use loop (fixes #14459) by @claudespice in #14635
- fix(runtime): don't end a tool-use stream on a stale tool_calls finish (fixes #13309) by @claudespice in #14548
- fix(smb): name listed locations under the share so directory datasets load (fixes #14060) by @claudespice in #14549
- fix(models): name a configured model's load failure in ai() and streaming nsql (fixes #14394) by @claudespice in #14632
- test(chbench): remove the mysql-cayenne[file]-adaptive-split HTAP arm by @sgrebnov in #14620
- fix(cayenne): record sharded CDC keys in the table-wide PK index by @lukekim in #14603
- fix(sqlite): keep TRY_CAST and every cast SQLite answers differently out of the federated plan (fixes #14398) by @claudespice in #14496
- ci: move remaining Ubuntu 22.04 references to 24.04 by @lukekim in #14677
- fix(snapshots): retry a failed snapshot attempt with backoff by @sgrebnov in #14676
- fix(snapshots): large snapshots no longer fail bootstrap on slow connection by @sgrebnov in #14571
- fix(refresh): start a refresh's source scan only when the sink reads it, so the full-text index keeps its rows (fixes #14619) by @claudespice in #14663
- test(cayenne): compare suite answers cell by cell against SQLite and chDB by @lukekim in #14473
- Evaluate any chat model via /v1/evaluate by @lukekim in #14568
- ci: quarantine the management API integration schedule until its dev OAuth client is reactivated (refs #12376) by @grokspice in #14684
- ci: stop scheduled Testoperator Ballista benchmarks by @phillipleblanc in #14682
- fix(cayenne): stop snapshotting datasets with a datalake tier, whose restored copies deleted each other's files by @lukekim in #14583
- fix(cayenne): count each Arrow allocation once for the mem-tier limit and checkpoint write sizing by @krinart in #14675
- fix(ci): classify a sign-off that reached its step budget as signalled, so an expiry publishes no verdict (fixes #13843) by @grokspice in #14685
- fix(runtime): retry a dataset whose source is unreachable at startup (fixes #14609) by @bjchambers in #14623
- fix(cayenne): archive only the snapshot directories a dataset snapshot references (fixes #14605) by @sgrebnov in #14627
- ci: route trunk-push macOS builds to the standard pool by @phillipleblanc in #14645
- ci: install strip for the retention OOM regression test by @phillipleblanc in #14701
- fix(smb): serve every share on a host from the one store registered for it (fixes #14550) by @grokspice in #14689
- fix(runtime): discard cached plans and results when a dataset is unloaded (refs #14251) by @claudespice in #14365
- feat(snapshots): reload snapshot-mode datasets from S3 event notifications (SQS) by @lukekim in #14335
- fix(cayenne): report runs per size tier when protected-snapshot compaction declines (fixes #13622) by @claudespice in #14476
- fix(cayenne): re-run a post-write compaction pass a concurrent append asked for (fixes #13906) by @claudespice in #14479
- fix(http): keep the request URL out of HTTP connector errors (fixes #13534) by @claudespice in #14490
- fix(runtime): load every localpod dataset that reads from one parent at startup (fixes #13087) by @claudespice in #14532
- fix(build): share one cargo metadata helper across the lint guards, so a broken cargo is never a violation (fixes #13121) by @claudespice in #14534
- fix(duckdb): push regexp_replace down again for a one-digit group reference, and refresh the Postgres ClickBench q35 plan (refs #13966) by @claudespice in #14552
- fix(duckdb): keep a cast into binary out of the federated plan (refs #14397) by @claudespice in #14633
- fix(runtime): keep finished async query jobs finished, and delete expired ones by @lukekim in #14585
- fix(federation): keep DataFusion's cast built-ins out of every federated plan (fixes #14444) by @claudespice in #14484
- test(cayenne): cover a swapped join filter through the sort-merge rewrite end to end (refs #14235) by @claudespice in #14551
- fix(runtime): decide every Spark/built-in function collision by name, and refuse an undecided one (fixes #14361) by @grokspice in #14400
- fix(ci): run every bin target's unit tests in the sign-off gate, and give spice connect its own --cloud-region refusal (fixes #13426) by @grokspice in #14406
- test(data_components): compile the federation unparser guards in every scoped test run (fixes #13625) by @grokspice in #14447
- fix(runtime-component): skip an inferred Cayenne index on a floating-point column instead of failing the load (fixes #14590) by @grokspice in #14599
- fix(cache): serve stale cached results during frequent table updates (stale_while_revalidate_ttl) by @sgrebnov in #14708
- Update Turso to 0.8.1 and retry write conflicts a metastore statement raises by @lukekim in #14680
- fix(snapshots): stop a snapshot dataset's load when a reload replaces it by @lukekim in #14673
- fix(cayenne): hand non-partition filters to every partition scan (fixes #12959) by @claudespice in #14370
- fix(postgres-accel): resolve secret references in the sidecar connection parameters (fixes #13296) by @claudespice in #14544
- fix(cayenne): drain the in-memory CDC tier before a full rewrite scans it (fixes #14450) by @claudespice in #14486
- feat(caching): warm SQL results cache on first refresh from persisted plan shapes by @lukekim in #14178
- fix(ci): retry only the container startup in the zero-retry MySQL CDC tests by @grokspice in #14727
- feat(key-index): immutable secondary index runs over compound Arrow keys by @bjchambers in #14592
- fix(federation): keep a fractional-to-integer cast out of the plans pushed to DuckDB, PostgreSQL, MySQL and BigQuery (fixes #14482) by @grokspice in #14601
- fix(testoperator): give accelerated bench configs a 300s ready_wait and name unready datasets on timeout (fixes #13973) by @claudespice in #14728
- fix(federation): keep arrow_typeof and its plan-introspection siblings local on every backend (fixes #14334) by @grokspice in #14695
- test(bench): refresh the tpch_q16 explain snapshots for the null-aware NOT IN plan (fixes #13977) by @claudespice in #14731
- ci: build trunk-push macOS legs on spiceai-macos-large again, keeping one-at-a-time coalescing by @grokspice in #14735
- ci(verus): derive the verified crates from cargo metadata; pin vstd once by @bjchambers in #14709
- docs: raise the test and evidence bar: differential first, exact assertions, performance always measured by @lukekim in #14723
- test(vortex): wait for a dropped segment cache to be freed before asserting it is gone (fixes #13295) by @claudespice in #14470
- ci(integration): report each integration part's own result in its required check by @grokspice in #14743
- fix(cayenne): back off futile bakes under a violated query goal too by @lukekim in #14736
- feat(object_store_occ): add transactional WAL and MVCC snapshots by @lukekim in #14732
- fix(snapshots): bootstrap readers from late snapshot publications by @phillipleblanc in #14468
- fix(cayenne): apply retention_sql to a
mode: memoryacceleration (fixes #14045) by @claudespice in #14344 - fix(duckdb, sqlite): keep the first copy of a key a write repeats under on_conflict: drop (fixes #14629) by @claudespice in #14748
- fix(cache): release a removed dataset's memory when the plan cache discards its plans (fixes #14251) by @claudespice in #14760
- build(deps): bump the aws-sdk group with 2 updates by @dependabot in #14764
- fix(llms): explain an Anthropic model's refusal of a sampling control instead of passing the bare 400 through (fixes #13564) by @grokspice in #14700
- build(deps-dev): bump dompurify by @dependabot in #14613
- fix(snapshots): allow cayenne_file_path and cayenne_metadata_dir on file_format: snapshot datasets (fixes #14696) by @sgrebnov in #14698
- test(snapshots): make snapshot integration tests robust under parallel CI runs by @sgrebnov in #14765
- feat(cayenne): keep the secondary index current across every write by @bjchambers in #14593
- fix: Update tpch benchmark snapshots for accelerated/file[parquet]-duckdb[memory].yaml by @app/github-actions in #14674
- Isolate Docker integration fixtures across concurrent processes by @phillipleblanc in #14639
- ci: run Substrait Mode A TPC-H at SF 1 on spiceai-macos in every merge-queue entry by @lukekim in #14522
- ci: stop repeating scheduled benchmarks: one source per accelerator, hosted sources weekly, a release commit once by @lukekim in #14681
- build(deps): bump nvidia/cuda by @dependabot in #14763
- fix(http): stop a non-2xx response body replacing an accelerated table's rows (fixes #13515) by @claudespice in #13538
- fix(chat-api): answer a tool-call-only assistant turn instead of panicking (fixes #13207) by @grokspice in #14232
- fix(sqlite, duckdb): keep a decimal AVG, and on SQLite a decimal SUM, out of the federated plan (fixes #14492) by @claudespice in #14670
- fix(spiceai, duckdb, sharepoint): name the Spicepod key a missing-parameter error asks for (refs #14446) by @claudespice in #14769
- Add table-bound ChangeSink ownership and backends by @phillipleblanc in #14704
- fix(ci): serialize rustup installs on shared macOS runners by @lukekim in #14717
- fix(cayenne): keep the per-shard PK index when the table-wide index is discarded by @lukekim in #14604
- Upgrade to DataFusion 55.1 and Arrow 59.3 by @krinart in #14612
- Update openapi.json by @app/github-actions in #14742
- build(deps): bump rustls from 0.23.40 to 0.23.45 by @dependabot in #14773
- Upgrade to iceberg-rust to 0.11.0 and DF 55.2 by @sgrebnov in #14771
- fix(cayenne): write an append into an empty unkeyed table as one load by @phillipleblanc in #14784
- fix(mysql_replication): don't re-send a live member's delivered commits after a reconnect by @lukekim in #14751
- ci(codeql): don't fail the SARIF upload when the merge queue already deleted its ref by @grokspice in #14793
- fix(acceleration): accept time_format iso8601 when the time_column is already a timestamp by @lukekim in #14777
- fix(cache): store SQL results that go stale mid-query by @lukekim in #14710
- fix(cayenne): serve filtered and global maintained aggregates by @lukekim in #14761
- perf(cayenne): build the checkpoint's tombstone union after releasing the capture locks by @lukekim in #14759
- ci: keep no artifacts or caches from PR and merge-queue checks, and run gating macOS jobs on spiceai-macos-large by @lukekim in #14804
- fix(cayenne): keep maintenance from deleting files a snapshot is archiving by @sgrebnov in #14789
- fix(snapshots): never overwrite shared snapshot metadata, and publish more than once to a file:// location by @lukekim in #14582
- perf(cdc): build deferred change rows in the reader while the apply runs by @lukekim in #14746
- fix: Update tpch benchmark snapshots for accelerated/on_zero_results/file[parquet]-cayenne[file]-on_zero_results.yaml by @app/github-actions in #14786
- test(testoperator): add a cold-start time-to-ready regression test by @phillipleblanc in #14785
- test(testoperator): make append tests more robust by @sgrebnov in #14817
- fix(runtime,data_components): wait for change-data-capture sources to record their positions before shutdown closes the accelerations (fixes #14523) by @grokspice in #14702
- fix(graphql): detect non-JSON responses and treat gateway errors as retryable by @lukekim in #14781
- fix(runtime): require keys for s3_auth: key, honor gs:// state location params, and leave newer rate-control state alone by @lukekim in #14584
- build(deps): bump hickory-resolver from 0.26.1 to 0.26.2 by @dependabot in #14774
- fix(listing): skip zero-byte objects (S3 folder markers) in format-selected listings by @sgrebnov in #14822
- Avoid unnecessary repartitioning in indexed dynamic-filter joins by @bjchambers in #14799
- feat(cayenne): keep one row per primary key automatically and deprecate on_conflict by @bjchambers in #14726
- fix(object-store): stop reading a response body at its first error by @phillipleblanc in #14831
- fix(ci): use byte-order collation in the benchmark Postgres container (refs #14815) by @sgrebnov in #14836
- fix(listing): keep a partition predicate as a residual filter on the
_locationfast path by @grokspice in #14790 - fix(ci): search only the unixodbc keg, not all of Homebrew's lib, in the macOS release builds by @grokspice in #14846
- fix(federation): keep date, timestamp and interval values local on SQLite reached through ADBC or ODBC (fixes #14753) by @claudespice in #14840
- chore: pin the fork at spiceai/datafusion#249 and guard the DataFusion fixes backported from 54 by @krinart in #14800
- Revert "fix(sqlite, duckdb): keep a decimal AVG, and on SQLite a decimal SUM, out of the federated plan (fixes #14492) (#14670)" by @krinart in #14825
- perf(runtime): skip EnsureRequirements while planning a point lookup by @lukekim in #14807
- fix(testoperator): report CH-benCH queries with no rows on either side as vacuous, not as matches by @lukekim in #14810
- feat(acceleration): make Cayenne the default acceleration engine by @phillipleblanc in #14837
- fix(cayenne): keep a join's LIMIT when the oversized-join rewrite makes it a sort-merge join by @lukekim in #14805
- test(bigquery): exclude subtrees an empty join build side never runs from the corpus job count (fixes #14848) by @sgrebnov in #14849
- fix(cayenne): fail contradictory write settings once, and order NULL times below every time by @bjchambers in #14847
- fix(runtime): serve an existing acceleration while its source is unavailable (fixes #14610) by @bjchambers in #14624
- Prune object-store file listing by metadata columns (#14264) by @Jeadie in #14303
- fix(vortex): fold partition values into the file-pruning predicate; test null-equal joins under mode:file by @lukekim in #14797
- feat(rate-control): adaptive, bounded and cluster-coordinated HTTP rate controls by @Jeadie in #14143
- fix(sql): align date_part('dow') with EXTRACT(dow) Sunday=0 by @lukekim in #14796
- fix(runtime): carry API-key principal into MCP tools/call by @lukekim in #14828
- perf(cayenne): hold maintained-aggregate retraction state in a compact shared index by @lukekim in #14762
- fix(deps): keep a hash join's order once it has a dynamic filter, so
vector_searchplans again by @Jeadie in #14857 - fix(cayenne): fence schema statistics and control maintenance tests by @bjchambers in #14856
- fix(github): retry nested GraphQL pages in place and fail incomplete nested connections by @sgrebnov in #14862
- fix(cayenne): load an append dataset that has retention_sql and orders versions by time by @phillipleblanc in #14878
- fix(cdc): Cayenne replication needs only a primary key, not on_conflict by @phillipleblanc in #14880
- feat(connector-huggingface): Hugging Face datasets data connector (
hf://datasets/...) by @lukekim in #14877 - fix(iceberg): Iceberg REST clients read the real table or get an error, never an empty one by @lukekim in #14588
Full Changelog: https://github.com/spiceai/spiceai/compare/v2.3.2...v2.4.0-rc.1

