Spice v1.10.0-rc.1 (Dec 2, 2025)

December 3, 2025 · 11 min read

Principal Software Engineer at Spice AI

Announcing the release of Spice v1.10.0-rc.1! ⚡

v1.10.0-rc1 is a release candidate for early testing of v1.10 features including an all new caching acceleration mode, tiny_lfu caching policy, a new DynamoDB Streams connector (Preview), improvements to the DynamoDB connector, faster distributed query execution, S3 connector improvements, and security hardening for v1.10.0-stable.

What's New in v1.10.0-rc1

Caching Acceleration Mode with SWR and TinyLFU

This release introduces a new caching acceleration mode that implements the stale-while-revalidate (SWR) pattern using Data Accelerators such as DuckDB or Cayenne, enabling queries to return file-persisted cached results immediately while asynchronously refreshing data in the background. Combined with the new TinyLFU cache eviction policy, Spice can now maintain higher cache hit rates while keeping memory usage predictable.

Key Features:

Stale-While-Revalidate (SWR): Returns cached data immediately while refreshing in the background
Data Accelerator Support: Cached accelerators can persist data to disk using DuckDB, SQLite, or Cayenne file modes.
TinyLFU Cache Policy: Probabilistic cache admission policy that maintains high hit rates with minimal overhead
Predictable Memory Usage: Configurable memory limits with automatic eviction of less frequently used entries

Example Spicepod.yml configuration:

runtime:
  caching:
    sql_results:
      enabled: true
      eviction_policy: tiny_lfu # default lru

datasets:
  - from: s3://my-bucket/data.parquet
    name: cached_data
    acceleration:
      enabled: true
      engine: duckdb
      mode: file # Persist cache to disk
      refresh_mode: caching
      refresh_check_interval: 10m

For more details, refer to the Data Acceleration Documentation and Caching Documentation.

DynamoDB Streams Data Connector in Preview

DynamoDB Connector now integrates with DynamoDB Streams which enables real-time streaming with support for both table bootstrapping and continuous change data capture (CDC). This connector automatically detects changes in DynamoDB tables and streams them into Spice for real-time query, search, and LLM-inference.

Key Features:

Real-Time CDC: Automatically captures inserts, updates, and deletes from DynamoDB tables
Table Bootstrapping: Initial full table load before streaming changes

Example Spicepod.yml configuration:

datasets:
  - from: dynamodb:my_table
    name: orders_stream
    acceleration:
      enabled: true
      refresh_mode: changes

For more details, refer to the DynamoDB Connector Documentation.

Cayenne Accelerator Enhancements

The Cayenne data accelerator now supports:

Sort Columns Configuration: Optimize inserts by pre-sorting data on specified columns for improved query performance

Example Spicepod.yml configuration:

datasets:
  - from: s3://my-bucket/data.parquet
    name: sorted_data
    acceleration:
      enabled: true
      engine: cayenne
      mode: file_create
      params:
        sort_columns: timestamp,region

For more details, refer to the Cayenne Documentation.

S3 Connector Improvements

S3 Location Predicate Pruning: The S3 data connector now supports location-based predicate pruning, dramatically reducing data scanned by pushing down predicates to S3 listing operations. This optimization is especially effective for partitioned datasets stored in S3.

AWS S3 Tables Write Support: Full read/write capability for AWS S3 Tables, enabling fast integration with AWS's table format for S3.

For more details, refer to the S3 Tables Data Connector Documentation and Glue Data Connection Documentation.

Faster Distributed Query Execution

Distributed query planning and execution have been significantly improved:

Fixed executor registration in cluster mode for more reliable distributed deployments
Improved hostname resolution for Flight server binding, enabling better executor discovery
Distributed accelerator registration: Data accelerators now properly register in distributed mode
Optimized query planning: DistributeFileScanOptimizer improvements for faster planning with large datasets

For more details, refer to the Distributed Query Documentation.

Search Improvements

Search capabilities have been improved with several performance and reliability enhancements:

Fixed FTS query blocking: Full-text search queries no longer block unnecessarily, improving query responsiveness
Optimized vector index operations: Eliminated unnecessary list_vectors calls for better performance
Improved limit pushdown: IndexerExec now properly handles limit pushdown for more efficient searches

For more details, refer to the Search Documentation.

Security Hardening

Multiple security improvements have been implemented:

SQL identifier quoting: Hardened SQL identifier quoting across all connectors to prevent injection attacks
Token redaction: Sensitive tokens are now fully redacted in debug output to prevent credential leakage
Path traversal prevention: Fixed tar extraction to prevent path traversal vulnerabilities
Input sanitization: Added validation for top_n_sample order_by parsing
Improved credential handling: Improved credential management in Glue connector

Developer Experience Improvements

Health probe metrics: Added health probe latency metrics for better observability
CLI improvements: Fixed .clear history command in the REPL to fully clear persisted history

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

No major cookbook updates. The Spice Cookbook still offers 82+ recipes to help you prototype quickly.

Upgrading

To try v1.10.0-rc1, use one of the following methods:

CLI:

spice upgrade --version 1.10.0-rc1

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.10.0-rc1 image:

docker pull spiceai/spiceai:1.10.0-rc1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 1.10.0-rc1

AWS Marketplace:

🎉 Spice is available in the AWS Marketplace.

What's Changed

Changelog

Test-operator: Add tpcds_q8 to the default row-count validation skip list by @sgrebnov in #8185
fix: Remove unwrap_used from test by @peasee in #8212
Run glue_iceberg_integration_test_catalog as part of main integration tests by @sgrebnov in #8222
Add TPCH sf100 testoperator spicepods with dispatch by @Jeadie in #8192
Build with CPU native flags by @lukekim in #8224
Make copilot check for empty copyright header by @krinart in #8245
fix: Apply assertion clippy in CI/Makefile only by @peasee in #8229
feat: Support running queries only in testoperator by @peasee in #8211
DuckDB query planning: aggregate pushdown by @mach-kernel in #8174
install.sh improvements by @lukekim in #8252
Fix .clear history by @lukekim in #8254
fix: Pushdown dynamic filters to partition scans by @peasee in #8240
Harden SQL identifier quoting in connectors by @phillipleblanc in #8276
Cayenne sort_columns on insert by @lukekim in #8091
Redact token debug output by @phillipleblanc in #8280
fix: Cayenne configuration options by @lukekim in #8281
Prevent path traversal in untar by @phillipleblanc in #8284
Fix cluster mode executor registration by @mach-kernel in #8292
Unignore s3_vectors_kafka_stream test by @Jeadie in #8289
Post-release house keeping by @krinart in #8293
Improve generate_changelog script by @krinart in #8273
Acceleration mode caching by @lukekim in #8237
Sanitization and security checks by @lukekim in #7854
Add health probe latency metric by @phillipleblanc in #8300
Add distributed registration for data accelerators by @phillipleblanc in #8299
Pass IndexedTableProvider down in 'changes_stream' and 'append_stream' by @Jeadie in #8295
Add dynamodb-streams crate by @krinart in #8283
Distributed query: resolve executor hostname when determining Flight server binding by @mach-kernel in #8304
Return computed embeddings from index for partitioned S3Vectors by @Jeadie in #8306
[DDB Streams] Skeleton for DynamoDB Streams by @krinart in #8296
DistributeFileScanOptimizer: Improve planning performance by @mach-kernel in #8305
feat: Add an ExactLeftAccumulator implementation by @peasee in #8302
deps: Upgrade Vortex to 0.56 by @peasee in #8311
DynamoDB table bootstrapping + streaming by @krinart in #8312
Avoid calling S3Vector list_vectors (or equivalent) when indexing into VectorIndexs by @Jeadie in #8282
Add on_conflict testing support to append benchmark by @sgrebnov in #8314
docker: Add valid home directory to fix duckdb extension loading issue by @phillipleblanc in #8318
Add GH Workflow to run Append benchmark test by @sgrebnov in #8321
Exclude MySQL SF100 from test-operator dispatch by @sgrebnov in #8320
feat: Update clippy lints by @peasee in #8317
Add S3 location predicate pruning to listing connector by @phillipleblanc
Review feedback for caching mode accelerator by @phillipleblanc in #8326
Also include Dockerfile home changes for release build by @phillipleblanc in #8327
Change communication channel from Discord to Slack by @Jeadie in #8330
Replace Discord link with Slack link in README by @Jeadie in #8331
fix(glue): Prevent OpenDAL from automatic loading of AWS credentials from environment by @sgrebnov in #8337
Block on index read for FTS queries by @Jeadie in #8339
Fix search query provider by @Jeadie in #8343
Support for writing into AWS S3 Tables by @sgrebnov in #8344
Acceleration file_create mode by @lukekim in #8347
Don't block on lock in FTS query path by @Jeadie in #8348
feat: Add an optimizer rule to replace join accumulator for Cayenne by @peasee in #8316
S3 Vectors limit updates by @lukekim in #8352
Sanitize top_n_sample order_by parsing by @phillipleblanc in #8356
Update version to v1.10.0-rc.1 by @ewgenius in #8362
Improve IndexerExec to properly handle limit pushdown by @sgrebnov in #8366
Fix Cayenne partition_by metadata flaky integration test by @phillipleblanc in #8367
Rework caching accelerator to use the stale-while-revalidate pattern. by @phillipleblanc in #8365
Add TinyLFU caching policy by @lukekim in #8370

Spice v1.9.1 (Nov 24, 2025)

November 24, 2025 · 7 min read

Viktor Yershov

Senior Software Engineer at Spice AI

Announcing the release of Spice v1.9.1!🔥

v1.9.1 introduces Amazon Bedrock Nova 2 Multimodal embeddings support with high-dimensional vectors up to 3,072 dimensions and purpose-optimized embeddings for semantic search and retrieval operations, DynamoDB timestamp filter pushdown for more efficient append-mode acceleration with configurable time formatting, HTTP Data Connector health probe configuration for improved endpoint validation reliability, and Spice .NET SDK v0.2 with expanded .NET version support and updated gRPC libraries. This release focuses on bug fixes, stability, and performance improvements.

Amazon Bedrock Nova 2 Multimodal embeddings

Spice now supports the Amazon Nova 2 Multimodal embeddings models via the Bedrock models provider, enabling high-quality text embeddings for semantic search and vector similarity operations. The Nova embeddings model offers configurable dimensions and advanced features like truncation modes and embedding purpose optimization.

Key Features:

High-Dimensional Embeddings: Support for up to 3,072 dimensions for rich semantic representations
Configurable Truncation: Control how input text is truncated when exceeding token limits (START, END, or NONE)
Purpose Optimization: Optimize embeddings for specific use cases (GENERIC_INDEX, GENERIC_RETRIEVAL, or CLASSIFICATION)
Multimodal Model: Leverages Amazon's Nova 2 multimodal architecture for consistent embeddings across different content types

Example spicepod.yml configuration:

embeddings:
  - from: bedrock:amazon.nova-2-multimodal-embeddings-v1:0
    name: nova_embeddings
    params:
      dimensions: '3072' # Required: Output dimensions
      truncation_mode: START # Optional: START, END, or NONE (default: NONE)
      embedding_purpose: GENERIC_RETRIEVAL # Optional. GENERIC_INDEX is default

For more details on the embedding parameters and configuration options, refer to the Amazon Nova Embeddings Documentation and the Spice Embeddings Documentation.

DynamoDB Timestamp Filter Pushdown

The DynamoDB Data Connector now supports timestamp filter pushdown, enabling more efficient append-mode acceleration refreshes by pushing timestamp filters directly to DynamoDB queries. Since DynamoDB stores timestamps as strings rather than native datetime types, this feature includes configurable timestamp formatting to ensure correct parsing and filtering.

Key Features:

Filters on timestamp columns are now pushed down to DynamoDB, reducing data transfer and improving query performance
Support for Go-style datetime formatting patterns to handle various timestamp string formats
Uses ISO 8601 format by default when no custom format is specified

Example spicepod.yml configuration:

datasets:
  - from: dynamodb:sales
    name: sales
    time_column: created_at
    time_format: timestamptz
    params:
      time_format: 2006-01-02T15:04:05.000Z07:00
    acceleration:
      enabled: true
      engine: duckdb
      refresh_mode: append

For more details, refer to the DynamoDB Data Connector Documentation.

HTTP Data Connector Health Probe Configuration

The HTTP Data Connector now supports configurable health probe paths for endpoint validation. Instead of using a random non-existent path, the system can now validate endpoints using a user-specified path, improving flexibility and reliability for health checks.

Example spicepod.yml configuration:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    params:
      file_format: json
      health_probe: /health-check

For more details, refer to the HTTP Data Connector Documentation.

Spice .NET SDK v0.2

The Spice .NET SDK has been upgraded with expanded .NET version support, custom User-Agent configuration, and updated gRPC libraries: spice-dotnet v0.2.0. The SDK is available on NuGet.

Key Features:

Expanded .NET Support: Now supports .NET Standard 2.0, .NET Core 8.0, 9.0, and 10.0.
Custom User-Agent: Configure custom User-Agent headers for client identification and telemetry.
Updated gRPC Libraries: Upgraded gRPC dependencies and netstandard for improved performance and reliability

Upgrade Example:

dotnet add package SpiceAI --version 0.2.0

For more details, refer to the .NET SDK Documentation.

Additional Improvements & Bug Fixes

Reliability: Fixed view loading to respect topological order, preventing dependency resolution errors.
Reliability: Migrated from deprecated trust_dns_resolver to hickory_resolver for improved DNS resolution reliability.
Security: Fixed arbitrary file access vulnerability during archive extraction ("Zip Slip") to prevent potential security exploits.
Distributed Query: Fixed object store initialization across scheduler/executor gap, improving reliability for distributed query execution.
Distributed Query: Optimized query routing by preventing runtime.* schema queries from being sent to the scheduler, improving performance for metadata queries.
Performance: Added Blake3 and xxHash support with xxh3_64 as the default caching hashing algorithm for improved cache and query performance.
Performance: Optimized default Zstd compression level to 6 for better balance between compression ratio and speed.
UX: Improved dataset loading output with clearer progress indicators and status messages.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

No major cookbook updates.

The Spice Cookbook includes 82 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.9.1, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.9.1 image:

docker pull spiceai/spiceai:1.9.1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Changelog

fix integration tests: order by the query to make snapshots deterministic by @phillipleblanc in #8198
Add health probe override by @lukekim in #8236
Use Moka optionally_get_with for SWR single-in-flight semantics by @lukekim in #8231
fix: Arbitrary file access during archive extraction ("Zip Slip") by @phillipleblanc in #8242
Migrate trust_dns_resolver to hickory_resolver by @phillipleblanc in #8243
fix: Deny assert macros in non-test code by @peasee in #8223
Distributed query: Object store initialization across scheduler/executor gap, misc bugfixes & improvements by @mach-kernel in #8009
Add Blake3, enable xxHash, set xxh3_64 as default, add bench by @lukekim in #8157
Make cache zstd default compression level 6 by @lukekim in #8234
Use seed for xxh3 by @lukekim in #8232
DynamoDB Timestamp Filter Pushdown by @krinart in #8235
Add ready_wait for mongo-arrow benchmarks by @krinart in #8246
Add support for amazon.nova-2-multimodal-embeddings-v1:0 by @Jeadie in #8225
Improve the output of dataset loading by @lukekim in #8256
Load views in topological order by @lukekim in #8255
Distributed query: Do not send runtime.* schema queries to scheduler by @mach-kernel in #8271
Remove input length check for Nova model. by @Jeadie in #8270

Spice v1.9.0 (Nov 19, 2025)

November 19, 2025 · 59 min read

Phillip LeBlanc

Co-Founder and CTO of Spice AI

Announcing the release of Spice v1.9.0-stable! 🌶

v1.9.0-stable introduces Spice Cayenne, a new high-performance data accelerator built on the Vortex columnar format that delivers better than DuckDB performance without single-file scaling limitations, and a preview of Multi-Node Distributed Query based on Apache Ballista. v1.9.0 also upgrades to DataFusion v50, DuckDB v1.4.2, and Delta-Kernel v0.16 for even higher query performance, expands search capabilities with full-text search on views and multi-column embeddings, and delivers many additional features and improvements.

What's New in v1.9.0

Cayenne Data Accelerator (Beta)

Introducing Cayenne: SQL as an Acceleration Format: A new high-performance Data Accelerator that simplifies multi-file data acceleration by using an embedded database (SQLite) for metadata while storing data in the Vortex columnar format, a Linux Foundation project. Cayenne delivers query and ingestion performance better than DuckDB's file-based acceleration without DuckDB's memory overhead and the scaling challenges of single DuckDB files.

Cayenne uses SQLite to manage acceleration metadata (schemas, snapshots, statistics, file tracking) through simple SQL transactions, while storing data in Vortex's compressed columnar format. This architecture provides:

Key Features:

SQLite + Vortex Architecture: All metadata is stored in SQLite tables with standard SQL transactions, while data lives in Vortex's compressed, chunked columnar format designed for zero-copy access and efficient scanning.
Simplified Operations: No complex file hierarchies, no JSON/Avro metadata files, no separate catalog servers—just SQL tables and Vortex data files. The entire metadata schema is intentionally simple for maximum reliability.
Fast Metadata Access: Single SQL query retrieves all metadata needed for query planning—no multiple round trips to storage, no S3 throttling, no reconstruction of metadata state from scattered files.
Efficient Small Changes: Dramatically reduces small file proliferation. Snapshots are just rows in SQLite tables, not new files on disk. Supports millions of snapshots without performance degradation.
High Concurrency: Changes consist of two steps: stage Vortex files (if any), then run a single SQL transaction. Much faster conflict resolution and support for many more concurrent updates than file-based formats.
Advanced Data Lifecycle: Full ACID transactions, delete support, and retention SQL execution on refresh commit.

Example Spicepod.yml configuration:

datasets:
  - from: s3:my_table
    name: accelerated_data_30d
    acceleration:
      enabled: true
      engine: cayenne
      mode: file
      refresh_mode: append
      retention_sql: DELETE FROM accelerated_data WHERE created_at < NOW() - INTERVAL '30 days'

Note, the Cayenne Data Accelerator is in Beta with limitations.

For more details, refer to the Cayenne Documentation, the Vortex project, and the DuckLake announcement that partly inspired this design.

Multi-Node Distributed Query (Preview)

Apache Ballista Integration: Spice now supports distributed query execution based on Apache Ballista, enabling distributed queries across multiple executor nodes for improved performance on large datasets. This feature is in preview in v1.9.0.

Architecture:

A distributed Spice cluster consists of:

Scheduler: Responsible for distributed query planning and work queue management for the executor fleet
Executors: One or more nodes responsible for running physical query plans

Getting Started:

Start a scheduler instance using an existing Spicepod. The scheduler is the only spiced instance that needs to be configured:

# Start scheduler (note the flight bind address override if you want it reachable outside localhost)
spiced --cluster-mode scheduler --flight 0.0.0.0:50051

Start one or more executors configured with the scheduler's flight URI:

# Start executor (automatically selects a free port if 50051 is taken)
spiced --cluster-mode executor --scheduler-url spiced://localhost:50051

Query Execution:

Queries run through the scheduler will now show a distributed_plan in EXPLAIN output, demonstrating how the query is distributed across executor nodes:

EXPLAIN SELECT count(id) FROM my_dataset;

Current Limitations:

Accelerated datasets are currently not supported. This feature is designed for querying partitioned data lake formats (Parquet, Delta Lake, Iceberg, etc.)
The feature is in preview and may have stability or performance limitations
Specific acceleration support is planned for future releases

For more details, refer to the Distributed Query Documentation.

DataFusion v50 Upgrade

Spice.ai is built on the Apache DataFusion query engine. The v50 release brings significant performance improvements and enhanced reliability:

Performance Improvements 🚀:

Dynamic Filter Pushdown: Enhanced dynamic filter pushdown for custom ExecutionPlans, ensuring filters propagate correctly through all physical operators for improved query performance.
Partition Pruning: Expanded partition pruning support ensures that unnecessary partitions are skipped when filters are not used, reducing data scanning overhead and improving query execution times.

Apache Spark Compatible Functions: Added support for Spark-compatible functions including array, bit_get/bit_count, bitmap_count, crc32/sha1, date_add/date_sub, if, last_day, like/ilike, luhn_check, mod/pmod, next_day, parse_url, rint, and width_bucket.

Bug Fixes & Reliability: Resolved issues with partition name validation and empty execution plans when vector index lists are empty. Fixed timestamp support for partition expressions, enabling better partitioning for time-series data.

See the Apache DataFusion 50.0.3 Release for more details.

DuckDB v1.4.2 Upgrade and Accelerator Improvements

DuckDB v1.4.2: DuckDB has been upgraded to v1.4.2, which includes several performance optimizations.

Composite ART Index Support: DuckDB in Spice now supports composite (multi-column) Adaptive Radix Tree (ART) indexes for accelerated table scans. When queries filter on multiple columns fully covered by a composite index, the optimizer automatically uses index scans instead of full table scans, delivering significant performance improvements for selective queries.

Example configuration:

datasets:
  - from: file://data.parquet
    name: sales
    acceleration:
      enabled: true
      engine: duckdb
      indexes:
        '(region, product_id)': enabled

Performance example with composite index on 7.5M rows:

SELECT * FROM sales WHERE region = 'US' AND product_id = 12345;

-- Without index: 0.282s
-- With composite index (region, product_id): 0.037s
-- Performance improvement: 7.6x faster with composite index

DuckDB Intermediate Materialization: Queries with indexes now use intermediate materialization (WITH ... AS MATERIALIZED) to leverage faster index scans. Currently supported for non-federated queries (query_federation: disabled) against a single table with indexes only. When predicates cover more columns than the index, the optimizer rewrites queries to first materialize index-filtered results, then apply remaining predicates. This optimization can deliver significant performance improvements for selective queries.

Example configuration:

datasets:
  - from: file://sales_data.parquet
    name: sales
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      params:
        query_federation: disabled # Required currently for intermediate materialization
      indexes:
        '(region, product_id)': enabled

Performance example:

-- Query with indexed columns (region, product_id) plus additional filter (amount)
SELECT * FROM sales
WHERE region = 'US' AND product_id = 12345 AND amount > 1000;

-- Optimized execution time: 0.031s (with intermediate materialization)
-- Standard execution time: 0.108s (without optimization)
-- Performance improvement: ~3.5x faster

The optimizer automatically rewrites the query to:

WITH _intermediate_materialize AS MATERIALIZED (
  SELECT * FROM sales WHERE region = 'US' AND product_id = 12345
)
SELECT * FROM _intermediate_materialize WHERE amount > 1000;

Parquet Buffering for Partitioned Writes: DuckDB partitioned writes in table mode now support Parquet buffering, reducing memory usage and improving write performance for large datasets.

Retention SQL on Refresh Commit: DuckDB accelerations now support running retention SQL on refresh commit, enabling automatic data cleanup and lifecycle management during refresh operations.

UTC Timezone for DuckDB: DuckDB now uses UTC as the default timezone, ensuring consistent behavior for time-based queries across different environments.

Example Spicepod.yml configuration:

datasets:
  - from: s3://my_bucket/large_table/
    name: partitioned_data
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      retention:
        sql: DELETE FROM partitioned_data WHERE event_time < NOW() - INTERVAL '7 days'

For more details, refer to the DuckDB Data Accelerator Documentation.

HTTP Data Connector

Querying endpoints as tables: The HTTP/HTTPS Data Connectors now supports querying HTTP endpoints directly as tables in SQL queries with dynamic filters. This feature transforms REST APIs into queryable data sources, making it easy to integrate external service data.
Query HTTP endpoint that returns structured data (JSON, CSV, etc.) as if it were a database table
Configurable retry logic, timeouts, and POST request support for more complex API interactions

Example Spicepod.yml configuration:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    params:
      file_format: json
      max_retries: 3
      client_timeout: 10s
      allowed_request_paths: /search/people
      request_query_filters: enabled
      request_body_filters: enabled

Example SQL query:

SELECT request_path, request_query, content
FROM tvmaze
WHERE request_path = '/search/people' and request_query = 'q=michael'
LIMIT 10;

If a request_body is supplied it will be posted to the endpoint:

Example SQL query:

SELECT request_path, request_query, content
FROM tvmaze
WHERE request_path = '/search/people' and request_query = 'q=michael' and request_body = '{"name": "michael"}'
LIMIT 10;

HTTP endpoints can be accelerated using refresh_sql:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    params:
      file_format: json
      allowed_request_paths: /search/people
      request_query_filters: enabled
      request_body_filters: enabled
    acceleration:
      enabled: true
      refresh_mode: full
      refresh_sql: |
        SELECT request_path, request_query, content 
        FROM tvmaze
        WHERE request_path = '/search/people'
          AND request_query IN ('q=michael', 'q=luke')

For more details, refer to the HTTP Data Connector Documentation.

DynamoDB Data Connector Improvements

Improved Query Performance: The DynamoDB Data Connector now includes improved filter handling for edge cases, parallel scan support for faster data ingestion, and better error handling for misconfigured queries. These improvements enable more reliable and performant access to DynamoDB data.

Example Spicepod.yml configuration:

datasets:
  - from: dynamodb:my_table
    name: ddb_data
    params:
      scan_segments: 10 # Default `auto` which calculates optimal segments based on number of rows

For more details, refer to the DynamoDB Data Connector Documentation.

S3 Data Connector Improvements

S3 Versioning Support: Spice now supports S3 Versioning for all connectors using object-store (S3, Delta Lake, etc.), ensuring range reads over versioned files are atomically correct. When S3 versioning is enabled, Spice automatically tracks version IDs during file discovery and uses them for all subsequent range reads, preventing inconsistencies from concurrent file modifications.

Current limitations:

Multi-file connections (e.g., partitioned datasets) do not yet support version tracking across all files
Version tracking is automatic when S3 versioning is enabled on the bucket

S3 Single-File Refresh Skipping: Spice now optimizes S3 single-file dataset refreshes by caching file metadata (ETag, Version ID, size, timestamp) and skipping unnecessary data fetches when the underlying file hasn't changed. This optimization dramatically reduces bandwidth usage and improves refresh performance for scenarios where data doesn't change frequently. The feature is enabled by default for accelerated S3 single-file datasets and includes metrics tracking for skipped refreshes.

Example configuration:

datasets:
  - from: s3://my-bucket/data.parquet
    name: s3_data
    acceleration:
      enabled: true
      engine: duckdb
      refresh_check_interval: 10s

When the file's metadata hasn't changed between refresh checks, Spice will skip the data fetch entirely, logging:

Skipping refresh for dataset 's3_data': file metadata unchanged

For more details, refer to the S3 Data Connector Documentation.

Search & Embeddings Enhancements

Full-Text Search on Views: Full-text search indexes are now supported on views, enabling advanced search scenarios over pre-aggregated or transformed data. This extends the power of Spice's search capabilities beyond base datasets.

Multi-Column Embeddings on Views: Views now support embedding columns, enabling vector search and semantic retrieval on view data. This is useful for search over aggregated or joined datasets.

Vector Engines on Views: Vector search engines are now available for views, enabling similarity search over complex queries and transformations.

Example Spicepod.yml configuration:

views:
  - name: aggregated_reviews
    sql: SELECT review_id, review_text FROM reviews WHERE rating > 4
    embeddings:
      - column: review_text
        model: openai:text-embedding-3-small

For more details, refer to the Search Documentation and Embeddings Documentation.

Dedicated Query Thread Pool (Now Enabled by Default)

Dedicated Query Thread Pool: Query execution and accelerated refreshes now run on their own dedicated thread pool, separate from the HTTP server. This prevents heavy query workloads from slowing down API responses, keeping health checks fast and avoiding unnecessary Kubernetes pod restarts under load.

This feature was opt-in in previous releases and is now enabled by default. To disable it and revert to the previous behavior, add the following spicepod.yaml configuration:

runtime:
  params:
    dedicated_thread_pool: none

For more details, refer to the Runtime Configuration Documentation.

Query Performance Optimizations

Stale-While-Revalidate Cache Control: Query results now support "stale-while-revalidate" cache control, allowing stale cached data to be served immediately while asynchronously refreshing the cache entry in the background. This improves response times for frequently-accessed queries while maintaining data freshness. Requires cache key type to be set to "sql (raw)" for proper operation.

Optimized Prepared Statements: Prepared statement handling has been optimized for better performance with parameterized queries, reducing planning overhead and improving execution time for repeated queries.

Large RecordBatch Chunking: Large Arrow RecordBatch objects are now automatically chunked to control memory usage during query execution, preventing memory exhaustion for queries returning large result sets.

Query Result Caching: Compressed Encoding, Stale-While-Revalidate Cache Control

Zstd Compression Encoding: Query result caching now supports optional Zstandard (zstd) compression encoding to reduce memory usage for cached query results. This is particularly beneficial for large result sets, reducing cache memory footprint while maintaining fast decompression times. Encoding can be configured via the encoding parameter with options none (default) or zstd.

Example configuration:

runtime:
  caching:
    sql_results:
      enabled: true
      max_size: 128MiB
      item_ttl: 1m
      encoding: zstd # Enable zstd compression

HTTP Cache-Control Support: The query result cache now supports the stale-while-revalidate Cache-Control directive, enabling faster response times by serving stale cached results immediately while asynchronously refreshing the cache in the background. This feature is particularly useful for applications that can tolerate slightly stale data in exchange for improved performance.

Example configuration:

runtime:
  caching:
    sql_results:
      enabled: true
      max_size: 128MiB
      item_ttl: 1m
      stale_while_revalidate_ttl: 1m # serve stale items for up to 1 minute after `item_ttl` expires

How it works:

When a cache entry is stale but within the stale-while-revalidate window, Spice will:

Immediately return the stale cached result to the client
Asynchronously re-execute the query in the background to refresh the cache
Future requests will use the refreshed data

Configuration:

Use the Cache-Control HTTP header with the stale-while-revalidate directive:

Cache-Control: max-age=300, stale-while-revalidate=60

This configuration caches results for 5 minutes (300 seconds), and allows serving stale results for an additional 60 seconds while refreshing in the background.

Requirements:

Must use plan or raw SQL cache keys (set cache_key_type to sql or plan in results_caching configuration)
Background revalidation re-executes queries through the normal query path
Timestamp tracking automatically determines cache entry age for staleness checks

Example configuration via HTTP header:

GET /v1/sql
Cache-Control: max-age=600, stale-while-revalidate=120
X-Cache-Key-Type: sql

This feature improves application responsiveness while ensuring data freshness through background updates.

For more details, refer to the Results Caching Documentation.

Security & Reliability Improvements

Enhanced HTTP Client Security: HTTP client usage across the runtime has been hardened with improved TLS validation, certificate pinning for critical endpoints, and better error handling for network failures.

ODBC Connector Improvements: Removed unwrap calls from the ODBC connector, improving error handling and reliability. Fixed secret handling and Kubernetes secret integration.

CLI Permissions Hardening: Tightened file permissions for the CLI and install script, ensuring secure defaults for configuration files and credentials.

Oracle Instant Client Pinning: Oracle Instant Client downloads are now pinned to specific SHAs, ensuring reproducible builds and preventing supply chain attacks.

AWS Authentication Improvements

Improved Credential Retry Logic: AWS SDK credential initialization has been significantly improved with more robust retry logic and better error handling. The system now automatically retries transient credential resolution failures using Fibonacci backoff, allowing Spice to tolerate extended AWS outages (up to ~48 hours) without manual intervention.

Key features:

Automatic retry with backoff: Implements Fibonacci backoff for transient credential failures (network issues, temporary AWS service disruptions)
Better error handling: Distinguishes between retryable errors (connector errors) and non-retryable errors (misconfiguration)
Unauthenticated access support: Properly supports unauthenticated access to public S3 buckets without requiring credentials
Improved error messages: Provides detailed logging with attempt numbers, retry intervals, and error context for better troubleshooting

The improvements ensure more reliable AWS service integration, particularly in environments with intermittent network connectivity or during AWS service degradations.

Observability & Tracing

DataFusion Log Emission: The Spice runtime now emits DataFusion internal logs, providing deeper visibility into query planning and execution for debugging and performance analysis.

AI Completions Tracing: Fixed tracing so that ai_completions operations are correctly parented under sql_query traces, improving observability for AI-powered queries.

Git Data Connector (Alpha)

Version-Controlled Data Access: The new Git Data Connector (Alpha) enables querying datasets stored in Git repositories. This connector is ideal for use cases involving configuration files, documentation, or any data tracked in version control.

Example Spicepod.yml configuration:

datasets:
  - from: git:https://github.com/myorg/myrepo
    name: git_metrics
    params:
      file_format: csv

For more details, refer to the Git Data Connector Documentation.

Spice Java SDK 0.4.0

The Spice Java SDK has been upgraded with support for configurable Arrow memory limit: spice-java v0.4.0

SpiceClient client = SpiceClient.builder()
    .withArrowMemoryLimitMB(1024) // 1GB limit
    .build();

For more details, refer to the Java SDK Documentation.

CLI Improvements

Install Specific Versions: The spice install command now supports installing specific versions of the Spice runtime and CLI. This enables easy version management, downgrading, or installation of specific releases for testing or compatibility requirements.

Usage:

# Install a specific version
spice install v1.8.3

# Install a specific version with AI flavor
spice install v1.8.3 ai

# Install latest version (existing behavior)
spice install
spice install ai

Note: Homebrew installations require manual version management via brew install spiceai/spiceai/spice@<version>.

Persistent Query History: The Spice CLI REPL (SQL, search, and chat interfaces) now persists command history to ~/.spice/query_history.txt, making your query history available across sessions. The history file is automatically created if it doesn't exist, with graceful fallback if the home directory cannot be determined.

New REPL Commands:

.clear - Clear the screen using ANSI escape codes for a clean workspace
.clear history - Clear and persist the query history, removing all stored commands

Tab Completion: Tab completion now includes suggestions based on your command history, making it faster to re-run or modify previous queries.

Example usage:

sql> SELECT * FROM my_table;
sql> .clear          # Clears the screen
sql> .clear history  # Clears command history
sql> # Use arrow keys or tab to access previous commands

For more details, refer to the CLI Documentation.

Additional Improvements & Bug Fixes

Reliability: Fixed refresh worker panics with recovery handling to prevent runtime crashes during acceleration refreshes.
Reliability: Improved error messages for missing or invalid spicepod.yaml files, providing actionable feedback for misconfiguration.
Reliability: Fixed DuckDB metadata pointer loading issues for snapshots.
Performance: Ensured ListingTable partitions are pruned correctly when filters are not used.
Reliability: Fixed vector dimension determination for partitioned indexes.
Search: Fixed casing issues in Reciprocal Rank Fusion (RRF) for hybrid search queries.
Search: Fixed search field handling as metadata for chunked search indexes.
Validation: Added timestamp support for partition expressions.
Validation: Fixed regexp_match function for DuckDB datasets.
Validation: Fixed partition name validation for improved reliability.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

New HTTP Data Connector Recipe: New recipe demonstrating how to query REST APIs and HTTP(s) endpoints. See HTTP Connector Recipe for details.

The Spice Cookbook includes 82 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.9.0, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.9.0 image:

docker pull spiceai/spiceai:1.9.0

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Dependencies

DataFusion: Upgraded to v50
Apache Arrow: Upgraded to v56
DuckDB: Upgraded to v1.4.2
Delta Kernel: Upgraded to v0.16.0

Changelog

Fix for search field as metadata for chunked search indexes by @Jeadie in #7429
Bump object_store from 0.12.3 to 0.12.4 by @app/dependabot in #7433
Properly respect disabling snapshots by @phillipleblanc in #7431
Revert "Properly respect disabling snapshots" by @sgrebnov in #7439
Revert "Disable snapshots by default" by @sgrebnov in #7438
Add preview warning for write access mode by @sgrebnov in #7440
fix: regexp_match for DuckDB datasets by @kczimm in #7443
Add feature is currently in preview warning for snapshots by @sgrebnov in #7442
[Logger] Also emit Datafusion logs by @mach-kernel in #7441
add missing snapshot by @kczimm in #7446
Fix tracing so that ai_completions are parented under sql_query by @lukekim in #7415
Enable snapshot acceleration by default by @phillipleblanc in #7451
Disable acceleration refresh metrics by @krinart in #7450
Add v1.8 release notes by @phillipleblanc in #7430
fix: partition name validation by @kczimm in #7452
Fix lint error due to ignore without reasons by @krinart in #7454
Add models and CUDA support to spiced install script by @lukekim in #7457
Post-release 1.8 updates by @phillipleblanc in #7455
Remove println in datafusion by @phillipleblanc in #7461
Update end_game.md to notify once release is done by @sgrebnov in #7460
Remove italics from snapshot logging by @phillipleblanc in #7463
Update openapi.json by @app/github-actions in #7466
Fix generate spicepod schema by @phillipleblanc in #7464
Fix generate acknowledements by @phillipleblanc in #7465
Update spicepod.schema.json by @app/github-actions in #7469
fix: Ensure ListingTable partitions are pruned when filters are not used by @peasee in #7471
Create runtime-secrets crate by @phillipleblanc in #7474
Create runtime-parameters crate by @phillipleblanc in #7475
Don't download the snapshot if the acceleration is present by @phillipleblanc in #7477
Bump hyper-util from 0.1.16 to 0.1.17 by @app/dependabot in #7434
Add missing remote CLI feature by @lukekim in #7478
Add 1.8.0 release analytics by @sgrebnov in #7481
CLI multi-line input support for spice sql by @lukekim in #7479
fix: duckdb partitioning cannot reload config by @kczimm in #7482
fix: Make search cache test use a slower uncached search by @peasee in #7473
Add support for S3 dataset params by @phillipleblanc in #7476
Add DuckDB TPC-H memory limit variations by @lukekim in #7484
Add better snapshot validation for incorrectly configured spicepods by @phillipleblanc in #7487
Move blocking/sync I/O to spawn blocking by @lukekim in #7462
Add DuckDB file accelerator 2G and 4G dispatches by @lukekim in #7491
Validate spicepod file exists before running tests by @lukekim in #7492
Make snapshot reading/writing more robust with Iceberg-like metadata.json by @phillipleblanc in #7486
Re-use build-testoperator and ensure it's cached by @lukekim in #7494
Fix duckdb test operator limit casing by @lukekim in #7498
fix: Update benchmark snapshots by @app/github-actions in #7499
Create runtime-request-context crate by @Jeadie in #7459
Add integration tests for Acceleration DB snapshotting by @phillipleblanc in #7489
Two minor fixes for AI udf tests by @krinart in #7503
Add model response timeout for ai udf tests by @krinart in #7504
Show error if FTS is misconfigured for datasets/views by @krinart in #7458
Add test for chunked search index with search field as metadata by @Jeadie in #7513
Add sccache for build test operator by @lukekim in #7515
Enhancement: Add spill_compression to runtime config by @krinart in #7505
Improve GitHub Data Connector by @lukekim in #7510
Add RequestTimeoutException to S3 client by @Jeadie in #7514
Add sha=<> to snapshot logging by @phillipleblanc in #7521
Add Type to GitHub Data Connector issues and fix double aliasing for project author by @lukekim in #7519
DuckDB acceleration: fix memory leak in duckdb_arrow_scan by @sgrebnov in #7524
Fix partition_by accelerations when a projection is applied on empty partition sets by @phillipleblanc in #7526
Nullable fields for index columns by @Jeadie in #7523
Fix missing winver dependency for Windows by @krinart in #7538
Update mongo config for benchmarks by @krinart in #7546
Add acceleration snapshots cookbook to template by @phillipleblanc in #7527
Bump github/codeql-action from 3 to 4 by @app/dependabot in #7535
Bump golang.org/x/sys from 0.36.0 to 0.37.0 by @app/dependabot in #7529
Make spice chat play nice with Unix pipes by @Jeadie in #7525
Configurable DuckDB duckdb_index_scan_percentage & duckdb_index_scan_max_count by @lukekim in #7551
[cherry-pick] Release notes for release 1.8.1 by @krinart in #7556
Fix 1.8.1 release notes by @krinart in #7558
FTS index with nonfilterable metadata; search field not metadata by default. by @Jeadie in #7548
Properly set auth headers in github_release.py by @krinart in #7560
project_schema when using EmptyExec by @kczimm in #7543
Add 1.8.1 release analytics by @kczimm in #7561
Bump golang.org/x/mod from 0.28.0 to 0.29.0 by @app/dependabot in #7530
Hive-style partitioning for DuckDB file mode by @kczimm in #7563
Vortex Data Accelerator (Dev grade) by @lukekim in #7566
Only load eval scorers when eval defined by @Jeadie in #7549
Bump octocrab from 0.45.0 to 0.47.0 by @app/dependabot in #7531
Bump regex from 1.11.3 to 1.12.1 by @app/dependabot in #7532
Fix custom file path for Vortex Data Accelerator by @phillipleblanc in #7570
Add List type support to Vortex Data Accelerator by @lukekim in #7569
Bump parking_lot from 0.12.4 to 0.12.5 by @app/dependabot in #7534
Bump tokio-postgres from 0.7.14 to 0.7.15 by @app/dependabot in #7533
Remove duplicate line from 1.8.1 release notes by @krinart in #7580
Upgrade Go from v1.24.2 to v1.25.3 by @lukekim in #7582
Fix race condition in S3 Vectors index and bucket creation by @kczimm in #7577
Add runtime-async crate with managed Tokio runtime by @phillipleblanc in #7575
Optimize GitHub Actions workflows by @lukekim in #7584
Use 'location' as primary key for document tables by @Jeadie in #7567
Extend query-related metrics by @krinart in #7571
Enabling acceleration refresh metrics by using runtime.metrics config by @krinart in #7583
S3Vector service metrics in client by @Jeadie in #7502
Fix score order for one test case by @Jeadie in #7595
ObjectMeta filter pushdown for ObjectStoreTextTable by @Jeadie in #7572
Return TableProvider from CandidateGeneration::search. by @Jeadie in #7559
EmptyHashJoinExecPhysicalOptimization, and use in VectorScanTableProvider by @Jeadie in #7587
Update official Docker builds to use release binaries by @phillipleblanc in #7597
New Generate Changelog workflow by @krinart in #7562
BytesProcessedExec to allow optimizer to do limit pushdown by @mach-kernel in #7539
GitHub Data Connector add Projects, improve rate-limiting and error handling by @lukekim in #7547
Add copilot-instructions to help improve Copilot reviews by @lukekim in #7606
Add support for DuckDB table-based partitioning by @sgrebnov in #7581
fix: Use nextest for integration models tests by @peasee in #7617
Fix license issue in table-providers by @phillipleblanc in #7620
Remove Build Docker Image from PR checks by @phillipleblanc in #7621
Combine PR Lint + Build checks by @phillipleblanc in #7623
Remove Cache Rust builds step by @phillipleblanc in #7625
Rename duckdb_partition_mode to partition_mode param by @sgrebnov in #7622
DuckDB table partitioning: delete partitions that no longer exist after full refresh by @sgrebnov in #7614
Build integration test binaries in single job by @phillipleblanc in #7624
Make DuckDB table partition data write threshold configurable by @sgrebnov in #7626
Handle table relations in HTTP v1/search by @Jeadie in #7615
Fix E2E Test sporadic failures on the macOS runners by @phillipleblanc in #7627
Emit query_active_count metric by @krinart in #7589
fix: Disable go cache in actions by @peasee in #7631
fix: Don't nullify DuckDB release callbacks for schemas by @peasee in #7628
Fix integration tests by reverting the use of batch inserts w/ prepared statements by @phillipleblanc in #7630
Split integration tests into 3 partitions by @phillipleblanc in #7635
Initial Pepper data accelerator by @lukekim in #7592
Only build the E2E Test CI binaries once by @phillipleblanc in #7633
Update BytesProcessedExec snapshots by @mach-kernel in #7637
Properly set RequestContext for stream execution in Flight by @krinart in #7591
Add task for creating release branch in docs by @kczimm in #7642
Add missing mongodb params by @krinart in #7647
fix: Update benchmark snapshots by @app/github-actions in #7649
Release notes for v1.8.2 by @Jeadie in #7645
docs: Update error handling in copilot instructions by @peasee in #7652
Pepper accelerator INSERT OVERWRITE support by @lukekim in #7643
fix: Update benchmark snapshots by @app/github-actions in #7650
Run Datafusion queries on a separate Tokio runtime by @phillipleblanc in #7586
Add explicit steps for docs DRI in end game by @kczimm in #7658
Add Release 1.8.2 QA Analytics by @krinart in #7661
Pepper full / append refresh support by @lukekim in #7662
Add 'client_timeout' for s3 vector by @Jeadie in #7501
Improvements to Endgame template by @krinart in #7660
Fix OSS docker release trigger when release marked as latest by @phillipleblanc in #7668
Use '#[serde(deny_unknown_fields)]' for base spicepod components by @Jeadie in #7669
DataFusion upgrade template to include Ballista by @mach-kernel in #7679
S3 Vector index spilling by @kczimm in #7613
Refresh request context bindings / fix trunk integration tests by @mach-kernel in #7680
Fix integration tests for refresh append by @lukekim in #7681
Distributed query support by @mach-kernel in #7585
Run acceleration refreshes on separate Tokio runtime by @phillipleblanc in #7671
Support DESCRIBE in clustered mode by @kczimm in #7686
use hyphen instead of period for index name spill separator by @kczimm in #7697
Add streaming option to /nsql endpoint by @kczimm in #7695
Task History min_sql_duration filter support by @lukekim in #7698
Spawn object_store IO tasks on the original Tokio runtime by @phillipleblanc in #7689
Update openapi.json by @app/github-actions in #7700
Adjust DataFusion runtime worker threads by @phillipleblanc in #7704
Gate dedicated SQL engine CPU runtime behind opt-in param dedicated_thread_pool by @phillipleblanc in #7705
Add TPCH S3 refresh spicepod by @phillipleblanc in #7706
DuckDB: include ANALYZE after write to update query optimizer statistics by @sgrebnov in #7714
Display execution time in Spice REPL for no results by @sgrebnov in #7713
Task History capture and store SQL query plans by @lukekim in #7701
Pepper TPC-H SF-1 benchmark by @lukekim in #7717
fix: Update benchmark snapshots by @app/github-actions in #7720
Optimize prepared statements (parameterized queries) by @lukekim in #7703
Pepper accelerator tests (Clickbench, TPC-H SF-5, SF-100) by @lukekim in #7721
Add support for DuckDB connection_pool_size param by @sgrebnov in #7716
Add health probing to testoperator runs by @phillipleblanc in #7709
Simplify AcceleratedTable by @Jeadie in #7724
Add some basic indexing tests for 'FullTextDatabaseIndex' by @Jeadie in #7688
DuckDB: on_refresh_recompute_statistics param + ANALYZE for table-based partitioning by @sgrebnov in #7719
Fix Windows builds by excluding pepper/vortex by @phillipleblanc in #7729
Enable separate CPU runtime thread pool for DataFusion by default by @phillipleblanc in #7732
'runtime-datafusion' crate for runtime related DataFusion components by @Jeadie in #7666
Pepper expanded append refresh support by @lukekim in #7670
Pepper basic partitioning by @lukekim in #7731
Stable pepper benchmark snapshots by @phillipleblanc in #7739
Delta Lake Connector: Support AWS_SESSION_TOKEN parameter by @mach-kernel in #7752
Pepper use SQLite WAL by @lukekim in #7757
v1.8.3 release notes by @mach-kernel in #7745
Increase testoperator health check threshold to 50ms by @phillipleblanc in #7767
Data Accelerator Graceful Shutdown by @lukekim in #7756
Remove Windows CUDA builds by @phillipleblanc in #7768
fix: Update benchmark snapshots by @app/github-actions in #7771
Bump actions/download-artifact from 5 to 6 by @app/dependabot in #7746
Bump serde from 1.0.226 to 1.0.228 by @app/dependabot in #7743
Fix casing for keywords and additional columns by @Jeadie in #7770
Bump actions/upload-artifact from 4 to 5 by @app/dependabot in #7750
Bump criterion from 0.5.1 to 0.7.0 by @app/dependabot in #7740
Bump rustls-native-certs from 0.8.1 to 0.8.2 by @app/dependabot in #7744
Git Data Connector (Alpha) by @lukekim in #7772
Pepper accelerator delete support by @lukekim in #7616
Update Helm chart instructions for Helm in end_game.md by @sgrebnov in #7776
Turso data accelerator by @lukekim in #7472
Apply retention SQL filter to refresh fetch by @phillipleblanc in #7778
Add Parquet buffering option for DuckDB partitioned writes (tables mode) by @sgrebnov in #7735
fix: EmptyExec when list indexes is empty by @kczimm in #7784
1.8.3 post-release housekeeping by @mach-kernel in #7783
feat: Upgrade to Datafusion v50 by @peasee in #7777
fix: Replace vortex datafusion with public crate by @peasee in #7791
Full-text search on views by @Jeadie in #7733
Revert "Apply retention SQL filter to refresh fetch (#7778)" by @phillipleblanc in #7796
fix: Add ingest duration and acceleration size metrics to testoperator by @peasee in #7792
Set local timezone to UTC for DuckDB by @phillipleblanc in #7797
add Timestamp support for partition expressions by @kczimm in #7803
Fix trunk lint by @krinart in #7804
Add missing mongodb params by @krinart in #7807
Embedding columns on view components by @Jeadie in #7795
Add Turso as a Pepper Catalog metastore by @lukekim in #7793
Run retention_sql on refresh commit for DuckDB by @lukekim in #7785
docs: Update datafusion upgrade checklist by @peasee in #7812
Vector engines on views by @Jeadie in #7808
Handle refresh worker panics and add recovery test by @phillipleblanc in #7815
chunk large record batches to control memory usage by @kczimm in #7802
fix: cannot determine vector dimension for partitioned indexes by @kczimm in #7818
Upgrade to Turso v0.3 by @lukekim in #7821
fix: Ensure custom *Exec ExecutionPlans push down dynamic filters by @peasee in #7811
handle casing in RRF by @Jeadie in #7825
Enable 'turso' for pepper acceleration by default by @sgrebnov in #7826
Improved DynamoDB Data Connector by @krinart in #7715
Initial support for llama.cpp as LLM inference backend by @lukekim in #7794
Pepper: Implement retention SQL on refresh commit by @phillipleblanc in #7814
Fix Dockerfiles for arm64 by @lukekim in #7834
[DynamoDB] Handle filter edge-cases by @krinart in #7830
[DynamoDB] Support parallelization for Scan request by @krinart in #7829
Don't feature gate Pepper by @lukekim in #7832
Fix llama.cpp static link by @lukekim in #7835
fix: docker nightly builds by @kczimm in #7837
Use GitHub-hosted macOS runner only for tag releases by @lukekim in #7836
Fix Bug: DuckDB INTERNAL Error: Failed to load metadata pointer by @sgrebnov in #7839
Fix docker arm64 build to use aegis in pure-rust mode by @lukekim in #7840
Revert "Use GitHub-hosted macOS runner only for tag releases" by @lukekim in #7843
Rename Pepper to Cayenne by @lukekim in #7844
Tighten CLI permissions and install script by @lukekim in #7845
Set mvcc for Cayenne Turso metastore by @lukekim in #7850
Optimize Prepared Statements by @lukekim in #7859
Remove unwrap from ODBC connector, fix secrets, and kuberenetes secre… by @lukekim in #7846
Improve and secure HTTP client usage by @lukekim in #7847
Pin Oracle Instant Client download to a SHA by @lukekim in #7851
Improve experience for missing or invalid Spicepod.yaml by @lukekim in #7849
chore: Fix PR linting by @peasee in #7865
Revert FlightIPC issues by @Jeadie in #7870
Bump Jimver/cuda-toolkit from 0.2.28 to 0.2.29 by @app/dependabot in #7878
Optimize macOS and Windows builds by @lukekim in #7863
Improve error message by adding 'cayenne' to the list of valid accelerator engines by @sgrebnov in #7882
fix: Kafka message delivery failed by @kczimm in #7883
fix: allow parameter index without dollar signs by @kczimm in #7887
docs: Update component criteria by @peasee in #7891
Temporary disable supports_limit_pushdown for SchemaCastScanExec by @sgrebnov in #7893
fix: Make integration run with no relevant changes, disable makefile targets on PR by @peasee in #7896
Add Cayenne benchmark and concurrency tests and remove indexes for Turso MVCC by @lukekim in #7879
Remove '.embeddings[].metadata' by @Jeadie in #7897
Revert llama.cpp engine by @lukekim in #7898
Make Cayenne snapshotting more robust by @sgrebnov in #7899
Release notes v1.9.0-rc1 by @Jeadie in #7902
Fix dataset_acceleration_last_refresh_time_ms unit to milliseconds in description by @ewgenius in #7901
Fix lint error in record_explain_plan functionality by @sgrebnov in #7906
Cleanup old snapshots after full refresh by @lukekim in #7908
Cayenne deletion vector caching support by @lukekim in #7903
Split filters into partition filters (for pruning) and data filters by @lukekim in #7889
fix: Update benchmark snapshots by @app/github-actions in #7911
fix: Update benchmark snapshots by @app/github-actions in #7912
fix: Update benchmark snapshots by @app/github-actions in #7913
Update spicepod.schema.json by @app/github-actions in #7916
fix: Update benchmark snapshots by @app/github-actions in #7917
Add Cayenne & Turso accelerators to E2E CI test matrix by @lukekim in #7922
Make preview warnings consistent by @lukekim in #7921
Filter and write optimizations by @lukekim in #7918
fix: Set sccache region explicitly by @peasee in #7928
fix: Enable integration test merge group checks by @peasee in #7927
Update testoperator release branch from 1.8 to 1.9 by @peasee in #7926
Update DuckDB to 1.4.1 with composite ART scans by @mach-kernel in #7884
Don't build Windows on trunk pushes by @lukekim in #7931
fix: Use correct minio secret in build binary push by @peasee in #7934
Update test-framework workers to use dedicated Flight client by @sgrebnov in #7938
Fix financebench, configure s3vectors for appropriate snapshotting by @Jeadie in #7935
Don't try to initialize accelerator if it is disabled by @lukekim in #7932
Add spark UDFs to Spice by @Jeadie in #7936
Fix extra async_trait in cayenne metadata catalog by @phillipleblanc in #7942
deps: Upgrade to Rust 1.90 by @peasee in #7941
Add cayenne accelerator to README.md by @ewgenius in #7905
fix: Update benchmark snapshots by @app/github-actions in #7948
Run integration tests with AWS_EC2_METADATA_DISABLED flag by @sgrebnov in #7952
Only retry credentials on ConnectorError by @kczimm in #7944
fix: Improve join reordering by ensuring JoinSelection is applied by @peasee in #7828
fix: Remove unused deps, consolidate workspace deps by @peasee in #7953
bump async-openai commit by @kczimm in #7929
deps: Use vortex fork by @peasee in #7954
Enable tracing in delta lake integration tests by @sgrebnov in #7951
Update datasets in S3 vectors test case by @Jeadie in #7956
Add spiced metrics scraping to test operator by @lukekim in #7937
Memoize S3 vectors ListIndex API call with configurable TTL by @kczimm in #7910
Cayenne performance optimizations by @lukekim in #7907
Setup HotFix issue template by @ewgenius in #7957
Fix AWS SDK credential cache retry handling by @phillipleblanc in #7943
Infer RRF join_key from TableProvider::constraints and implement SearchQueryProvider::constraints. by @Jeadie in #7959
[Optimizer]: DuckDB intermediate materialization (non-federated) by @mach-kernel in #7964
1.7.3 post-release housekeeping by @ewgenius in #7962
Fix digest_many UDF for ColumnarValue::Array. by @Jeadie in #7960
Fix spiced metrics reporting as part of benchmark tests by @sgrebnov in #7967
Avoid pushing down Spice specific UDFs to accelerators during federation by @Jeadie in #7909
CLI file persisted history with .clear and .clear history commands by @lukekim in #7970
ResultsCache Cache-Control stale-while-revalidate by @lukekim in #7963
Use GetVectors API instead of returnData by @kczimm in #7083
Make DuckDB intermediate materialization logic more robust by @sgrebnov in #7971
[Cayenne] Configurable target Vortex file size by @lukekim in #7972
fix: Update benchmark snapshots by @app/github-actions in #7974
Bump github.com/klauspost/compress from 1.17.11 to 1.18.1 by @app/dependabot in #7872
fix: Update benchmark snapshots by @app/github-actions in #7978
fix: Update benchmark snapshots by @app/github-actions in #7982
Run Integration tests on spiceai-dev-runners by @sgrebnov in #7985
[CLI] Fix cursor issue due to flush by @lukekim in #7981
fix: Support S3 versioning, Vortex dynamic filter pushdown by @peasee in #7984
Make cluster a default feature by @lukekim in #7994
Optimize DuckDB Intermediate Index Materialization for No-Index Case by @sgrebnov in #7998
HTTP connector with dynamic filter support by @lukekim in #7969
Revert federation 'can_execute_plan' by @Jeadie in #7999
Fix stale caching by @lukekim in #7995
Fix count(*) for http connector by @krinart in #8001
[CLI] Install specific version by @lukekim in #8006
Fix stale with revalidate request/response by @lukekim in #8005
Fallback RequestContext for cluster queries by @Jeadie in #8007
Use use_rustls_tls for Spice Cloud /connect by @lukekim in #8008
Use delta-kernel-rs 0.16x + Parquet reader with object meta API changes by @mach-kernel in #8011
fix: Update datafusion & arrow-rs with S3 versioning fix by @lukekim in #8012
Add 1.9.0-rc.2 release notes by @sgrebnov in #7993
Update Datafusion version by @sgrebnov in #8014
[Acceleration] DuckDB tables mode partitioner + CTE rewrite optimizer by @mach-kernel in #8013
Update spicepod.schema.json by @app/github-actions in #8015
Update acknowledgements by @app/github-actions in #8016
Upgrade shutdown signal Ordering by @krinart in #8017
Set max-age: 0 during stale by @lukekim in #8018
Add E2E test release for Helm by @lukekim in #8023
Bump github.com/olekukonko/tablewriter from 0.0.5 to 1.1.1 by @app/dependabot in #7989
Bump schemars from 0.9.0 to 1.0.4 by @app/dependabot in #7877
Update generate_changelog script by @krinart in #8028
Update QA analytics Release 1.9.0-rc.2 by @krinart in #8027
[CLI] Improve auto-complete by @lukekim in #8022
Improve verify helm workflow by @lukekim in #8024
Bump azure_core from 0.28.0 to 0.30.0 by @app/dependabot in #7986
Test operator load test row count validation by @lukekim in #8036
fix: Revert HTTP response offloading by @peasee in #8041
Disable advanced filters pruning for partitioned tables by @sgrebnov in #8037
fix: Ensure Vortex UncompressedSizeInBytes is calculated by @peasee in #8044
Add 1.9.0-rc.3 release notes by @sgrebnov in #8048
fix: Update test snapshots by @app/github-actions in #8046
add benchmark spicepods by @Jeadie in #8047
DynamoDB TPC-H SF1 Benchmarks by @krinart in #8043
Bump github.com/AzureAD/microsoft-authentication-library-for-go from 1.5.0 to 1.6.0 by @app/dependabot in #7988
Bump golang.org/x/sys from 0.37.0 to 0.38.0 by @app/dependabot in #7987
v1.9.0-rc.2 README updates by @lukekim in #8035
Bump suppaftp from 5.4.0 to 6.3.0 by @app/dependabot in #7875
Bump ctor from 0.5.0 to 0.6.0 by @app/dependabot in #7873
WW README Update by @wyattwenzel in #8058
Reenable dynamic federation support by @Jeadie in #8026
fix: Prevent SortExec from ordering below SchemaCastScanExec by @peasee in #8061
Skip logging and return OK() on error during shutdown by @krinart in #8057
Partition pruning with complex expressions by @lukekim in #8040
Update openapi.json by @app/github-actions in #8064
Make DynamoDB snapshots consistent by @krinart in #8069
Add check for error log by @krinart in #8070
Fix tracing of 's3_vector_query_and_get' by @Jeadie in #8065
DuckDB v1.4.2 by @mach-kernel in #8073
Fix failing OpenAI test by @krinart in #8076
Enable 'test_recency_scoring' by @Jeadie in #8068
Test operator: avoid duplicate Flight requests when using --http-clients by @sgrebnov in #8071
Update load tests to use truth percentile values by @sgrebnov in #8079
Update DynamoDB to RC by @krinart in #8060
CachedQueryVector to avoid recomputing embedding vector for spilling/partitioned vector indexes. by @Jeadie in #8059
Fix DuckDB on_commit sink race by @lukekim in #8081
Add partitioned duckdb by @lukekim in #8083
[CLI] Security and santization by @lukekim in #8082
fix: Update benchmark snapshots by @app/github-actions in #8084
Fix partition_by expression by @lukekim in #8087
Data Components security fixes and sanitization by @lukekim in #8086
Runtime security and sanitization by @lukekim in #8088
Add spicepod-validator tool and fix spicepods by @lukekim in #8089
Skip data fetches for S3 single file refreshes by @lukekim in #8072
MCP security and sanitization by @lukekim in #8090
Update spicepod.schema.json by @app/github-actions in #8099
Update acknowledgements by @app/github-actions in #8098
Add install-dev target back to Makefile by @Jeadie in #8100
fix 'testoperator run search' by @Jeadie in #8101
Update datafusion-table-providers - fix nullability inferences for MySQL and PostgreSQL, and fix full text search for PostgreSQL by @ewgenius in #8092
Remove duplicate install-with-models by @phillipleblanc in #8107
Improve Cayenne partitioning by @lukekim in #8097
Testoperator dispatch: respect verify_results dispatch configuration by @sgrebnov in #8106
Include 'match' column only if chunk offsets found in seach query 'LogicalPlan' by @Jeadie in #8102
Fix validation path by @lukekim in #8109
Fix dispatch paths by @lukekim in #8110
Fix dispatch spicepod paths by @lukekim in #8112
fix: Update benchmark snapshots by @app/github-actions in #8113
fix: Update benchmark snapshots by @app/github-actions in #8114
fix: Update benchmark snapshots by @app/github-actions in #8116
Update test Spicepods by @lukekim in #8131
Add validation to reference schema by @lukekim in #8111
Include root error when failing to find latest timestamp in accelerated table by @sgrebnov in #8132
fix: HTTP Connector validation, query and body by @lukekim in #8115
Update nsql model list by @lukekim in #8141
Update DynamoDB Benchmarks by @krinart in #8135
Fix Dremio E2E test by @sgrebnov in #8139
fix: Update MongoDB benchmark snapshots by @app/github-actions in #8143
fix: Update DynamoDB benchmark snapshots by @app/github-actions in #8142
fix: Update benchmark snapshots by @app/github-actions in #8145
fix: Update iceberg[catalog] benchmark snapshots by @app/github-actions in #8144
Improve HTTP Connector UX by @lukekim in #8146
QueryOverrides for DynamoDB benchmarks by @krinart in #8151
test-framework: add row count validation skipping with TPC-DS defaults by @sgrebnov in #8149
fix: Update benchmark snapshots by @app/github-actions in #8148
fix: Update benchmark snapshots by @app/github-actions in #8154
fix: Update benchmark snapshots by @app/github-actions in #8155
fix: Update test snapshots by @app/github-actions in #8160
Suppress delta_kernel::listed_log_files warnings by @phillipleblanc in #8158
Update table providers to fix warning by @phillipleblanc in #8156
Suppress MCP limit log by @phillipleblanc in #8159
Remove incorrect tool name validation by @Jeadie in #8161
Disable results validation for federated/glue[csv].yaml by @phillipleblanc in #8163
fix: Update benchmark snapshots by @app/github-actions in #8164
Fix dynamodb overrides by @phillipleblanc in #8165
Fix dynamo db overrides again by @phillipleblanc in #8166
few more dynamodb overrides by @phillipleblanc in #8167
Add stub release notes for v1.9.0-rc.4 by @phillipleblanc in #8168
Add v1.9.0-rc.4 release notes by @lukekim in #8169
fix: Cayenne concurrent table creation by @lukekim in #8176
fix: Avoid pruning bucket partitions for != and NOT IN operators by @sgrebnov in #8177

Spice v1.9.0-rc.4 (Nov 18, 2025)

November 18, 2025 · 22 min read

Phillip LeBlanc

Co-Founder and CTO of Spice AI

Announcing the release of Spice v1.9.0-rc.4! 🌶

This release candidate brings DuckDB v1.4.2, Cayenne partitioning improvements, and comprehensive security hardening across the CLI, data connectors, runtime, and MCP. v1.9.0-rc.4 also includes MySQL and PostgreSQL connector improvements with fixed nullability inferences and full-text search support, DynamoDB consistency improvements, HTTP connector validation and UX enhancements, and numerous reliability and performance optimizations. Significant improvements were also made to test and automation infrastructure to ensure high quality releases.

v1.9.0 introduces Spice Cayenne, a new high-performance data accelerator built on the Vortex columnar format that delivers better than DuckDB performance without single-file scaling limitations, and a preview of Multi-Node Distributed Query based on Apache Ballista. v1.9.0 also upgrades to DataFusion v50 for even higher query performance, expands search capabilities with full-text search on views and multi-column embeddings, and delivers many additional features and improvements.

What's New in v1.9.0

Cayenne Data Accelerator (Beta)

Key Features:

SQLite + Vortex Architecture: All metadata is stored in SQLite tables with standard SQL transactions, while data lives in Vortex's compressed, chunked columnar format designed for zero-copy access and efficient scanning.
Simplified Operations: No complex file hierarchies, no JSON/Avro metadata files, no separate catalog servers—just SQL tables and Vortex data files. The entire metadata schema is intentionally simple for maximum reliability.
Fast Metadata Access: Single SQL query retrieves all metadata needed for query planning—no multiple round trips to storage, no S3 throttling, no reconstruction of metadata state from scattered files.
Efficient Small Changes: Dramatically reduces small file proliferation. Snapshots are just rows in SQLite tables, not new files on disk. Supports millions of snapshots without performance degradation.
High Concurrency: Changes consist of two steps: stage Vortex files (if any), then run a single SQL transaction. Much faster conflict resolution and support for many more concurrent updates than file-based formats.
Advanced Data Lifecycle: Full ACID transactions, delete support, and retention SQL execution on refresh commit.

Example Spicepod.yml configuration:

datasets:
  - from: s3:my_table
    name: accelerated_data_30d
    acceleration:
      enabled: true
      engine: cayenne
      mode: file
      refresh_mode: append
      retention_sql: DELETE FROM accelerated_data WHERE created_at < NOW() - INTERVAL '30 days'

Note, the Cayenne Data Accelerator is in Beta with limitations.

For more details, refer to the Cayenne Documentation, the Vortex project, and the DuckLake announcement that partly inspired this design.

Multi-Node Distributed Query (Preview)

Architecture:

A distributed Spice cluster consists of:

Scheduler: Responsible for distributed query planning and work queue management for the executor fleet
Executors: One or more nodes responsible for running physical query plans

Getting Started:

Start a scheduler instance using an existing Spicepod. The scheduler is the only spiced instance that needs to be configured:

# Start scheduler (note the flight bind address override if you want it reachable outside localhost)
spiced --cluster-mode scheduler --flight 0.0.0.0:50051

Start one or more executors configured with the scheduler's flight URI:

# Start executor (automatically selects a free port if 50051 is taken)
spiced --cluster-mode executor --scheduler-url spiced://localhost:50051

Query Execution:

Queries run through the scheduler will now show a distributed_plan in EXPLAIN output, demonstrating how the query is distributed across executor nodes:

EXPLAIN SELECT count(id) FROM my_dataset;

Current Limitations:

Accelerated datasets are currently not supported. This feature is designed for querying partitioned data lake formats (Parquet, Delta Lake, Iceberg, etc.)
The feature is in preview and may have stability or performance limitations
Specific acceleration support is planned for future releases

DataFusion v50 Upgrade

Spice.ai is built on the Apache DataFusion query engine. The v50 release brings significant performance improvements and enhanced reliability:

Performance Improvements 🚀:

Dynamic Filter Pushdown: Enhanced dynamic filter pushdown for custom ExecutionPlans, ensuring filters propagate correctly through all physical operators for improved query performance.
Partition Pruning: Expanded partition pruning support ensures that unnecessary partitions are skipped when filters are not used, reducing data scanning overhead and improving query execution times.

See the Apache DataFusion 50.0.3 Release for more details.

DuckDB v1.4.2 Upgrade and Accelerator Improvements

DuckDB v1.4.2: DuckDB has been upgraded to v1.4.2, which includes several performance optimizations.

Example configuration:

datasets:
  - from: file://data.parquet
    name: sales
    acceleration:
      enabled: true
      engine: duckdb
      indexes:
        '(region, product_id)': enabled

Performance example with composite index on 7.5M rows:

SELECT * FROM sales WHERE region = 'US' AND product_id = 12345;

-- Without index: 0.282s
-- With composite index (region, product_id): 0.037s
-- Performance improvement: 7.6x faster with composite index

Example configuration:

datasets:
  - from: file://sales_data.parquet
    name: sales
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      params:
        query_federation: disabled # Required currently for intermediate materialization
      indexes:
        '(region, product_id)': enabled

Performance example:

-- Query with indexed columns (region, product_id) plus additional filter (amount)
SELECT * FROM sales
WHERE region = 'US' AND product_id = 12345 AND amount > 1000;

-- Optimized execution time: 0.031s (with intermediate materialization)
-- Standard execution time: 0.108s (without optimization)
-- Performance improvement: ~3.5x faster

The optimizer automatically rewrites the query to:

WITH _intermediate_materialize AS MATERIALIZED (
  SELECT * FROM sales WHERE region = 'US' AND product_id = 12345
)
SELECT * FROM _intermediate_materialize WHERE amount > 1000;

Parquet Buffering for Partitioned Writes: DuckDB partitioned writes in table mode now support Parquet buffering, reducing memory usage and improving write performance for large datasets.

Retention SQL on Refresh Commit: DuckDB accelerations now support running retention SQL on refresh commit, enabling automatic data cleanup and lifecycle management during refresh operations.

UTC Timezone for DuckDB: DuckDB now uses UTC as the default timezone, ensuring consistent behavior for time-based queries across different environments.

Example Spicepod.yml configuration:

datasets:
  - from: s3://my_bucket/large_table/
    name: partitioned_data
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      retention:
        sql: DELETE FROM partitioned_data WHERE event_time < NOW() - INTERVAL '7 days'

HTTP Data Connector

Querying endpoints as tables: The HTTP/HTTPS Data Connectors now supports querying HTTP endpoints directly as tables in SQL queries with dynamic filters. This feature transforms REST APIs into queryable data sources, making it easy to integrate external service data.
Query HTTP endpoint that returns structured data (JSON, CSV, etc.) as if it were a database table
Configurable retry logic, timeouts, and POST request support for more complex API interactions

Example Spicepod.yml configuration:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    params:
      file_format: json
      max_retries: 3
      client_timeout: 10s

Example SQL query:

SELECT request_path, request_query, content
FROM tvmaze
WHERE request_path = '/search/people' and request_query = 'q=michael'
LIMIT 10;

If a request_body is supplied it will be posted to the endpoint:

Example SQL query:

SELECT request_path, request_query, content
FROM tvmaze
WHERE request_path = '/search/people' and request_query = 'q=michael' and request_body = '{"name": "michael"}'
LIMIT 10;

HTTP endpoints can be accelerated using refresh_sql:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    acceleration:
      enabled: true
      refresh_mode: full
      refresh_sql: |
        SELECT request_path, request_query, content 
        FROM tvmaze
        WHERE request_path = '/search/people'
          AND request_query IN ('q=michael', 'q=luke')

DynamoDB Data Connector Improvements

Example Spicepod.yml configuration:

datasets:
  - from: dynamodb:my_table
    name: ddb_data
    params:
      scan_segments: 10 # Default `auto` which calculates optimal segments based on number of rows

S3 Versioning Support

Atomic Range Reads for Versioned Files: Spice now supports S3 Versioning for all connectors using object-store (S3, Delta Lake, etc.), ensuring range reads over versioned files are atomically correct. When S3 versioning is enabled, Spice automatically tracks version IDs during file discovery and uses them for all subsequent range reads, preventing inconsistencies from concurrent file modifications.

Current limitations:

Multi-file connections (e.g., partitioned datasets) do not yet support version tracking across all files
Version tracking is automatic when S3 versioning is enabled on the bucket

Search & Embeddings Enhancements

Multi-Column Embeddings on Views: Views now support embedding columns, enabling vector search and semantic retrieval on view data. This is useful for search over aggregated or joined datasets.

Vector Engines on Views: Vector search engines are now available for views, enabling similarity search over complex queries and transformations.

Example Spicepod.yml configuration:

views:
  - name: aggregated_reviews
    sql: SELECT review_id, review_text FROM reviews WHERE rating > 4
    embeddings:
      - column: review_text
        model: openai:text-embedding-3-small

Dedicated Query Thread Pool (Now Enabled by Default)

This feature was opt-in in previous releases and is now enabled by default. To disable it and revert to the previous behavior, add the following spicepod.yaml configuration:

runtime:
  params:
    dedicated_thread_pool: none

Query Performance Optimizations

Query Result Cache: Stale-While-Revalidate

How it works:

When a cache entry is stale but within the stale-while-revalidate window, Spice will:

Immediately return the stale cached result to the client
Asynchronously re-execute the query in the background to refresh the cache
Future requests will use the refreshed data

Configuration:

Use the Cache-Control HTTP header with the stale-while-revalidate directive:

Cache-Control: max-age=300, stale-while-revalidate=60

This configuration caches results for 5 minutes (300 seconds), and allows serving stale results for an additional 60 seconds while refreshing in the background.

Requirements:

Must use plan or raw SQL cache keys (set cache_key_type to sql or plan in results_caching configuration)
Background revalidation re-executes queries through the normal query path
Timestamp tracking automatically determines cache entry age for staleness checks

Example configuration via HTTP header:

GET /v1/sql
Cache-Control: max-age=600, stale-while-revalidate=120
X-Cache-Key-Type: sql

This feature improves application responsiveness while ensuring data freshness through background updates.

Security & Reliability Improvements

ODBC Connector Improvements: Removed unwrap calls from the ODBC connector, improving error handling and reliability. Fixed secret handling and Kubernetes secret integration.

CLI Permissions Hardening: Tightened file permissions for the CLI and install script, ensuring secure defaults for configuration files and credentials.

Oracle Instant Client Pinning: Oracle Instant Client downloads are now pinned to specific SHAs, ensuring reproducible builds and preventing supply chain attacks.

AWS Authentication Improvements

Key features:

Automatic retry with backoff: Implements Fibonacci backoff for transient credential failures (network issues, temporary AWS service disruptions)
Configurable retry limits: Supports up to 300 retry attempts with a maximum retry interval of 600 seconds
Better error handling: Distinguishes between retryable errors (connector errors) and non-retryable errors (misconfiguration)
Unauthenticated access support: Properly supports unauthenticated access to public S3 buckets without requiring credentials
Improved error messages: Provides detailed logging with attempt numbers, retry intervals, and error context for better troubleshooting

The improvements ensure more reliable AWS service integration, particularly in environments with intermittent network connectivity or during AWS service degradations.

Observability & Tracing

DataFusion Log Emission: The Spice runtime now emits DataFusion internal logs, providing deeper visibility into query planning and execution for debugging and performance analysis.

AI Completions Tracing: Fixed tracing so that ai_completions operations are correctly parented under sql_query traces, improving observability for AI-powered queries.

Git Data Connector (Alpha)

Example Spicepod.yml configuration:

datasets:
  - from: git:https://github.com/myorg/myrepo
    name: git_metrics
    params:
      file_format: csv

For more details, refer to the Git Data Connector Documentation.

Spice Java SDK 0.4.0

The Spice Java SDK have been upgraded with support configurable Arrow memory limit: spice-java v0.4.0

SpiceClient client = SpiceClient.builder()
    .withArrowMemoryLimitMB(1024) // 1GB limit
    .build();

CLI Improvements

Usage:

# Install a specific version
spice install v1.8.3

# Install a specific version with AI flavor
spice install v1.8.3 ai

# Install latest version (existing behavior)
spice install
spice install ai

Note: Homebrew installations require manual version management via brew install spiceai/spiceai/spice@<version>.

New REPL Commands:

.clear - Clear the screen using ANSI escape codes for a clean workspace
.clear history - Clear and persist the query history, removing all stored commands

Tab Completion: Tab completion now includes suggestions based on your command history, making it faster to re-run or modify previous queries.

Example usage:

sql> SELECT * FROM my_table;
sql> .clear          # Clears the screen
sql> .clear history  # Clears command history
sql> # Use arrow keys or tab to access previous commands

Additional Improvements & Bug Fixes

Reliability: Fixed refresh worker panics with recovery handling to prevent runtime crashes during acceleration refreshes.
Reliability: Improved error messages for missing or invalid spicepod.yaml files, providing actionable feedback for misconfiguration.
Reliability: Fixed DuckDB metadata pointer loading issues for snapshots.
Performance: Ensured ListingTable partitions are pruned correctly when filters are not used.
Reliability: Fixed vector dimension determination for partitioned indexes.
Search: Fixed casing issues in Reciprocal Rank Fusion (RRF) for hybrid search queries.
Search: Fixed search field handling as metadata for chunked search indexes.
Validation: Added timestamp support for partition expressions.
Validation: Fixed regexp_match function for DuckDB datasets.
Validation: Fixed partition name validation for improved reliability.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

New HTTP Data Connector Recipe: New recipe demonstrating how to query REST APIs and HTTP(s) endpoints. See HTTP Connector Recipe for details.

The Spice Cookbook includes 82 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.9.0-rc.4, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.9.0-rc.4 image:

docker pull spiceai/spiceai:1.9.0-rc.4

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Dependencies

DataFusion: Upgraded to v50
Apache Arrow: Upgraded to v56
DuckDB: Upgraded to v1.4.2
Delta Kernel: Upgraded to v0.16.0

Changelog (rc.4)

Upgrade shutdown signal Ordering by @krinart in #8017
Set max-age: 0 during stale by @lukekim in #8018
Add E2E test release for Helm by @lukekim in #8023
Update generate_changelog script by @krinart in #8028
[CLI] Improve auto-complete by @lukekim in #8022
Improve verify helm workflow by @lukekim in #8024
fix: Ensure Vortex UncompressedSizeInBytes is calculated by @peasee in #8044
WW README Update by @wyattwenzel in #8058
Reenable dynamic federation support by @Jeadie in #8026
fix: Prevent SortExec from ordering below SchemaCastScanExec by @peasee in #8061
Skip logging and return OK() on error during shutdown by @krinart in #8057
Partition pruning with complex expressions by @lukekim in #8040
Make DynamoDB snapshots consistent by @krinart in #8069
Add check for error log by @krinart in #8070
Fix tracing of 's3_vector_query_and_get' by @Jeadie in #8065
DuckDB v1.4.2 by @mach-kernel in #8073
Fix failing OpenAI test by @krinart in #8076
Enable 'test_recency_scoring' by @Jeadie in #8068
Test operator: avoid duplicate Flight requests when using --http-clients by @sgrebnov in #8071
Update load tests to use truth percentile values by @sgrebnov in #8079
Update DynamoDB to RC by @krinart in #8060
CachedQueryVector to avoid recomputing embedding vector for spilling/partitioned vector indexes. by @Jeadie in #8059
Fix DuckDB on_commit sink race by @lukekim in #8081
Add partitioned duckdb by @lukekim in #8083
[CLI] Security and santization by @lukekim in #8082
Fix partition_by expression by @lukekim in #8087
Data Components security fixes and sanitization by @lukekim in #8086
Runtime security and sanitization by @lukekim in #8088
Add spicepod-validator tool and fix spicepods by @lukekim in #8089
Skip data fetches for S3 single file refreshes by @lukekim in #8072
MCP security and sanitization by @lukekim in #8090
Add install-dev target back to Makefile by @Jeadie in #8100
fix 'testoperator run search' by @Jeadie in #8101
Update datafusion-table-providers - fix nullability inferences for MySQL and PostgreSQL, and fix full text search for PostgreSQL by @ewgenius in #8092
Remove duplicate install-with-models by @phillipleblanc in #8107
Improve Cayenne partitioning by @lukekim in #8097
Testoperator dispatch: respect verify_results dispatch configuration by @sgrebnov in #8106
Include 'match' column only if chunk offsets found in seach query 'LogicalPlan' by @Jeadie in #8102
Fix validation path by @lukekim in #8109
Fix dispatch paths by @lukekim in #8110
Fix dispatch spicepod paths by @lukekim in #8112
Update test Spicepods by @lukekim in #8131
Add validation to reference schema by @lukekim in #8111
Include root error when failing to find latest timestamp in accelerated table by @sgrebnov in #8132
fix: HTTP Connector validation, query and body by @lukekim in #8115
Update nsql model list by @lukekim in #8141
Update DynamoDB Benchmarks by @krinart in #8135
Fix Dremio E2E test by @sgrebnov in #8139
Improve HTTP Connector UX by @lukekim in #8146
QueryOverrides for DynamoDB benchmarks by @krinart in #8151
test-framework: add row count validation skipping with TPC-DS defaults by @sgrebnov in #8149
Suppress delta_kernel::listed_log_files warnings by @phillipleblanc in #8158
Update table providers to fix warning by @phillipleblanc in #8156
Suppress MCP limit log by @phillipleblanc in #8159
Remove incorrect tool name validation by @Jeadie in #8161
Disable results validation for federated/glue[csv].yaml by @phillipleblanc in #8163
Fix dynamodb overrides by @phillipleblanc in #8165
Fix dynamo db overrides again by @phillipleblanc in #8166
few more dynamodb overrides by @phillipleblanc in #8167
Add stub release notes for v1.9.0-rc.4 by @phillipleblanc in #8168

Spice v1.9.0-rc.2 (Nov 11, 2025)

November 11, 2025 · 32 min read

Sergei Grebnov

Senior Software Engineer at Spice AI

Announcing the release of Spice v1.9.0-rc.2! 🌶

This is the second release candidate for v1.9.0, which introduces Spice Cayenne, a new high-performance data accelerator built on the Vortex columnar format that delivers better than DuckDB performance without single-file scaling limitations and a preview of Multi-Node Distributed Query based on Apache Ballista. v1.9.0-rc.2 also upgrades to DataFusion v50 and DuckDB v1.4.1 for even higher query performance, expands search capabilities with full-text search on views and multi-column embeddings, includes significant DynamoDB and DuckDB accelerator improvements, expands the HTTP data connector to support endpoints as tables, and delivers many security and reliability improvements.

What's New in v1.9.0-rc.2

Cayenne Data Accelerator (Beta)

Key Features:

SQLite + Vortex Architecture: All metadata is stored in SQLite tables with standard SQL transactions, while data lives in Vortex's compressed, chunked columnar format designed for zero-copy access and efficient scanning.
Simplified Operations: No complex file hierarchies, no JSON/Avro metadata files, no separate catalog servers—just SQL tables and Vortex data files. The entire metadata schema is intentionally simple for maximum reliability.
Fast Metadata Access: Single SQL query retrieves all metadata needed for query planning—no multiple round trips to storage, no S3 throttling, no reconstruction of metadata state from scattered files.
Efficient Small Changes: Dramatically reduces small file proliferation. Snapshots are just rows in SQLite tables, not new files on disk. Supports millions of snapshots without performance degradation.
High Concurrency: Changes consist of two steps: stage Vortex files (if any), then run a single SQL transaction. Much faster conflict resolution and support for many more concurrent updates than file-based formats.
Advanced Data Lifecycle: Full ACID transactions, delete support, and retention SQL execution on refresh commit.

Example Spicepod.yml configuration:

datasets:
  - from: s3:my_table
    name: accelerated_data_30d
    acceleration:
      enabled: true
      engine: cayenne
      mode: file
      refresh_mode: append
      retention_sql: DELETE FROM accelerated_data WHERE created_at < NOW() - INTERVAL '30 days'

Note, the Cayenne Data Accelerator is in Beta with limitations.

For more details, refer to the Cayenne Documentation, the Vortex project, and the DuckLake announcement that partly inspired this design.

Multi-Node Distributed Query (Preview)

Architecture:

A distributed Spice cluster consists of:

Scheduler: Responsible for distributed query planning and work queue management for the executor fleet
Executors: One or more nodes responsible for running physical query plans

Getting Started:

Start a scheduler instance using an existing Spicepod. The scheduler is the only spiced instance that needs to be configured:

# Start scheduler (note the flight bind address override if you want it reachable outside localhost)
spiced --cluster-mode scheduler --flight 0.0.0.0:50051

Start one or more executors configured with the scheduler's flight URI:

# Start executor (automatically selects a free port if 50051 is taken)
spiced --cluster-mode executor --scheduler-url spiced://localhost:50051

Query Execution:

Queries run through the scheduler will now show a distributed_plan in EXPLAIN output, demonstrating how the query is distributed across executor nodes:

EXPLAIN SELECT count(id) FROM my_dataset;

Current Limitations:

Accelerated datasets are currently not supported. This feature is designed for querying partitioned data lake formats (Parquet, Delta Lake, Iceberg, etc.)
The feature is in preview and may have stability or performance limitations
Specific acceleration support is planned for future releases

DataFusion v50 Upgrade

Spice.ai is built on the Apache DataFusion query engine. The v50 release brings significant performance improvements and enhanced reliability:

Performance Improvements 🚀:

Dynamic Filter Pushdown: Enhanced dynamic filter pushdown for custom ExecutionPlans, ensuring filters propagate correctly through all physical operators for improved query performance.
Partition Pruning: Expanded partition pruning support ensures that unnecessary partitions are skipped when filters are not used, reducing data scanning overhead and improving query execution times.

See the Apache DataFusion 50.0.0 Release for more details.

DuckDB v1.4.1 Upgrade and Accelerator Improvements

DuckDB v1.4.1: DuckDB has been upgraded to v1.4.1, which includes several performance optimizations.

Example configuration:

datasets:
  - from: file://data.parquet
    name: sales
    acceleration:
      enabled: true
      engine: duckdb
      indexes:
        '(region, product_id)': enabled

Performance example with composite index on 7.5M rows:

SELECT * FROM sales WHERE region = 'US' AND product_id = 12345;

-- Without index: 0.282s
-- With composite index (region, product_id): 0.037s
-- Performance improvement: 7.6x faster with composite index

Example configuration:

datasets:
  - from: file://sales_data.parquet
    name: sales
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      params:
        query_federation: disabled # Required currently for intermediate materialization
      indexes:
        '(region, product_id)': enabled

Performance example:

-- Query with indexed columns (region, product_id) plus additional filter (amount)
SELECT * FROM sales
WHERE region = 'US' AND product_id = 12345 AND amount > 1000;

-- Optimized execution time: 0.031s (with intermediate materialization)
-- Standard execution time: 0.108s (without optimization)
-- Performance improvement: ~3.5x faster

The optimizer automatically rewrites the query to:

WITH _intermediate_materialize AS MATERIALIZED (
  SELECT * FROM sales WHERE region = 'US' AND product_id = 12345
)
SELECT * FROM _intermediate_materialize WHERE amount > 1000;

Parquet Buffering for Partitioned Writes: DuckDB partitioned writes in table mode now support Parquet buffering, reducing memory usage and improving write performance for large datasets.

Retention SQL on Refresh Commit: DuckDB accelerations now support running retention SQL on refresh commit, enabling automatic data cleanup and lifecycle management during refresh operations.

UTC Timezone for DuckDB: DuckDB now uses UTC as the default timezone, ensuring consistent behavior for time-based queries across different environments.

Example Spicepod.yml configuration:

datasets:
  - from: s3://my_bucket/large_table/
    name: partitioned_data
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      retention:
        sql: DELETE FROM partitioned_data WHERE event_time < NOW() - INTERVAL '7 days'

HTTP Data Connector

Querying endpoints as tables: The HTTP/HTTPS Data Connectors now supports querying HTTP endpoints directly as tables in SQL queries with dynamic filters. This feature transforms REST APIs into queryable data sources, making it easy to integrate external service data.
Query HTTP endpoint that returns structured data (JSON, CSV, etc.) as if it were a database table
Configurable retry logic, timeouts, and POST request support for more complex API interactions

Example Spicepod.yml configuration:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    params:
      file_format: json
      max_retries: 3
      client_timeout: 10s

Example SQL query:

SELECT request_path, request_query, content
FROM tvmaze
WHERE request_path = '/search/people' and request_query = 'q=michael'
LIMIT 10;

If a request_body is supplied it will be posted to the endpoint:

Example SQL query:

SELECT request_path, request_query, content
FROM tvmaze
WHERE request_path = '/search/people' and request_query = 'q=michael' and request_body = '{"name": "michael"}'
LIMIT 10;

HTTP endpoints can be accelerated using refresh_sql:

datasets:
  - from: https://api.tvmaze.com
    name: tvmaze
    acceleration:
      enabled: true
      refresh_mode: full
      refresh_sql: |
        SELECT request_path, request_query, content 
        FROM tvmaze
        request_path = '/search/people'
          AND request_query IN ('q=michael', 'q=luke')

DynamoDB Data Connector Improvements

Example Spicepod.yml configuration:

datasets:
  - from: dynamodb:my_table
    name: ddb_data
    params:
      scan_segments: 10 # Default `auto` which calculates optimal segments based on number of rows

S3 Versioning Support

Current limitations:

Multi-file connections (e.g., partitioned datasets) do not yet support version tracking across all files
Version tracking is automatic when S3 versioning is enabled on the bucket

Search & Embeddings Enhancements

Multi-Column Embeddings on Views: Views now support embedding columns, enabling vector search and semantic retrieval on view data. This is useful for search over aggregated or joined datasets.

Vector Engines on Views: Vector search engines are now available for views, enabling similarity search over complex queries and transformations.

Example Spicepod.yml configuration:

views:
  - name: aggregated_reviews
    sql: SELECT review_id, review_text FROM reviews WHERE rating > 4
    embeddings:
      - column: review_text
        model: openai:text-embedding-3-small

Dedicated Query Thread Pool (Now Enabled by Default)

This feature was opt-in in previous releases and is now enabled by default in v1.9.0-rc.2. To disable it and revert to the previous behavior, add the following spicepod.yaml configuration:

runtime:
  params:
    dedicated_thread_pool: none

Query Performance Optimizations

Query Result Cache: Stale-While-Revalidate

How it works:

When a cache entry is stale but within the stale-while-revalidate window, Spice will:

Immediately return the stale cached result to the client
Asynchronously re-execute the query in the background to refresh the cache
Future requests will use the refreshed data

Configuration:

Use the Cache-Control HTTP header with the stale-while-revalidate directive:

Cache-Control: max-age=300, stale-while-revalidate=60

This configuration caches results for 5 minutes (300 seconds), and allows serving stale results for an additional 60 seconds while refreshing in the background.

Requirements:

Must use plan or raw SQL cache keys (set cache_key_type to sql or plan in results_caching configuration)
Background revalidation re-executes queries through the normal query path
Timestamp tracking automatically determines cache entry age for staleness checks

Example configuration via HTTP header:

GET /v1/sql
Cache-Control: max-age=600, stale-while-revalidate=120
X-Cache-Key-Type: sql

This feature improves application responsiveness while ensuring data freshness through background updates.

Security & Reliability Improvements

ODBC Connector Improvements: Removed unwrap calls from the ODBC connector, improving error handling and reliability. Fixed secret handling and Kubernetes secret integration.

CLI Permissions Hardening: Tightened file permissions for the CLI and install script, ensuring secure defaults for configuration files and credentials.

Oracle Instant Client Pinning: Oracle Instant Client downloads are now pinned to specific SHAs, ensuring reproducible builds and preventing supply chain attacks.

AWS Authentication Improvements

Key features:

Automatic retry with backoff: Implements Fibonacci backoff for transient credential failures (network issues, temporary AWS service disruptions)
Configurable retry limits: Supports up to 300 retry attempts with a maximum retry interval of 600 seconds
Better error handling: Distinguishes between retryable errors (connector errors) and non-retryable errors (misconfiguration)
Unauthenticated access support: Properly supports unauthenticated access to public S3 buckets without requiring credentials
Improved error messages: Provides detailed logging with attempt numbers, retry intervals, and error context for better troubleshooting

The improvements ensure more reliable AWS service integration, particularly in environments with intermittent network connectivity or during AWS service degradations.

Observability & Tracing

DataFusion Log Emission: The Spice runtime now emits DataFusion internal logs, providing deeper visibility into query planning and execution for debugging and performance analysis.

AI Completions Tracing: Fixed tracing so that ai_completions operations are correctly parented under sql_query traces, improving observability for AI-powered queries.

Git Data Connector (Alpha)

Example Spicepod.yml configuration:

datasets:
  - from: git:https://github.com/myorg/myrepo
    name: git_metrics
    params:
      file_format: csv

For more details, refer to the Git Data Connector Documentation.

Spice Java SDK 0.4.0

The Spice Java SDK have been upgraded with support configurable Arrow memory limit: spice-java v0.4.0

SpiceClient client = SpiceClient.builder()
    .withArrowMemoryLimitMB(1024) // 1GB limit
    .build();

CLI Improvements

Usage:

# Install a specific version
spice install v1.8.3

# Install a specific version with AI flavor
spice install v1.8.3 ai

# Install latest version (existing behavior)
spice install
spice install ai

Note: Homebrew installations require manual version management via brew install spiceai/spiceai/spice@<version>.

New REPL Commands:

.clear - Clear the screen using ANSI escape codes for a clean workspace
.clear history - Clear and persist the query history, removing all stored commands

Tab Completion: Tab completion now includes suggestions based on your command history, making it faster to re-run or modify previous queries.

Example usage:

sql> SELECT * FROM my_table;
sql> .clear          # Clears the screen
sql> .clear history  # Clears command history
sql> # Use arrow keys or tab to access previous commands

Additional Improvements & Bug Fixes

Reliability: Fixed refresh worker panics with recovery handling to prevent runtime crashes during acceleration refreshes.
Reliability: Improved error messages for missing or invalid spicepod.yaml files, providing actionable feedback for misconfiguration.
Reliability: Fixed DuckDB metadata pointer loading issues for snapshots.
Performance: Ensured ListingTable partitions are pruned correctly when filters are not used.
Reliability: Fixed vector dimension determination for partitioned indexes.
Search: Fixed casing issues in Reciprocal Rank Fusion (RRF) for hybrid search queries.
Search: Fixed search field handling as metadata for chunked search indexes.
Validation: Added timestamp support for partition expressions.
Validation: Fixed regexp_match function for DuckDB datasets.
Validation: Fixed partition name validation for improved reliability.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

New HTTP Data Connector Recipe: New recipe demonstrating how to query REST APIs and HTTP(s) endpoints. See HTTP Connector Recipe for details.

The Spice Cookbook includes 82 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.9.0-rc.2, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.9.0-rc.2 image:

docker pull spiceai/spiceai:1.9.0-rc.2

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Dependencies

DataFusion: Upgraded to v50
Apache Arrow: Upgraded to v56
DuckDB: Upgraded to v1.4.1
Delta Kernel: Upgraded to v0.16.0

Changelog

Fix for search field as metadata for chunked search indexes by @Jeadie in #7429
Bump object_store from 0.12.3 to 0.12.4 by @app/dependabot in #7433
Properly respect disabling snapshots by @phillipleblanc in #7431
Revert "Properly respect disabling snapshots" by @sgrebnov in #7439
Revert "Disable snapshots by default" by @sgrebnov in #7438
Add preview warning for write access mode by @sgrebnov in #7440
fix: regexp_match for DuckDB datasets by @kczimm in #7443
Add feature is currently in preview warning for snapshots by @sgrebnov in #7442
[Logger] Also emit Datafusion logs by @mach-kernel in #7441
add missing snapshot by @kczimm in #7446
Fix tracing so that ai_completions are parented under sql_query by @lukekim in #7415
Enable snapshot acceleration by default by @phillipleblanc in #7451
Disable acceleration refresh metrics by @krinart in #7450
Add v1.8 release notes by @phillipleblanc in #7430
fix: partition name validation by @kczimm in #7452
Fix lint error due to ignore without reasons by @krinart in #7454
Add models and CUDA support to spiced install script by @lukekim in #7457
Post-release 1.8 updates by @phillipleblanc in #7455
Remove println in datafusion by @phillipleblanc in #7461
Update end_game.md to notify once release is done by @sgrebnov in #7460
Remove italics from snapshot logging by @phillipleblanc in #7463
Update openapi.json by @app/github-actions in #7466
Fix generate spicepod schema by @phillipleblanc in #7464
Fix generate acknowledements by @phillipleblanc in #7465
Update spicepod.schema.json by @app/github-actions in #7469
fix: Ensure ListingTable partitions are pruned when filters are not used by @peasee in #7471
Create runtime-secrets crate by @phillipleblanc in #7474
Create runtime-parameters crate by @phillipleblanc in #7475
Don't download the snapshot if the acceleration is present by @phillipleblanc in #7477
Fix casing for keywords and additional columns by @Jeadie in #7770
Bump actions/upload-artifact from 4 to 5 by @app/dependabot in #7750
Bump criterion from 0.5.1 to 0.7.0 by @app/dependabot in #7740
Bump rustls-native-certs from 0.8.1 to 0.8.2 by @app/dependabot in #7744
Git Data Connector (Alpha) by @lukekim in #7772
Pepper accelerator delete support by @lukekim in #7616
Update Helm chart instructions for Helm in end_game.md by @sgrebnov in #7776
Turso data accelerator by @lukekim in #7472
Apply retention SQL filter to refresh fetch by @phillipleblanc in #7778
Add Parquet buffering option for DuckDB partitioned writes (tables mode) by @sgrebnov in #7735
fix: EmptyExec when list indexes is empty by @kczimm in #7784
1.8.3 post-release housekeeping by @mach-kernel in #7783
feat: Upgrade to Datafusion v50 by @peasee in #7777
fix: Replace vortex datafusion with public crate by @peasee in #7791
Full-text search on views by @Jeadie in #7733
Revert "Apply retention SQL filter to refresh fetch (#7778)" by @phillipleblanc in #7796
fix: Add ingest duration and acceleration size metrics to testoperator by @peasee in #7792
Set local timezone to UTC for DuckDB by @phillipleblanc in #7797
add Timestamp support for partition expressions by @kczimm in #7803
Fix trunk lint by @krinart in #7804
Add missing mongodb params by @krinart in #7807
Embedding columns on view components by @Jeadie in #7795
Add Turso as a Pepper Catalog metastore by @lukekim in #7793
Run retention_sql on refresh commit for DuckDB by @lukekim in #7785
docs: Update datafusion upgrade checklist by @peasee in #7812
Vector engines on views by @Jeadie in #7808
Handle refresh worker panics and add recovery test by @phillipleblanc in #7815
chunk large record batches to control memory usage by @kczimm in #7802
fix: cannot determine vector dimension for partitioned indexes by @kczimm in #7818
Upgrade to Turso v0.3 by @lukekim in #7821
fix: Ensure custom *Exec ExecutionPlans push down dynamic filters by @peasee in #7811
handle casing in RRF by @Jeadie in #7825
Enable 'turso' for pepper acceleration by default by @sgrebnov in #7826
Improved DynamoDB Data Connector by @krinart in #7715
Initial support for llama.cpp as LLM inference backend by @lukekim in #7794
Pepper: Implement retention SQL on refresh commit by @phillipleblanc in #7814
Fix Dockerfiles for arm64 by @lukekim in #7834
[DynamoDB] Handle filter edge-cases by @krinart in #7830
[DynamoDB] Support parallelization for Scan request by @krinart in #7829
Don't feature gate Pepper by @lukekim in #7832
Fix llama.cpp static link by @lukekim in #7835
fix: docker nightly builds by @kczimm in #7837
Use GitHub-hosted macOS runner only for tag releases by @lukekim in #7836
Fix Bug: DuckDB INTERNAL Error: Failed to load metadata pointer by @sgrebnov in #7839
Fix docker arm64 build to use aegis in pure-rust mode by @lukekim in #7840
Revert "Use GitHub-hosted macOS runner only for tag releases" by @lukekim in #7843
Rename Pepper to Cayenne by @lukekim in #7844
Tighten CLI permissions and install script by @lukekim in #7845
Set mvcc for Cayenne Turso metastore by @lukekim in #7850
Optimize Prepared Statements by @lukekim in #7859
Remove unwrap from ODBC connector, fix secrets, and kuberenetes secre… by @lukekim in #7846
Improve and secure HTTP client usage by @lukekim in #7847
Pin Oracle Instant Client download to a SHA by @lukekim in #7851
Improve experience for missing or invalid Spicepod.yaml by @lukekim in #7849
chore: Fix PR linting by @peasee in #7865
Revert FlightIPC issues by @Jeadie in #7870
Improve error message by adding 'cayenne' to the list of valid accelerator engines by @sgrebnov in #7882
fix: allow parameter index without dollar signs by @kczimm in #7887
Temporary disable supports_limit_pushdown for SchemaCastScanExec by @sgrebnov in #7893
Remove '.embeddings[].metadata' by @Jeadie in #7897
Optimize macOS and Windows builds by @lukekim in #7863
fix: Kafka message delivery failed by @kczimm in #7883
docs: Update component criteria by @peasee in #7891
fix: Make integration run with no relevant changes, disable makefile targets on PR by @peasee in #7896
Add Cayenne benchmark and concurrency tests and remove indexes for Turso MVCC by @lukekim in #7879
Revert llama.cpp engine by @lukekim in #7898
Make Cayenne snapshotting more robust by @sgrebnov in #7899
Release notes v1.9.0-rc1 by @Jeadie in #7902
Fix dataset_acceleration_last_refresh_time_ms unit to milliseconds in description by @ewgenius in #7901
Fix lint error in record_explain_plan functionality by @sgrebnov in #7906
Cleanup old snapshots after full refresh by @lukekim in #7908
Cayenne deletion vector caching support by @lukekim in #7903
Split filters into partition filters (for pruning) and data filters by @lukekim in #7889
fix: Update benchmark snapshots by @app/github-actions in #7911
fix: Update benchmark snapshots by @app/github-actions in #7912
fix: Update benchmark snapshots by @app/github-actions in #7913
Update spicepod.schema.json by @app/github-actions in #7916
fix: Update benchmark snapshots by @app/github-actions in #7917
Add Cayenne & Turso accelerators to E2E CI test matrix by @lukekim in #7922
Make preview warnings consistent by @lukekim in #7921
Filter and write optimizations by @lukekim in #7918
fix: Set sccache region explicitly by @peasee in #7928
fix: Enable integration test merge group checks by @peasee in #7927
Update testoperator release branch from 1.8 to 1.9 by @peasee in #7926
Update DuckDB to 1.4.1 with composite ART scans by @mach-kernel in #7884
Don't build Windows on trunk pushes by @lukekim in #7931
fix: Use correct minio secret in build binary push by @peasee in #7934
Update test-framework workers to use dedicated Flight client by @sgrebnov in #7938
Fix financebench, configure s3vectors for appropriate snapshotting by @Jeadie in #7935
Don't try to initialize accelerator if it is disabled by @lukekim in #7932
Add spark UDFs to Spice by @Jeadie in #7936
Fix extra async_trait in cayenne metadata catalog by @phillipleblanc in #7942
deps: Upgrade to Rust 1.90 by @peasee in #7941
Add cayenne accelerator to README.md by @ewgenius in #7905
fix: Update benchmark snapshots by @app/github-actions in #7948
Run integration tests with AWS_EC2_METADATA_DISABLED flag by @sgrebnov in #7952
Only retry credentials on ConnectorError by @kczimm in #7944
fix: Improve join reordering by ensuring JoinSelection is applied by @peasee in #7828
fix: Remove unused deps, consolidate workspace deps by @peasee in #7953
bump async-openai commit by @kczimm in #7929
deps: Use vortex fork by @peasee in #7954
Enable tracing in delta lake integration tests by @sgrebnov in #7951
Update datasets in S3 vectors test case by @Jeadie in #7956
Add spiced metrics scraping to test operator by @lukekim in #7937
Memoize S3 vectors ListIndex API call with configurable TTL by @kczimm in #7910
Cayenne performance optimizations by @lukekim in #7907
Setup HotFix issue template by @ewgenius in #7957
Fix AWS SDK credential cache retry handling by @phillipleblanc in #7943
Infer RRF join_key from TableProvider::constraints and implement SearchQueryProvider::constraints. by @Jeadie in #7959
[Optimizer]: DuckDB intermediate materialization (non-federated) by @mach-kernel in #7964
1.7.3 post-release housekeeping by @ewgenius in #7962
Fix digest_many UDF for ColumnarValue::Array. by @Jeadie in #7960
Fix spiced metrics reporting as part of benchmark tests by @sgrebnov in #7967
Avoid pushing down Spice specific UDFs to accelerators during federation by @Jeadie in #7909
CLI file persisted history with .clear and .clear history commands by @lukekim in #7970
ResultsCache Cache-Control stale-while-revalidate by @lukekim in #7963
Use GetVectors API instead of returnData by @kczimm in #7083
Make DuckDB intermediate materialization logic more robust by @sgrebnov in #7971
[Cayenne] Configurable target Vortex file size by @lukekim in #7972
fix: Update benchmark snapshots by @app/github-actions in #7974
Bump github.com/klauspost/compress from 1.17.11 to 1.18.1 by @app/dependabot in #7872
fix: Update benchmark snapshots by @app/github-actions in #7978
fix: Update benchmark snapshots by @app/github-actions in #7982
Run Integration tests on spiceai-dev-runners by @sgrebnov in #7985
[CLI] Fix cursor issue due to flush by @lukekim in #7981
fix: Support S3 versioning, Vortex dynamic filter pushdown by @peasee in #7984
Make cluster a default feature by @lukekim in #7994
Optimize DuckDB Intermediate Index Materialization for No-Index Case by @sgrebnov in #7998
HTTP connector with dynamic filter support by @lukekim in #7969
Revert federation 'can_execute_plan' by @Jeadie in #7999
Fix stale caching by @lukekim in #7995
Fix count(*) for http connector by @krinart in #8001
[CLI] Install specific version by @lukekim in #8006
Fix stale with revalidate request/response by @lukekim in #8005
Fallback RequestContext for cluster queries by @Jeadie in #8007
Use use_rustls_tls for Spice Cloud /connect by @lukekim in #8008
Use delta-kernel-rs 0.16x + Parquet reader with object meta API changes by @mach-kernel in #8011
fix: Update datafusion & arrow-rs with S3 versioning fix by @lukekim in #8012

Spice v1.9.0-rc.1 (Nov 4, 2025)

November 4, 2025 · 16 min read

William Croxson

Senior Software Engineer at Spice AI

This is the first release candidate for v1.9.0, which introduces Cayenne, a new high-performance data accelerator built on the Vortex columnar format that delivers DuckDB-comparable performance without scaling limitations. This release also upgrades to DataFusion v50 for improved query performance, expands search capabilities with full-text search on views and multi-column embeddings, includes significant DynamoDB and DuckDB accelerator improvements, and delivers security and reliability enhancements.

What's New in v1.9.0-rc.1

Cayenne Data Accelerator (Alpha)

Introducing Cayenne: SQL as an Acceleration Format: A new high-performance data accelerator that simplifies multi-file data acceleration by using an embedded database (SQLite) for metadata while storing data in the Vortex columnar format. Cayenne delivers query and ingestion performance comparable or better to DuckDB's file-based acceleration without DuckDB's memory overhead and the scaling challenges of single DuckDB files.

Cayenne uses SQLite to manage acceleration metadata (schemas, snapshots, statistics, file tracking) through simple SQL transactions, while storing actual data in Vortex's compressed columnar format. This architecture provides:

Key Features:

SQLite + Vortex Architecture: All metadata is stored in SQLite tables with standard SQL transactions, while data lives in Vortex's compressed, chunked columnar format designed for zero-copy access and efficient scanning.
Simplified Operations: No complex file hierarchies, no JSON/Avro metadata files, no separate catalog servers—just SQL tables and Vortex data files. The entire metadata schema is intentionally simple for maximum reliability.
Fast Metadata Access: Single SQL query retrieves all metadata needed for query planning—no multiple round trips to storage, no S3 throttling, no reconstruction of metadata state from scattered files.
Efficient Small Changes: Dramatically reduces small file proliferation. Snapshots are just rows in SQLite tables, not new files on disk. Supports millions of snapshots without performance degradation.
High Concurrency: Changes consist of two steps: stage Vortex files (if any), then run a single SQL transaction. Much faster conflict resolution and support for many more concurrent updates than file-based formats.
Advanced Data Lifecycle: Full ACID transactions, delete support, and retention SQL execution on refresh commit.

Example Spicepod.yml configuration:

datasets:
  - from: s3:my_table
    name: accelerated_data
    acceleration:
      enabled: true
      engine: cayenne
      retention:
        sql: DELETE FROM accelerated_data WHERE created_at < NOW() - INTERVAL '30 days'

Note, the Cayenne Data Accelerator is in Alpha with limitations.

For more details, refer to the Cayenne Documentation, the Vortex project, and the DuckLake announcement that partly inspired this design.

DataFusion v50 Upgrade

Spice.ai is built on the DataFusion query engine. The v50 release brings significant performance improvements and enhanced reliability:

Performance Improvements 🚀:

Dynamic Filter Pushdown: Enhanced dynamic filter pushdown for custom ExecutionPlans, ensuring filters propagate correctly through all physical operators for improved query performance.
Partition Pruning: Expanded partition pruning support ensures that unnecessary partitions are skipped when filters are not used, reducing data scanning overhead and improving query execution times.

See the Apache DataFusion 50.0.0 Release for more details.

DynamoDB Data Connector Improvements

Example Spicepod.yml configuration:

datasets:
  - from: dynamodb:my_table
    name: ddb_data
    params:
      scan_segments: 10 # Default `auto` which calculates optimal segments based on number of rows

Search & Embeddings Enhancements

Multi-Column Embeddings on Views: Views now support embedding columns, enabling vector search and semantic retrieval on view data. This is useful for search over aggregated or joined datasets.

Vector Engines on Views: Vector search engines are now available for views, enabling similarity search over complex queries and transformations.

Example Spicepod.yml configuration:

views:
  - name: aggregated_reviews
    sql: SELECT review_id, review_text FROM reviews WHERE rating > 4
    embeddings:
      - column: review_text
        model: openai:text-embedding-3-small

DuckDB Accelerator Improvements

Parquet Buffering for Partitioned Writes: DuckDB partitioned writes in table mode now support Parquet buffering, reducing memory usage and improving write performance for large datasets.

Retention SQL on Refresh Commit: DuckDB accelerations now support running retention SQL on refresh commit, enabling automatic data cleanup and lifecycle management during refresh operations.

UTC Timezone for DuckDB: DuckDB now uses UTC as the default timezone, ensuring consistent behavior for time-based queries across different environments.

Example Spicepod.yml configuration:

datasets:
  - from: s3://my_bucket/large_table/
    name: partitioned_data
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      retention:
        sql: DELETE FROM partitioned_data WHERE event_time < NOW() - INTERVAL '7 days'

Query Performance Optimizations

Security & Reliability Improvements

ODBC Connector Improvements: Removed unwrap calls from the ODBC connector, improving error handling and reliability. Fixed secret handling and Kubernetes secret integration.

CLI Permissions Hardening: Tightened file permissions for the CLI and install script, ensuring secure defaults for configuration files and credentials.

Oracle Instant Client Pinning: Oracle Instant Client downloads are now pinned to specific SHAs, ensuring reproducible builds and preventing supply chain attacks.

Observability & Tracing

DataFusion Log Emission: The Spice runtime now emits DataFusion internal logs, providing deeper visibility into query planning and execution for debugging and performance analysis.

AI Completions Tracing: Fixed tracing so that ai_completions operations are correctly parented under sql_query traces, improving observability for AI-powered queries.

Git Data Connector (Alpha)

Example Spicepod.yml configuration:

datasets:
  - from: git:https://github.com/myorg/myrepo
    name: git_metrics
    params:
      file_format: csv

For more details, refer to the Git Data Connector Documentation.

Additional Improvements & Bug Fixes

Reliability: Fixed refresh worker panics with recovery handling to prevent runtime crashes during acceleration refreshes.
Reliability: Improved error messages for missing or invalid spicepod.yaml files, providing actionable feedback for misconfiguration.
Reliability: Fixed DuckDB metadata pointer loading issues for snapshots.
Performance: Ensured ListingTable partitions are pruned correctly when filters are not used.
Reliability: Fixed vector dimension determination for partitioned indexes.
Search: Fixed casing issues in Reciprocal Rank Fusion (RRF) for hybrid search queries.
Search: Fixed search field handling as metadata for chunked search indexes.
Validation: Added timestamp support for partition expressions.
Validation: Fixed regexp_match function for DuckDB datasets.
Validation: Fixed partition name validation for improved reliability.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

No major cookbook updates.

The Spice Cookbook includes 81 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.9.0-rc.1, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.9.0-rc.1 image:

docker pull spiceai/spiceai:1.9.0-rc.1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Changelog

Fix for search field as metadata for chunked search indexes by @Jeadie in #7429
Bump object_store from 0.12.3 to 0.12.4 by @app/dependabot in #7433
Properly respect disabling snapshots by @phillipleblanc in #7431
Revert "Properly respect disabling snapshots" by @sgrebnov in #7439
Revert "Disable snapshots by default" by @sgrebnov in #7438
Add preview warning for write access mode by @sgrebnov in #7440
fix: regexp_match for DuckDB datasets by @kczimm in #7443
Add feature is currently in preview warning for snapshots by @sgrebnov in #7442
[Logger] Also emit Datafusion logs by @mach-kernel in #7441
add missing snapshot by @kczimm in #7446
Fix tracing so that ai_completions are parented under sql_query by @lukekim in #7415
Enable snapshot acceleration by default by @phillipleblanc in #7451
Disable acceleration refresh metrics by @krinart in #7450
Add v1.8 release notes by @phillipleblanc in #7430
fix: partition name validation by @kczimm in #7452
Fix lint error due to ignore without reasons by @krinart in #7454
Add models and CUDA support to spiced install script by @lukekim in #7457
Post-release 1.8 updates by @phillipleblanc in #7455
Remove println in datafusion by @phillipleblanc in #7461
Update end_game.md to notify once release is done by @sgrebnov in #7460
Remove italics from snapshot logging by @phillipleblanc in #7463
Update openapi.json by @app/github-actions in #7466
Fix generate spicepod schema by @phillipleblanc in #7464
Fix generate acknowledements by @phillipleblanc in #7465
Update spicepod.schema.json by @app/github-actions in #7469
fix: Ensure ListingTable partitions are pruned when filters are not used by @peasee in #7471
Create runtime-secrets crate by @phillipleblanc in #7474
Create runtime-parameters crate by @phillipleblanc in #7475
Don't download the snapshot if the acceleration is present by @phillipleblanc in #7477
Fix casing for keywords and additional columns by @Jeadie in #7770
Bump actions/upload-artifact from 4 to 5 by @app/dependabot in #7750
Bump criterion from 0.5.1 to 0.7.0 by @app/dependabot in #7740
Bump rustls-native-certs from 0.8.1 to 0.8.2 by @app/dependabot in #7744
Git Data Connector (Alpha) by @lukekim in #7772
Pepper accelerator delete support by @lukekim in #7616
Update Helm chart instructions for Helm in end_game.md by @sgrebnov in #7776
Turso data accelerator by @lukekim in #7472
Apply retention SQL filter to refresh fetch by @phillipleblanc in #7778
Add Parquet buffering option for DuckDB partitioned writes (tables mode) by @sgrebnov in #7735
fix: EmptyExec when list indexes is empty by @kczimm in #7784
1.8.3 post-release housekeeping by @mach-kernel in #7783
feat: Upgrade to Datafusion v50 by @peasee in #7777
fix: Replace vortex datafusion with public crate by @peasee in #7791
Full-text search on views by @Jeadie in #7733
Revert "Apply retention SQL filter to refresh fetch (#7778)" by @phillipleblanc in #7796
fix: Add ingest duration and acceleration size metrics to testoperator by @peasee in #7792
Set local timezone to UTC for DuckDB by @phillipleblanc in #7797
add Timestamp support for partition expressions by @kczimm in #7803
Fix trunk lint by @krinart in #7804
Add missing mongodb params by @krinart in #7807
Embedding columns on view components by @Jeadie in #7795
Add Turso as a Pepper Catalog metastore by @lukekim in #7793
Run retention_sql on refresh commit for DuckDB by @lukekim in #7785
docs: Update datafusion upgrade checklist by @peasee in #7812
Vector engines on views by @Jeadie in #7808
Handle refresh worker panics and add recovery test by @phillipleblanc in #7815
chunk large record batches to control memory usage by @kczimm in #7802
fix: cannot determine vector dimension for partitioned indexes by @kczimm in #7818
Upgrade to Turso v0.3 by @lukekim in #7821
fix: Ensure custom *Exec ExecutionPlans push down dynamic filters by @peasee in #7811
handle casing in RRF by @Jeadie in #7825
Enable 'turso' for pepper acceleration by default by @sgrebnov in #7826
Improved DynamoDB Data Connector by @krinart in #7715
Initial support for llama.cpp as LLM inference backend by @lukekim in #7794
Pepper: Implement retention SQL on refresh commit by @phillipleblanc in #7814
Fix Dockerfiles for arm64 by @lukekim in #7834
[DynamoDB] Handle filter edge-cases by @krinart in #7830
[DynamoDB] Support parallelization for Scan request by @krinart in #7829
Don't feature gate Pepper by @lukekim in #7832
Fix llama.cpp static link by @lukekim in #7835
fix: docker nightly builds by @kczimm in #7837
Use GitHub-hosted macOS runner only for tag releases by @lukekim in #7836
Fix Bug: DuckDB INTERNAL Error: Failed to load metadata pointer by @sgrebnov in #7839
Fix docker arm64 build to use aegis in pure-rust mode by @lukekim in #7840
Revert "Use GitHub-hosted macOS runner only for tag releases" by @lukekim in #7843
Rename Pepper to Cayenne by @lukekim in #7844
Tighten CLI permissions and install script by @lukekim in #7845
Set mvcc for Cayenne Turso metastore by @lukekim in #7850
Optimize Prepared Statements by @lukekim in #7859
Remove unwrap from ODBC connector, fix secrets, and kuberenetes secre… by @lukekim in #7846
Improve and secure HTTP client usage by @lukekim in #7847
Pin Oracle Instant Client download to a SHA by @lukekim in #7851
Improve experience for missing or invalid Spicepod.yaml by @lukekim in #7849
chore: Fix PR linting by @peasee in #7865
Revert FlightIPC issues by @Jeadie in #7870
Improve error message by adding 'cayenne' to the list of valid accelerator engines by @sgrebnov in #7882
fix: allow parameter index without dollar signs by @kczimm in #7887
Temporary disable supports_limit_pushdown for SchemaCastScanExec by @sgrebnov in #7893
Remove '.embeddings[].metadata' by @Jeadie in #7897

Spice v1.8.3 (Oct 27, 2025)

October 27, 2025 · 5 min read

David Stancu

Principal Software Engineer at Spice AI

Announcing the release of Spice v1.8.3! ⚡

Spice v1.8.3 is a patch release focused on performance, reliability, and observability. This release delivers optimizations for DuckDB acceleration, parameterized queries, and query plans. A new opt-in dedicated thread pool for queries is now in preview.

What's New in v1.8.3

DuckDB Data Accelerator Improvements

Connection Pool Sizing: The DuckDB accelerator now supports a configurable connection_pool_size parameter, supporting fine-grained control over concurrent query execution. This enables tuning for high-concurrency workloads and improved resource utilization.

Example Spicepod.yaml snippet:

datasets:
  - from: postgres:my_table
    name: my_table
    acceleration:
      enabled: true
      engine: duckdb
      params:
        connection_pool_size: 10

Automatic Statistics Recomputation: The new on_refresh_recompute_statistics parameter, on by default, triggers automatic ANALYZE execution after refreshes. This keeps DuckDB optimizer statistics up-to-date, ensuring efficient query plans and optimal performance.

Example Spicepod.yaml snippet:

datasets:
  - from: postgres:my_table
    name: my_table
    acceleration:
      enabled: true
      engine: duckdb
      params:
        on_refresh_recompute_statistics: disabled # default enabled

Task History SQL Query Plan Capture & Configuration

Spice now supports automated SQL query plan capture and store (via EXPLAIN or EXPLAIN ANALYZE) in the task history, enabling deeper analysis and debugging of query execution. This feature is configurable, supporting control of which queries are included based on duration thresholds and plan type.

New Configuration Options:
- task_history.captured_plan: Controls which plan is captured (none, explain, or explain analyze). Default none.
- task_history.min_sql_duration: Minimum query duration before a plan is captured.
- task_history.min_plan_duration: Minimum plan execution duration before a plan is captured.

Example spicepod.yaml snippet:

runtime:
  task_history:
    captured_plan: explain analyze
    min_sql_duration: 5s
    min_plan_duration: 10s

Query plans are captured asynchronously to avoid blocking query execution. The result of the plan is stored in the standard sql_query output in the task history.

Learn more in the Task History Documentation.

Query Performance Optimizations

Optimized Prepared Statements (Parameterized Queries): Prepared statement caching for parameterized SQL queries has been improved, reducing planning overhead for repeated queries with different parameters. This results in faster execution and lower latency for workloads that reuse query structures.
Limit Pushdown via BytesProcessedExec: Introduces the BytesProcessedExec physical operator, enabling limit pushdown for large datasets. This optimization reduces the amount of data processed and improves top-k query performance.

Dedicated Query Thread Pool (Opt-In)

Spice now supports running query execution and accelerated refreshes on a dedicated thread pool, separate from the HTTP server. This prevents heavy query workloads from slowing down API responses, keeping health and readiness checks fast. Opt-In for v1.8.3: This feature is opt-in for this release and will become enabled by default (opt-out) in v1.9.

Example Spicepod.yaml snippet:

runtime:
  params:
    dedicated_thread_pool: sql_engine # Default: disabled

Validation & Reliability Improvements

Selective Evaluation Scorer Loading: Evaluation scorers are now loaded only when evaluation is explicitly defined, reducing unnecessary initialization and improving startup performance.
Improved Error Reporting: Enhanced error messages for misconfigured full-text search (FTS) on datasets and views, providing actionable feedback for configuration issues.

REPL & Usability

Execution Time Display: The Spice REPL now displays query execution time even when queries return no results, improving user feedback and diagnostics.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

No major cookbook updates.

The Spice Cookbook includes 81 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.8.3, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.8.3 image:

docker pull spiceai/spiceai:1.8.3

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Changelog

Fix generate spicepod schema by @phillipleblanc in #7464
Only load eval scorers when eval defined by @Jeadie in #7549
BytesProcessedExec to allow optimizer to do limit pushdown by @mach-kernel in #7539
Enhancement: Add spill_compression to runtime config by @krinart in #7505
Task History min_sql_duration filter support by @lukekim in #7698
Show error if FTS is misconfigured for datasets/views by @krinart in #7458
project_schema when using EmptyExec by @kczimm in #7543
Fix score order for one test case by @Jeadie in #7595
Fix license issue in table-providers by @phillipleblanc in #7620
Split integration tests into 3 partitions by @phillipleblanc in #7635
Fix OSS docker release trigger when release marked as latest by @phillipleblanc in #7668
Properly set auth headers in github_release.py by @krinart in #7560
Run Datafusion queries on a separate Tokio runtime by @phillipleblanc in #7586
Update BytesProcessedExec snapshots by @mach-kernel in #7637
Display execution time in Spice REPL for no results by @sgrebnov in #7713
Add support for DuckDB connection_pool_size param by @sgrebnov in #7716
Task History capture and store SQL query plans by @lukekim in #7701

Spice v1.8.2 (Oct 21, 2025)

October 21, 2025 · 5 min read

Jack Eadie

Token Plumber at Spice AI

Announcing the release of Spice v1.8.2! 🔍

Spice v1.8.2 is a patch release focused on reliability, validation, performance, and bug fixes, with improvements across DuckDB acceleration, S3 Vectors, document tables, and HTTP search.

What's New in v1.8.2

Support Table Relations in `/v1/search` HTTP Endpoint

Spice now supports table relations for the additional_columns and where parameters in the /v1/search endpoint. This enables improved search for multi-dataset use cases, where filters and columns can be used on specific datasets.

Example:

curl 'http://localhost:8090/v1/search' \
    -H 'Content-Type: application/json' \
    -H 'Accept: application/json' -d '{
        "text": "hello world",
        "additional_columns": ["tbl1.foo", "tbl2.bar", "baz"],
        "where": "tbl1.foo > 100000",
        "limit": 5
    }'

In this example, search results from the tbl1 dataset will include columns foo and baz, where foo > 100000. For tbl2, columns bar and baz will be returned.

DuckDB Data Accelerator Table Partitioning & Indexing

Configurable DuckDB Index Scan: DuckDB acceleration now supports configurable duckdb_index_scan_percentage and duckdb_index_scan_max_count parameters, supporting fine-tuning of index scan behavior for improved query performance.

Example:

datasets:
  - from: postgres:my_table
    name: my_table
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      params:
        # When combined, DuckDB will use an index scan when the number of qualifying rows is less than the maximum of these two thresholds
        duckdb_index_scan_percentage: '0.10' # 10% as decimal
        duckdb_index_scan_max_count: '1000'

Hive-Style Partitioning: In file-partitioned mode, the DuckDB data accelerator uses Hive-style partitioning for more efficient file management.
Table-Based Partitioning: Spice now supports partitioning DuckDB accelerations within a single file. This approach maintains ACID guarantees for full and append mode refreshes, while optimizing resource usage and improving query performance. Configure via the partition_mode parameter:

datasets:
  - from: file:test_data.parquet
    name: test_data
    params:
      file_format: parquet
    acceleration:
      enabled: true
      engine: duckdb
      mode: file
      params:
        partition_mode: tables
      partition_by:
        - bucket(100, Field1)

S3 Vectors Reliability

Race Condition Fix: Resolved a race condition in S3 Vectors index and bucket creation. The runtime also now checks if an index or bucket exists after a ConflictException, ensuring robust error handling during index creation and improving reliability for large-scale multi-index vector search.

Document Table Improvements

Primary Key Update: Document tables now use the location column as the primary key, improving performance, consistency, and query reliability.

Additional Improvements & Bugfixes

Reliability: Improved error handling and resource checks for S3 Vectors and DuckDB acceleration.
Validation: Expanded validation for partitioning and index creation.
Performance: Optimized partition refresh and index scan logic.
Bugfix: Don't nullify DuckDB release callbacks for schemas.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

No major cookbook updates.

The Spice Cookbook includes 81 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.8.2, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.8.2 image:

docker pull spiceai/spiceai:1.8.2

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Changelog

Update mongo config for benchmarks by @krinart in #7546
Configurable DuckDB duckdb_index_scan_percentage & duckdb_index_scan_max_count by @lukekim in #7551
Fix race condition in S3 Vectors index and bucket creation by @kczimm in #7577
Use 'location' as primary key for document tables by @Jeadie in #7567
Update official Docker builds to use release binaries by @phillipleblanc in #7597
Hive-style partitioning for DuckDB file mode by @kczimm in #7563
New Generate Changelog workflow by @krinart in #7562
Add support for DuckDB table-based partitioning by @sgrebnov in #7581
DuckDB table partitioning: delete partitions that no longer exist after full refresh by @sgrebnov in #7614
Rename duckdb_partition_mode to partition_mode param by @sgrebnov in #7622
Fix license issue in table-providers by @phillipleblanc in #7620
Make DuckDB table partition data write threshold configurable by @sgrebnov in #7626
fix: Don't nullify DuckDB release callbacks for schemas by @peasee in #7628
Fix integration tests by reverting the use of batch inserts w/ prepared statements by @phillipleblanc in #7630
Return TableProvider from CandidateGeneration::search by @Jeadie in #7559
Handle table relations in HTTP v1/search by @Jeadie in #7615

Spice v1.8.1 (Oct 13, 2025)

October 14, 2025 · 5 min read

Viktor Yershov

Senior Software Engineer at Spice AI

Announcing the release of Spice v1.8.1! 🚀

Spice v1.8.1 is a patch release that adds Acceleration Snapshots Indexes, and includes a number of bug fixes and performance improvements.

What's New in v1.8.1

Acceleration Snapshot Indexes

Management of Acceleration Snapshots has been improved by adopting an Iceberg-inspired metadata.json, which now encodes pointer IDs, schema serialization, and robust checksum and size, which is validate before loading the snapshot.
Acceleration Snapshot Metrics: The following metrics are now available for Acceleration Snapshots:
dataset_acceleration_snapshot_bootstrap_duration_ms: The time it took the runtime to download the snapshot - only emitted when it initially downloads the snapshot.
dataset_acceleration_snapshot_bootstrap_bytes: The number of bytes downloaded to bootstrap the acceleration from the snapshot.
dataset_acceleration_snapshot_bootstrap_checksum: The checksum of the snapshot used to bootstrap the acceleration.
dataset_acceleration_snapshot_failure_count: Number of failures encountered when writing a new snapshot at the end of the refresh cycle. A snapshot failure does not prevent the refresh from completing.
dataset_acceleration_snapshot_write_timestamp: Unix timestamp in seconds when the last snapshot was completed.
dataset_acceleration_snapshot_write_duration_ms: The time it took to write the snapshot to object storage.
dataset_acceleration_snapshot_write_bytes: The number of bytes written on the last snapshot write.
dataset_acceleration_snapshot_write_checksum: The SHA256 checksum of the last snapshot write.

To learn more, see the Acceleration Snapshots Documentation and the Metrics Documentation.

Improved Regular Expression for DuckDB acceleration

Regular expression support has been expanded when using DuckDB acceleration for functions like regexp-like and regexp_match.

For more details, refer to the SQL Reference for the list of available regular expression functions.

Additional Improvements & Bugfixes

Reliability: Resolved an issue with partitioning on empty partition sets.
Validation: Added better validation for incorrectly configured Spicepods.
Reliability: Fixed partition_by accelerations when a projection is applied on empty partition sets.
Performance: Ensured ListingTable partitions are pruned when filters are not used.
Performance: Don't download acceleration snapshots if the acceleration is already present.
Performance: Refactored some blocking I/O and synchronization in the async codebase by moving operations to tokio::task::spawn_blocking, replacing blocking locks with async-friendly variants.
Bugfix: Nullable fields are now supported for S3 Vectors index columns.

Contributors

Breaking Changes

No breaking changes.

Cookbook Updates

New Accelerated Snapshots Recipe - The recipe shows how to bootstrap DuckDB accelerations from object storage to skip cold starts.

The Spice Cookbook includes 81 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.8.1, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.8.1 image:

docker pull spiceai/spiceai:1.8.1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Changelog

Remove println in datafusion by @phillipleblanc in #7461
fix: Ensure ListingTable partitions are pruned when filters are not used by @peasee in #7471
Create runtime-secrets crate by @phillipleblanc in #7474
Create runtime-parameters crate by @phillipleblanc in #7475
Don't download the snapshot if the acceleration is present by @phillipleblanc in #7477
Add support for S3 dataset params by @phillipleblanc in #7476
Add better snapshot validation for incorrectly configured spicepods by @phillipleblanc in #7487
Move blocking/sync I/O to spawn blocking by @lukekim in #7462
Validate spicepod file exists before running tests by @lukekim in #7492
Make snapshot reading/writing more robust with Iceberg-like metadata.json by @phillipleblanc in #7486
Create runtime-request-context crate by @Jeadie in #7459
Two minor fixes for AI udf tests by @krinart in #7503
Add model response timeout for ai udf tests by @krinart in #7504
Add sccache for build test operator by @lukekim in #7515
Fix partition_by accelerations when a projection is applied on empty partition sets by @phillipleblanc in #7526
Nullable fields for index columns by @Jeadie in #7523

Spice v1.8.0 (Oct 6, 2025)

October 7, 2025 · 20 min read

Phillip LeBlanc

Co-Founder and CTO of Spice AI

Announcing the release of Spice v1.8.0! 🧊

Spice v1.8.0 delivers major advances in data writes, scalable vector search, and now in preview—managed acceleration snapshots for fast cold starts. This release introduces write support for Iceberg tables using standard SQL INSERT INTO, partitioned S3 Vector indexes for petabyte-scale vector search, and preview of the AI SQL function for direct LLM integration in SQL. Additional improvements include improved reliability, and the v3.0.3 release of the Spice.js Node.js SDK.

What's New in v1.8.0

Iceberg Table Write Support (Preview)

Append Data to Iceberg Tables with SQL INSERT INTO: Spice now supports writing to Iceberg tables and catalogs using standard SQL INSERT INTO statements. This enables data ingestion, transformation, and pipeline use cases—no Spark or external writer required.

Append-only: Initial version targets appends; no overwrite or delete.
Schema validation: Inserted data must match the target table schema.
Secure by default: Writes are only enabled for datasets or catalogs explicitly marked with access: read_write.

Example Spicepod configuration:

catalogs:
  - from: iceberg:https://glue.ap-northeast-3.amazonaws.com/iceberg/v1/catalogs/111111/namespaces
    name: ice
    access: read_write

datasets:
  - from: iceberg:https://iceberg-catalog-host.com/v1/namespaces/my_namespace/tables/my_table
    name: iceberg_table
    access: read_write

Example SQL usage:

-- Insert from another table
INSERT INTO iceberg_table
SELECT * FROM existing_table;

-- Insert with values
INSERT INTO iceberg_table (id, name, amount)
VALUES (1, 'John', 100.0), (2, 'Jane', 200.0);

-- Insert into catalog table
INSERT INTO ice.sales.transactions
VALUES (1001, '2025-01-15', 299.99, 'completed');

Note: Only Iceberg datasets and catalogs with access: read_write support writes. Internal Spice tables and other connectors remain read-only.

Learn more in the Iceberg Data Connector documentation.

Acceleration Snapshots for Fast Cold Starts (Preview)

Bootstrap Managed Accelerations from Object Storage: Spice now supports managed acceleration snapshots in preview, enabling datasets accelerated with file-based engines (DuckDB or SQLite) to bootstrap from a snapshot stored in object storage (such as S3) if the local acceleration file does not exist on startup. This dramatically reduces cold start times and enables ephemeral storage for accelerations with persistent recovery.

Key features:

Rapid readiness: Datasets can become ready in seconds by downloading a pre-built snapshot, skipping lengthy initial acceleration.
Hive-style partitioning: Snapshots are organized by month, day, and dataset for easy retention and management.
Flexible bootstrapping: Configurable fallback and retry behavior if a snapshot is missing or corrupted.

Example Spicepod configuration:

snapshots:
  enabled: true
  location: s3://some_bucket/some_folder/ # Folder for storing snapshots
  bootstrap_on_failure_behavior: warn # Options: warn, retry, fallback
  params:
    s3_auth: iam_role # All S3 dataset params accepted here

datasets:
  - from: s3://some_bucket/some_table/
    name: some_table
    params:
      file_format: parquet
      s3_auth: iam_role
    acceleration:
      enabled: true
      snapshots: enabled # Options: enabled, disabled, bootstrap_only, create_only
      engine: duckdb
      mode: file
      params:
        duckdb_file: /nvme/some_table.db

How it works:

On startup, if the acceleration file does not exist, Spice checks the snapshot location for the latest snapshot and downloads it.
Snapshots are stored as: s3://some_bucket/some_folder/month=2025-09/day=2025-09-30/dataset=some_table/some_table_<timestamp>.db
If no snapshot is found, a new acceleration file is created as usual.
Snapshots are written after each refresh (unless configured otherwise).

Supported snapshot modes:

enabled: Download and write snapshots.
bootstrap_only: Only download on startup, do not write new snapshots.
create_only: Only write snapshots, do not download on startup.
disabled: No snapshotting.

Note: This feature is only supported for file-based accelerations (DuckDB or SQLite) with dedicated files.

Why use acceleration snapshots?

Faster cold starts: Skip waiting for full acceleration on startup.
Ephemeral storage: Use fast local disks (e.g., NVMe) for acceleration, with persistent recovery from object storage.
Disaster recovery: Recover from federated source outages by bootstrapping from the latest snapshot.

Partitioned S3 Vector Indexes

Efficient, Scalable Vector Search with Partitioning: Spice now supports partitioning Amazon S3 Vector indexes and scatter-gather queries using a partition_by expression in the dataset vector engine configuration. Partitioned indexes enable faster ingestion, lower query latency, and scale to billions of vectors.

Example Spicepod configuration:

datasets:
  - name: reviews
    vectors:
      enabled: true
      engine: s3_vectors
      params:
        s3_vectors_bucket: my-bucket
        s3_vectors_index: base-embeddings
      partition_by:
        - 'bucket(50, PULocationID)'
    columns:
      - name: body
        embeddings:
          from: bedrock_titan
      - name: title
        embeddings:
          from: bedrock_titan

See the Amazon S3 Vectors documentation for details.

AI SQL function for LLM Integration (Preview)

LLMs Directly In SQL: A new asynchronous ai SQL function enables direct calls to LLMs from SQL queries for text generation, translation, classification, and more. This feature is released in preview and supports both default and model-specific invocation.

Example Spicepod model configuration:

models:
  - name: gpt-4o
    from: openai:gpt-4o
    params:
      openai_api_key: ${secrets:openai_key}

Example SQL usage:

-- basic usage with default model
SELECT ai('hi, this prompt is directly from SQL.');

-- basic usage with specified model
SELECT ai('hi, this prompt is directly from SQL.', 'gpt-4o');

-- Using row data as input to the prompt
SELECT ai(concat_ws(' ', 'Categorize the zone', Zone, 'in a single word. Only return the word.')) AS category
FROM taxi_zones
LIMIT 10;

Learn more in the SQL Reference AI documentation.

Remote Endpoint Support for Spice CLI

Run CLI Commands Remotely: The Spice CLI now supports connecting to remote Spice instances, enabling you to run spice sql, spice search, and spice chat commands from your local machine against a remote spiced daemon or to Spice Cloud. Previously, these commands required running on the same machine as the runtime. Now, new flags allow remote execution:

--cloud: Connect to a Spice Cloud instance (requires --api-key).
--endpoint <endpoint>: Connect to a remote Spice instance via HTTP or Arrow Flight SQL (gRPC). Supports http://, https://, grpc://, or grpc+tls:// schemes.

Examples:

# Run SQL queries against a remote Spice instance
spice sql --endpoint http://remote-host:8090

# Use Spice Cloud for chat or search
spice chat --cloud --api-key <your-api-key>
spice search --cloud --api-key <your-api-key>

Supported CLI Commands:

spice sql --cloud / spice sql --endpoint <endpoint>
spice search --cloud / spice search --endpoint <endpoint>
spice chat --cloud / spice chat --endpoint <endpoint>

Additional Flags:

--headers: Pass custom HTTP headers to the remote endpoint.
--tls-root-certificate-file: Specify a root certificate for TLS verification.
--user-agent: Set a custom user agent for requests.

For more details, see the Spice CLI Command Reference.

Spice.js v3.0.3 SDK

Spice.js v3.0.3 Released: The official Spice.ai Node.js/JavaScript SDK has been updated to v3.0.3, bringing cross-platform support, new APIs, and improved reliability for both Node.js and browser environments.

Modern Query Methods: Use sql(), sqlJson(), and nsql() for flexible querying, streaming, and natural language to SQL.
Browser Support: SDK now works in browsers and web applications, automatically selecting the optimal transport (gRPC or HTTP).
Health Checks & Dataset Refresh: Easily monitor Spice runtime health and trigger dataset refreshes on demand.
Automatic HTTP Fallback: If gRPC/Flight is unavailable, the SDK falls back to HTTP automatically.
Migration Guidance: v3 requires Node.js 20+, uses camelCase parameters, and introduces a new package structure.

Example usage:

import { SpiceClient } from '@spiceai/spice'

const client = new SpiceClient(apiKey)
const table = await client.sql('SELECT * FROM my_table LIMIT 10')
console.table(table.toArray())

See Spice.js SDK documentation for full details, migration tips, and advanced usage.

Additional Improvements

Reliability: Improved logging, error handling, and network readiness checks across connectors (Iceberg, Databricks, etc.).
Vector search durability and scale: Refined logging, stricter default limits, safeguards against index-only scans and duplicate results, and always-accessible metadata for robust queryability at scale.
Cache behavior: Tightened cache logic for modification queries.
Full-Text Search: FTS metadata columns now usable in projections; max search results increased to 1000.
RRF Hybrid Search: Reciprocal Rank Fusion (RRF) UDTF enhancements for advanced hybrid search scenarios.

Contributors

Breaking Changes

This release introduces two breaking changes associated with the search observability and tooling.

Firstly, the document_similarity tool has been renamed to search. This has the equivalent change to tracing of these tool calls:

## Old: v1.7.1
>> spice trace tool_use::document_similarity
>> curl -XPOST http://localhost:8090/v1/tools/document_similarity \
  -d '{
    "datasets": ["my_tbl"],
    "text": "Welcome to another Spice release"
  }'

## New: v1.8.0
>> spice trace tool_use::search
>> curl -XPOST http://localhost:8090/v1/tools/search \
  -d '{
    "datasets": ["my_tbl"],
    "text": "Welcome to another Spice release"
  }'

Secondly, the vector_search task in runtime.task_history has been renamed to search.

Cookbook Updates

Added new AI SQL function recipe for invoking LLMs within SQL queries.
Updated Iceberg Catalog Connector recipe for Iceberg Writes.
Updated Spice.js JavaScript (Node.js) SDK for v3.0.3 with examples and v2 to v3 migration guide.

The Spice Cookbook now includes 80 recipes to help you get started with Spice quickly and easily.

Upgrading

To upgrade to v1.8.0, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:1.8.0 image:

docker pull spiceai/spiceai:1.8.0

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai

AWS Marketplace:

🎉 Spice is now available in the AWS Marketplace!

What's Changed

Dependencies

iceberg-rust: Upgraded to v0.7.0-rc.1
mimalloc: Upgraded from 0.1.47 to 0.1.48
azure_core: Upgraded from 0.27.0 to 0.28.0
Jimver/cuda-toolkit: Upgraded from 0.2.27 to 0.2.28

Changelog

Add #[cfg(feature = "postgres")] to acceleration refresh tests by @Jeadie in #7241
fix: Update benchmark snapshots by @github-actions[bot] in #7267
fix: Update benchmark snapshots by @github-actions[bot] in #7268
fix: Update benchmark snapshots by @github-actions[bot] in #7269
Update the tpch benchmark snapshots for: federated/databricks[sql_warehouse].yaml by @github-actions[bot] in #7270
EmbeddingInput cache keys to include model name by @mach-kernel in #7275
ensure FTS metadata columns can be used in projection by @Jeadie in #7282
Use 8-core runners for Windows CUDA builds by @sgrebnov in #7284
Make search test more robust by @krinart in #7283
Post-release housekeeping by @sgrebnov in #7272
fix: Use median cached response duration for test search cache by @peasee in #7286
Bump dirs from 5.0.1 to 6.0.0 by @dependabot[bot] in #7244
Bump indexmap from 2.11.0 to 2.11.4 by @dependabot[bot] in #7248
Fix JOIN level filters not having columns in schema by @Jeadie in #7287
use SessionContext::new_empty in RRF by @kczimm in #7291
Use rust:1.89-slim-bookworm for build, more places to bump rust version by @sgrebnov in #7293
Update openapi.json by @github-actions[bot] in #7290
Enable chunking in SearchIndex by @Jeadie in #7143
Add index name and remove duplicate records string to S3 Vectors log by @lukekim in #7260
Use file-based fts index by @Jeadie in #7024
Remove 'PostApplyCandidateGeneration' by @Jeadie in #7288
RRF: Rank and recency boosting by @mach-kernel in #7294
Update ROADMAP.md by removing v1.7 milestone by @sgrebnov in #7297
RRF: Preserve base ranking when results differ -> FULL OUTER JOIN does not produce time column by @mach-kernel in #7300
chore: remove unused Dataset methods by @kczimm in #7295
fix removing embedding column by @Jeadie in #7302
fix: Add feature flag for using object store in spicepod by @peasee in #7303
Upgrade to iceberg-rust v0.7.0-rc1 by @sgrebnov in #7296
Enable DML Update SQL operations for datasets configured as access: read_write by @sgrebnov in #7304
Create and parse partitioned S3 vector index names by @kczimm in #7198
RRF: Fix decay for disjoint result sets by @mach-kernel in #7305
RRF: Project top scores, do not yield duplicate results by @mach-kernel in #7306
RRF: Case sensitive column/ident handling by @mach-kernel in #7309 in #7309
For vector_search, use a default limit of 1000 if no limit specified by @lukekim in #7311
Don’t cache modification queries (DDL, DML, COPY) by @sgrebnov in #7316
Fix Anthropic model regex and add validation tests by @ewgenius in #7319
Enhancement: Implement before/after/lag metrics for acceleration refresh by @krinart in #7310
Refactor chat model health check to lower tokens usage for reasoning models by @ewgenius in #7317
Add support for writing into Iceberg tables by @sgrebnov in #7315
Fix lint warnings by @lukekim in #7327
Use logical plan in SearchQueryProvider by @Jeadie in #7314
FTS max search results 100 -> 1000 by @Jeadie in #7331
Improve Databricks SQL Warehouse Error Handling by @sgrebnov in #7332
Use spicepod embedding model name for model_name() by @Jeadie in #7333
Handle async queries for Databricks SQL Warehouse API by @phillipleblanc in #7335
Enable DML (INSERT INTO) operations for catalogs configured as access:read_write by @sgrebnov in #7330
Bump regex from 1.11.2 to 1.11.3 by @dependabot[bot] in #7336
Update qa_analytics.csv with 1.7.0 release data by @sgrebnov in #7337
RRF: Fix ident resolution for struct fields, autohashed join key for varying types by @mach-kernel in #7339
v1.7.1 release notes by @kczimm in #7348
Bump Jimver/cuda-toolkit from 0.2.27 to 0.2.28 by @dependabot[bot] in #7343
Add support for writing into Glue (Iceberg) tables and catalogs by @sgrebnov in #7355
Bump mimalloc from 0.1.47 to 0.1.48 by @dependabot[bot] in #7342
Add ai async UDF by @lukekim in #7328
Use self-hosted and spiceai-macos runners for workflows where possible by @lukekim in #7371
Several updates for improved search testing by @Jeadie in #7358
Update supported versions in SECURITY.md by @Jeadie in #7377
1.7.1 release analytics by @mach-kernel in #7380
Add acceleration_file_path helper and refactor spice_sys to use Snafu errors by @phillipleblanc in #7376 in #7376
fix: Update benchmark snapshots by @github-actions[bot] in #7353
Robust search test by @Jeadie in #7381
[bug] Fix ai UDF bug of mismatched column length by @lukekim in #7383
Add OpenOption to spice_sys acceleration tables by @phillipleblanc in #7379
Add new snapshots Spicepod configuration by @phillipleblanc in #7384
Update naming of tool_use::document_similarity and vector_search spans by @Jeadie in #7273
fix: Update benchmark snapshots by @github-actions[bot] in #7354
Make ai UDF a models only feature by @lukekim in #7387
Add new runtime_acceleration crate; create SnapshotManager; implement SnapshotManager::download_latest_snapshot by @phillipleblanc in #7386
Refactor 'VectorScanTableProvider' to use just 'VectorIndex::list_table_provider' by @Jeadie in #7318
Fix embed logs by @Jeadie in #7382
Enable spicepod dependencies in testoperator by @Jeadie in #7334
ai UDF security and performance optimizations by @lukekim in #7392
Wire up the snapshot download on dataset startup by @phillipleblanc in #7389
Implement initial snapshot creation logic in SnapshotManager by @phillipleblanc in #7391
Make tool_use::table_schema output model-friendly by @krinart in #7393
Fix minor lint warnings by @lukekim in #7395
Enable metadata columns in document-based object store datasets by @Jeadie in #7397
Core dependencies of financebench by @Jeadie in #7400
Add S3vector variant to financebench by @Jeadie in #7399
Set PostgreSQL unsupported_spice_action=string by default by @lukekim in #7398
Use non-blocking connection check for verify_ns_lookup_and_tcp_connect by @phillipleblanc in #7401
Bump moka from 0.12.10 to 0.12.11 by @dependabot[bot] in #7340
Bump tokio-postgres from 0.7.13 to 0.7.14 by @dependabot[bot] in #7344
Bump azure_core from 0.27.0 to 0.28.0 by @dependabot[bot] in #7338
Forbid INSERT OVERWRITE DML operations by @sgrebnov in #7402
Make database connection pool sizes consistent by @lukekim in #7403
Disable vector index only scans by @Jeadie in #7405
Make CLI --endpoint and --cloud args & table output consistent by @lukekim in #7396
Write new snapshots at the end of an accelerated refresh by @phillipleblanc in #7410
Read and write partitioned S3 indexes by @kczimm in #7313
Fix partial data writes in Iceberg data connector by @sgrebnov in #7411
Remove nix by @phillipleblanc in #7414
Use DataFusion JoinSetTracer for async context propagation by @lukekim in #7416
Implement cache invalidation for DML (INSERT INTO) operations by @sgrebnov in #7394
Make cleanup disk GH action; use in integration tests by @Jeadie in #7418
Move S3Vector to 'search' crate by @Jeadie in #7373
Use LogicalPlan builder API for LogicalPlans by @Jeadie in #7408
Use hive-style partitioned paths for DB snapshots by @phillipleblanc in #7422
Limit results from SearchIndex::query_table_provider by @Jeadie in #7421
Delay initial readiness if snapshots are enabled with an append-mode refresh by @phillipleblanc in #7425
Disable snapshots by default by @phillipleblanc in #7426
Rewrite ChunkedNonIndexVectorGeneration to use LogicalPlanBuilder (instead of string formatting) by @Jeadie in #7413
Fix for search field as metadata for chunked search indexes by @Jeadie in #7429
Add feature is currently in preview warning for read_write access mode by @sgrebnov in #7440
Add feature is currently in preview warning for snapshots by @sgrebnov in #7442
Fix tracing so that ai_completions are parented under sql_query by @lukekim in #7415
Disable acceleration refresh metrics by @krinart in #7450
Enable snapshot acceleration by default by @phillipleblanc in #7451
fix: partition name validation by @kczimm in #7452

What's New in v1.10.0-rc1​

Caching Acceleration Mode with SWR and TinyLFU​

DynamoDB Streams Data Connector in Preview​

Cayenne Accelerator Enhancements​

S3 Connector Improvements​

Faster Distributed Query Execution​

Search Improvements​

Security Hardening​

Developer Experience Improvements​

Contributors​

Breaking Changes​

Cookbook Updates​

Upgrading​

What's Changed​

Changelog​

Amazon Bedrock Nova 2 Multimodal embeddings​

DynamoDB Timestamp Filter Pushdown​

HTTP Data Connector Health Probe Configuration​

Spice .NET SDK v0.2​

Additional Improvements & Bug Fixes​

Contributors​

Breaking Changes​

Cookbook Updates​

Upgrading​

What's Changed​

Changelog​

What's New in v1.9.0​

Cayenne Data Accelerator (Beta)​

Multi-Node Distributed Query (Preview)​

DataFusion v50 Upgrade​

DuckDB v1.4.2 Upgrade and Accelerator Improvements​

HTTP Data Connector​

DynamoDB Data Connector Improvements​

S3 Data Connector Improvements​

Search & Embeddings Enhancements​

Dedicated Query Thread Pool (Now Enabled by Default)​

Query Performance Optimizations​

Query Result Caching: Compressed Encoding, Stale-While-Revalidate Cache Control​

Security & Reliability Improvements​

AWS Authentication Improvements​

Observability & Tracing​

Git Data Connector (Alpha)​

Spice Java SDK 0.4.0​

CLI Improvements​

Additional Improvements & Bug Fixes​

Contributors​

Breaking Changes​

Cookbook Updates​

Upgrading​

What's Changed​

Dependencies​

Changelog​

What's New in v1.9.0​

Cayenne Data Accelerator (Beta)​

Multi-Node Distributed Query (Preview)​

DataFusion v50 Upgrade​

DuckDB v1.4.2 Upgrade and Accelerator Improvements​

HTTP Data Connector​

DynamoDB Data Connector Improvements​

S3 Versioning Support​

Search & Embeddings Enhancements​

Dedicated Query Thread Pool (Now Enabled by Default)​

Query Performance Optimizations​

Query Result Cache: Stale-While-Revalidate​

Security & Reliability Improvements​

AWS Authentication Improvements​

Observability & Tracing​

Git Data Connector (Alpha)​

Spice Java SDK 0.4.0​

CLI Improvements​

Additional Improvements & Bug Fixes​

Contributors​

Breaking Changes​

Cookbook Updates​

Upgrading​

What's Changed​

Dependencies​

Changelog (rc.4)​

What's New in v1.9.0-rc.2​

Cayenne Data Accelerator (Beta)​

What's New in v1.10.0-rc1

Caching Acceleration Mode with SWR and TinyLFU

DynamoDB Streams Data Connector in Preview

Cayenne Accelerator Enhancements

S3 Connector Improvements

Faster Distributed Query Execution

Search Improvements

Security Hardening

Developer Experience Improvements

Contributors

Breaking Changes

Cookbook Updates

Upgrading

What's Changed

Changelog

Amazon Bedrock Nova 2 Multimodal embeddings

DynamoDB Timestamp Filter Pushdown

HTTP Data Connector Health Probe Configuration

Spice .NET SDK v0.2

Additional Improvements & Bug Fixes

Contributors

Breaking Changes

Cookbook Updates

Upgrading

What's Changed

Changelog

What's New in v1.9.0

Cayenne Data Accelerator (Beta)

Multi-Node Distributed Query (Preview)

DataFusion v50 Upgrade

DuckDB v1.4.2 Upgrade and Accelerator Improvements

HTTP Data Connector

DynamoDB Data Connector Improvements

S3 Data Connector Improvements

Search & Embeddings Enhancements

Dedicated Query Thread Pool (Now Enabled by Default)

Query Performance Optimizations

Query Result Caching: Compressed Encoding, Stale-While-Revalidate Cache Control

Security & Reliability Improvements

AWS Authentication Improvements

Observability & Tracing

Git Data Connector (Alpha)

Spice Java SDK 0.4.0

CLI Improvements

Additional Improvements & Bug Fixes

Contributors

Breaking Changes

Cookbook Updates

Upgrading

What's Changed

Dependencies

Changelog

What's New in v1.9.0

Cayenne Data Accelerator (Beta)

Multi-Node Distributed Query (Preview)

DataFusion v50 Upgrade

DuckDB v1.4.2 Upgrade and Accelerator Improvements

HTTP Data Connector

DynamoDB Data Connector Improvements

S3 Versioning Support

Search & Embeddings Enhancements

Dedicated Query Thread Pool (Now Enabled by Default)

Query Performance Optimizations

Query Result Cache: Stale-While-Revalidate

Security & Reliability Improvements

AWS Authentication Improvements

Observability & Tracing

Git Data Connector (Alpha)

Spice Java SDK 0.4.0

CLI Improvements

Additional Improvements & Bug Fixes

Contributors

Breaking Changes

Cookbook Updates

Upgrading

What's Changed

Dependencies

Changelog (rc.4)

What's New in v1.9.0-rc.2

Cayenne Data Accelerator (Beta)