Snapshots
Acceleration Snapshots are available in the Spice Enterprise edition.
Spicepod Example​
snapshots:
enabled: true
location: s3://some_bucket/some_folder/
bootstrap_on_failure_behavior: warn
params:
s3_auth: iam_role
datasets:
- name: some_table
acceleration:
engine: duckdb
mode: file
snapshots: enabled
snapshots_trigger: refresh_complete
params:
duckdb_file: /nvme/some_table.db
Overview​
Acceleration snapshots let Spice reuse a pre-built acceleration file on startup instead of waiting for a full refresh. When a dataset uses a file-mode acceleration engine (DuckDB, SQLite, Cayenne, or Turso) and the local file is missing (for example on first boot or when using ephemeral NVMe storage), Spice downloads the most recent snapshot from object storage and moves the dataset straight to a ready state.
How it works​
- On startup, Spice checks whether the file supplied in
acceleration.params(for exampleduckdb_file) exists. - If the file is missing and snapshots are enabled, Spice looks under the configured snapshot location and downloads the newest snapshot for that dataset. The download runs as consecutive 8 MiB range requests, four at a time, each with its own retries, so a slow connection does not restart the whole download. Each download holds at most 32 MiB in memory.
- If no snapshot is available, the acceleration boots empty and refreshes from the source. A
refresh_mode: snapshotreader instead waits for the first snapshot. - Spice creates new snapshots based on the configured
snapshots_triggermode. A snapshot that fails for a transient reason, such as a network or object store error, is retried up to three times with backoff. A permanent failure, such as a schema mismatch, an unreadablemetadata.json, or a store that does not enforce conditional writes, is not retried.
Snapshots are organized with Hive-style partitioning so they are easy to retain and prune. For a dataset named my_dataset accelerated with DuckDB, Spice writes files such as:
s3://some_bucket/some_folder/month=2025-09/day=2025-09-30/dataset=my_dataset/my_dataset_20250919T134522Z.duckdb
The timestamp is recorded in UTC using ISO 8601 without punctuation. The file extension names the acceleration engine that wrote the snapshot: .duckdb, .sqlite, .cayenne, or .turso. The engine is also recorded in the snapshot metadata, and a snapshot whose engine differs from the dataset's current engine is rejected at bootstrap rather than restored.
Every accelerated dataset must write to its own file (for example, /nvme/my_dataset.db). Sharing a single file across multiple datasets is not supported.
What each snapshot contains​
Each snapshot is a complete copy of that dataset's acceleration file at the time Spice writes it. The object under the snapshot location is the whole accelerated dataset — DuckDB, SQLite, Cayenne, or Turso — ready for a reader to download and open. Spice writes that full file on every snapshot. It does not publish an incremental delta of the rows or objects that changed since the previous snapshot.
A bootstrap, and a refresh_mode: snapshot reload, replaces the local acceleration file with that copy. Every upload moves the full file, and bucket or object replication of the snapshot prefix moves the full file again. Storage grows with the size of the accelerated dataset times how often snapshots are written, until a lifecycle rule expires older objects. DuckDB snapshots_compaction can shrink each copy. The uploaded object is still a full file.
When the workload needs incremental replication of changed data between regions or storage tiers, use a path that publishes the changed objects:
- Cayenne cold / datalake tier.
cayenne_datalake_locationgraduates data onto standard object storage (a general-purpose S3 or S3-compatible bucket). Promotion carries unchanged cold files forward and rewrites the cold files a change can touch. The tier requiresrefresh_mode: changesorappendand a primary key. It is a data tier, separate fromsnapshots.location. - Iceberg batched writes with merge-on-read. Iceberg writes append new data files, and deletes are equality delete files that readers merge on scan. A catalog commit publishes those objects. Use this path when the dataset lives in Iceberg and readers apply delete files during scan.
Neither path bootstraps a file-mode accelerator. Keep snapshots for that.
Configure snapshot storage​
Snapshots are controlled with a top-level snapshots block in the Spicepod. The location can point to S3, Azure ADLS Gen2, Google Cloud Storage, or the local filesystem.
snapshots:
enabled: true
location: s3://some_bucket/some_folder/ # Folder where snapshots are written
bootstrap_on_failure_behavior: warn # retry | fallback | warn
params:
s3_auth: iam_role # Defaults to iam_role for snapshots
Supported storage backends​
| Backend | URL scheme | Environment variables |
|---|---|---|
| Amazon S3 | s3:// | Standard AWS credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, etc.) |
| Azure ADLS Gen2 | abfss://, abfs:// | AZURE_STORAGE_ACCOUNT_NAME, AZURE_STORAGE_ACCOUNT_KEY, AZURE_CLIENT_ID/AZURE_TENANT_ID/AZURE_FEDERATED_TOKEN_FILE |
| Google Cloud Storage | gs:// | GOOGLE_APPLICATION_CREDENTIALS, Workload Identity |
| Local filesystem | file:// | N/A |
location must be a URI with a scheme — a bare filesystem path such as /nvme/snapshots is not a valid URI, so it fails to parse and snapshots are disabled with an error logged. Use file:///nvme/snapshots/ for a local folder.
When the location is an S3 bucket, params accepts these S3 parameters: s3_region, s3_endpoint, s3_auth (iam_role or key; default iam_role), s3_key, s3_secret, s3_session_token, s3_queue_url (an SQS queue that receives the location's S3 event notifications, so refresh_mode: snapshot readers reload on publication), client_timeout, and allow_http. Spice ignores other keys and logs a warning for each. With s3_auth: key, both s3_key and s3_secret must be set; without them, Spice logs an error and does not use the snapshot location, rather than connecting with credentials from the environment. Azure and GCS locations also accept their respective connector parameters under params for explicit credential overrides. When no explicit credentials are supplied, Spice reads standard environment variables for each cloud provider. Values in params can reference secrets with ${secrets:<name>} for S3, Azure, and GCS locations.
An invalid S3 parameter value, such as s3_auth: public, disables snapshots for the dataset. Spice logs an error that names the dataset, the location, and the parameter, and does not fall back to default parameters. Those defaults would drop every other configured parameter, including s3_endpoint, s3_region, and the credentials, so snapshots could be written to a different store or with a different identity than the one configured. When s3_queue_url is set, the dataset fails to register with the parameter error instead. Fix the parameter and restart Spice.
Conditional writes on the snapshot location​
Each location holds one metadata.json document that records the snapshots of every dataset publishing there, from every Spice instance. Spice changes this document only with conditional writes. It creates the document only if it does not exist yet (If-None-Match: *), and updates it only if it still has the version Spice read (If-Match on its ETag). Google Cloud Storage expresses both conditions with the object generation (x-goog-if-generation-match). When another writer changes the document first, the store refuses the write, and Spice reads the document again and reapplies its change. Concurrent publishers therefore do not overwrite each other's entries.
The location's store must enforce both conditions. Amazon S3, Azure ADLS Gen2, Google Cloud Storage, and a local file:// directory do. An S3-compatible store must honor If-None-Match and If-Match on PUT; Spice checks this before it publishes, as described below.
Before each Spice instance creates a dataset's first snapshot, it probes the store by creating, updating, and deleting one .spice-conditional-write-probe-<id> object under the location, in at most five requests. When the store rejects conditional writes, or accepts a write whose condition does not hold, Spice uploads no snapshot and does not write metadata.json. The snapshot fails with an error that names the dataset, the location, and the reason. Spice keeps this result and does not probe the store again, so restart Spice after moving location to a store that enforces conditional writes.
When the probe cannot reach a result, for example after a transient error or because the credentials cannot write the probe object, the snapshot is not published and the next attempt probes again. The error asks for permission to create, update, and delete .spice-conditional-write-probe-* objects under the location. Every instance that publishes snapshots probes with its own credentials, and each needs permission to create and update these objects. Permission to delete them is needed only for cleanup: a failed delete does not change the probe result and leaves one small probe object behind.
Spice does not fall back to an unconditional overwrite of metadata.json. If the store returns no ETag or version for the document, the publish fails after its snapshot file is uploaded, and the file is not recorded.
When other writers keep changing metadata.json, Spice makes up to 10 attempts to update it, waiting a randomized interval before each retry. The first wait is 12 to 25 ms, and the upper bound doubles with each retry up to 2 seconds. If every attempt loses to another writer, the snapshot file stays uploaded but is not recorded, and the error names the dataset; the dataset's next snapshot tries again. An update that reached the store but whose response was lost is recognized on the next read and recorded once.
POST /v1/datasets/{name}/acceleration/snapshots/current probes the store on every request before it changes the current snapshot. When the store fails the probe, or the probe cannot reach a result, the request returns 500 Internal Server Error with the reason and metadata.json is not changed. Selecting the snapshot that is already current leaves metadata.json unwritten. The probe still runs first, so the request still creates, updates, and tries to delete a probe object.
Restoring a snapshot at startup, and reloading one with refresh_mode: snapshot or file_format: snapshot, only reads metadata.json and does not probe the store.
In Spice v2.3.2 and earlier, a file:// location records only its first publish. Every later publish, from any dataset, uploads its snapshot file and then fails to update metadata.json with the error Operation `put_opts` with mode `PutMode::Update` not yet implemented by LocalFileSystem.
Several instances creating one dataset's snapshots​
When more than one Spice instance creates snapshots of the same dataset at the same location, only one of them uploads. The instances race for a lease object, leases/<dataset>.json under the location, with the same conditional writes as metadata.json. The holder renews the lease each time it is about to create a snapshot. The other instances skip their uploads and log a message such as:
Dataset 'orders' is not creating snapshots while instance 'spice-1' holds its snapshot writer lease; this instance takes over if that lease goes 2m without renewal. See: https://spiceai.org/docs/features/data-acceleration/snapshots
A lease lasts twice the dataset's snapshot interval, kept between 30 seconds and 24 hours. The interval is snapshots_trigger_threshold with snapshots_trigger: time_interval and refresh_check_interval with refresh_complete, and otherwise 10 minutes. Another instance takes over a lease that has gone unrenewed that long, as measured on its own clock, so clock skew between instances does not expire a live lease. A holder whose snapshot fails releases the lease at once. A holder that loses the lease during an upload does not publish that snapshot over a newer one from the instance that took over.
Each instance is identified by the SPICE_INSTANCE_ID environment variable, or by its host name when the variable is unset, so a restarted instance resumes the lease it held. Give each replica a distinct SPICE_INSTANCE_ID: two processes with the same identity both upload, and each logs a warning naming the identity.
Snapshot location and the Cayenne data tier​
Keep location on standard object storage: Amazon S3 (s3://), Google Cloud Storage (gs://), or Azure ADLS Gen2 (abfss://, abfs://). A local file:// folder is valid on a single machine. Standard cloud storage is what bucket and object replication, and readers in another region, use.
location is independent of Cayenne's S3 Express One Zone data tier. cayenne_file_path and the cayenne_s3_* parameters (cayenne_s3_region, cayenne_s3_zone_ids, cayenne_s3_auth, and the related keys) store Cayenne Vortex files on an Express One Zone directory bucket. That bucket is single-zone storage for the accelerator's data files. It does not substitute for the snapshot bucket, and it does not provide the replication or multi-region reads a standard snapshot location does. See S3 Express One Zone storage.
The Cayenne cold tier uses a general-purpose S3 or S3-compatible bucket, set by cayenne_datalake_location, not GCS or ADLS. That prefix is the cold data tier, separate from snapshots.location.
Failure behavior​
bootstrap_on_failure_behavior controls what Spice does when it cannot load the most recent snapshot.
retry– keep retrying the newest snapshot until it succeeds.fallback– try older snapshot files until one loads successfully.warn– log a warning and continue with an empty acceleration. (Default.)
Enable snapshots per dataset​
Each dataset opts into snapshotting through the acceleration.snapshots field. Four modes are available:
enabled– download snapshots on startup and write a new snapshot after each refresh.bootstrap_only– only download snapshots; never write new ones.create_only– write new snapshots after refreshes, but never download them on startup.disabled– disable snapshot usage for this dataset. (Default.)
Complete configuration:
acceleration:
snapshots: enabled | disabled # default: disabled
snapshots_trigger: <trigger_mode> # see trigger modes below
snapshots_trigger_threshold: <value> # threshold for time_interval or stream_batches
snapshots_compaction: enabled | disabled # default: disabled (DuckDB only)
snapshots_reset_expiry_on_load: enabled | disabled # default: disabled (DuckDB only with Caching refresh mode)
Snapshot triggers​
The snapshots_trigger setting controls when Spice creates new snapshots. The available triggers depend on the dataset's refresh mode.
Batch-based datasets​
Datasets using refresh_mode: full, or refresh_mode: append with a time_column, support the following triggers:
| Trigger | Description |
|---|---|
refresh_complete | Create a snapshot after each data refresh completes. (Default.) |
time_interval | Create snapshots at a fixed time interval. |
Example with default trigger:
datasets:
- from: s3://some_bucket/some_table/
name: some_table
acceleration:
enabled: true
engine: duckdb
mode: file
snapshots: enabled
# snapshots_trigger defaults to refresh_complete
params:
duckdb_file: /nvme/some_table.db
Example with time-based trigger:
datasets:
- from: s3://some_bucket/some_table/
name: some_table
acceleration:
enabled: true
engine: duckdb
mode: file
snapshots: enabled
snapshots_trigger: time_interval
snapshots_trigger_threshold: 30m
params:
duckdb_file: /nvme/some_table.db
Caching datasets​
Datasets using refresh_mode: caching support a single trigger:
| Trigger | Description |
|---|---|
time_interval | Create snapshots at a fixed time interval. (Default, and the only supported trigger.) |
The interval is set with snapshots_trigger_threshold and defaults to 10m. Setting snapshots_trigger to refresh_complete or stream_batches on a caching dataset is a configuration error.
Stream-based datasets​
Datasets using refresh_mode: changes, or refresh_mode: append without a time_column, support the following triggers:
| Trigger | Description |
|---|---|
time_interval | Create snapshots at a fixed time interval. (Default: 10m.) |
stream_batches | Create a snapshot after a specified number of batches are processed. |
Example with time-based trigger (default):
datasets:
- from: debezium:cdc_source
name: cdc_table
acceleration:
enabled: true
engine: duckdb
mode: file
refresh_mode: changes
snapshots: enabled
# snapshots_trigger defaults to time_interval
# snapshots_trigger_threshold defaults to 10m
params:
duckdb_file: /nvme/cdc_table.db
Example with batch-based trigger:
datasets:
- from: debezium:cdc_source
name: cdc_table
acceleration:
enabled: true
engine: duckdb
mode: file
refresh_mode: changes
snapshots: enabled
snapshots_trigger: stream_batches
snapshots_trigger_threshold: 300
params:
duckdb_file: /nvme/cdc_table.db
Snapshot compaction​
For DuckDB-based accelerations, enable snapshots_compaction to compact the database before uploading. This uses DuckDB's internal mechanism (COPY DATABASE) to reduce file size and improve read performance.
acceleration:
enabled: true
engine: duckdb
mode: file
snapshots: enabled
snapshots_compaction: enabled
params:
duckdb_file: /nvme/some_table.db
Compaction is only available for the DuckDB acceleration engine.
Snapshot Resetting Expiry on Load​
When using Caching refresh mode with DuckDB-based acceleration, you can enable snapshots_reset_expiry_on_load to extend the data's expiry to now() + TTL each time a snapshot is loaded.
acceleration:
enabled: true
engine: duckdb
mode: file
refresh_mode: caching
snapshots: enabled
snapshots_reset_expiry_on_load: enabled
params:
caching_ttl: 1m
caching_stale_while_revalidate_ttl: 1m
Complete example​
snapshots:
enabled: true
location: s3://some_bucket/some_folder/
bootstrap_on_failure_behavior: warn
params:
s3_auth: iam_role
datasets:
# Batch dataset with refresh-triggered snapshots
- from: s3://some_bucket/batch_table/
name: batch_table
params:
file_format: parquet
s3_auth: iam_role
acceleration:
enabled: true
engine: duckdb
mode: file
snapshots: enabled
snapshots_trigger: refresh_complete
snapshots_compaction: enabled
params:
duckdb_file: /nvme/batch_table.db
# Stream dataset with time-interval snapshots
- from: debezium:cdc_source
name: stream_table
acceleration:
enabled: true
engine: duckdb
mode: file
refresh_mode: changes
snapshots: enabled
snapshots_trigger: time_interval
snapshots_trigger_threshold: 5m
params:
duckdb_file: /nvme/stream_table.db
Append-mode accelerations that define a time_column wait to report ready until the first append refresh completes after snapshot bootstrap. This keeps the dataset out of rotation until the freshest data is available while still benefiting from the snapshot-assisted startup. See Fast Cold Starts for additional context.
Serve a dataset from published snapshots​
A Spice instance can serve the snapshots another instance publishes without any connection to the original source. Set from to the S3 prefix that holds the snapshots' metadata.json and set file_format: snapshot:
datasets:
- from: s3://some_bucket/some_folder/ # The prefix that holds metadata.json
name: some_table # Selects the some_table entry in metadata.json
params:
file_format: snapshot
s3_region: us-east-1
s3_auth: iam_role
The dataset name selects the entry in metadata.json, so it must match the name of the dataset that publishes the snapshots. Spice reads metadata.json, detects the engine that created the dataset's current snapshot (Cayenne, DuckDB, SQLite, or Turso), restores that snapshot into the same engine, and serves queries from it. It then checks for newer snapshots and swaps each one in, as the snapshot refresh mode does.
A snapshot dataset needs no acceleration block, no refresh_mode, no acceleration.snapshots setting, and no top-level snapshots section. The snapshot location comes from the dataset's own from and s3_* params. An optional acceleration block can set refresh_check_interval (default 1m) and engine params.
Loading and readiness​
The dataset reports Ready only after the current process has restored a snapshot. A local copy left from an earlier run is never served as current. Until the first snapshot is published, the dataset reports an error status and Spice logs a warning such as Dataset 'some_table' has no snapshot to load yet, so it cannot be queried until one is published. Spice keeps checking, with a backoff capped at refresh_check_interval. A query that reaches the dataset before a snapshot loads returns an error, not an empty result.
Spice keeps the local copy under .spice/data/. DuckDB, SQLite, and Turso copies are named for the dataset and a hash of the from location, so a dataset pointed at a new location never reopens the previous location's copy. Cayenne keeps its own layout: a data directory per dataset and one shared catalog, both under .spice/data/ by default. Each parameter moves its own part. cayenne_file_path moves the data directory, and when it is a local path the catalog follows it to {cayenne_file_path}/metadata. cayenne_metadata_dir moves the catalog alone, and the data directory keeps its default. Setting both lets a reader keep its copy on a chosen volume, in the same layout as the writer:
datasets:
- from: s3://some_bucket/some_folder/
name: some_table
params:
file_format: snapshot
s3_region: us-east-1
s3_auth: iam_role
acceleration:
params:
cayenne_file_path: /data/some_table/
cayenne_metadata_dir: /data/metadata/
Snapshot datasets follow the same metastore location rule as other Cayenne datasets: datasets that set different cayenne_file_path values must all set the same cayenne_metadata_dir, or the Spicepod is rejected at startup.
Configuration constraints​
The dataset is read-only. Spice rejects a configuration that contradicts reading snapshots, with an error that names the setting to remove:
frommust be ans3://location. Other connectors are not supported.paramsaccepts onlyfile_format,s3_region,s3_endpoint,s3_auth,s3_key,s3_secret,s3_session_token,client_timeout, andallow_http.s3_authmust beiam_roleorkey.accessmust beread(the default).embeddings,vectors, andfull_text_searchare not supported. Configure them on the dataset that publishes the snapshots.- In the
accelerationblock, Spice rejectsenabled: false, anyrefresh_modeother thansnapshot,mode: file_createormode: file_update,snapshots: enabledorsnapshots: create_only,refresh_sql, theretention_*settings,on_zero_results: use_source, and anengineother than the one that created the snapshots. - Engine path params (
duckdb_file,duckdb_data_dir,sqlite_file,turso_file, andcayenne_s3_zone_ids) are rejected, because Spice chooses where the local copy lives.cayenne_file_pathandcayenne_metadata_dirare accepted when the snapshots were created by Cayenne, and rejected by name when another engine created them.
HTTP API behavior​
GET /v1/datasets/{name}/acceleration/snapshots lists the snapshots at the dataset's own location. POST /v1/datasets/{name}/acceleration/snapshots/current returns 400 Bad Request, because a reader never changes the metadata it reads; set the current snapshot on the instance that publishes the snapshots. POST /v1/datasets/{name}/acceleration/refresh checks for a newer snapshot.
Best practices​
- Budget storage for a full file on every write. Each snapshot is a complete copy of the accelerated dataset. See What each snapshot contains and Snapshot location and the Cayenne data tier.
- Treat the snapshot interval as a freshness gap. A reader that bootstraps from object storage serves the last successful snapshot until its own next refresh.
refresh_completeis as fresh as the writer's last refresh;time_intervalcan be older still. Size the trigger against the freshness the readers are allowed to serve, and keep a durable volume when that gap is too wide. See Read/Write Separation. - Pair with ephemeral storage: Deployments commonly place the acceleration file on fast ephemeral disks (such as NVMe instance storage) while relying on snapshots for persistence across restarts. Local NVMe is the recommended medium for accelerations — see Storage for the tiers, the instance-store lifetime, and the capacity figures.
- Enable compaction for large datasets: Use
snapshots_compaction: enabledfor DuckDB accelerations to reduce snapshot size and improve bootstrap performance. - Tune trigger thresholds for stream datasets: For high-throughput streaming datasets, balance snapshot frequency against I/O overhead by adjusting
snapshots_trigger_threshold. - Align retention policies: Apply an object storage lifecycle rule that mirrors the desired snapshot retention policy.
- Monitor bootstraps: Track warning logs emitted when Spice falls back to an empty acceleration so operators can respond quickly if snapshot loading fails.
- Search indexes are not all inside the accelerator file. A DuckDB HNSW index lives in the DuckDB file, so it is part of that dataset's snapshot. A built-in full-text index does not: the default in-memory Tantivy index is rebuilt on every start, and
index_store: filewrites a separate directory (.spice/data/fts/...unlessindex_directoryis set) that acceleration snapshots do not upload. For restart parity of full-text search, keep that directory on the same durable volume as the acceleration file, or build the index once on a central tier and serve sidecars from the search results cache. See Where indexes are built.
For the full reference, see snapshots in the Spicepod specification and acceleration.snapshots.
For the production deployment pattern that uses snapshots to separate ingest from read workloads, see Read/Write Separation.
- Only datasets are supported for snapshots. Views are not supported.
- Partitioned Cayenne accelerations (
engine: cayennewithpartition_by) skip periodic and pre-recreate snapshots, with a warning naming the dataset. - Cayenne with a cold tier (
cayenne_datalake_locationset) neither creates nor bootstraps from snapshots. The dataset loads from its source, with a warning naming the dataset. See Cold Object-Store Tier.
Cookbook​
- A cookbook recipe to configure snapshots for file-mode accelerations so datasets avoid cold starts and recover quickly after restarts. Accelerated Snapshots
