Skip to main content
Version: Next

YAML syntax for Spicepod manifests

Spicepod manifests use YAML syntax. By default they are stored in the root directory of the application and named spicepod.yaml or spicepod.yml — when the runtime is given a directory path, it looks for one of those two filenames. The runtime also accepts any YAML file path directly (e.g. spiced ./configs/my-app.yaml); files passed by path must declare kind: Spicepod and a recognized version and are otherwise rejected with an explicit error.

Tip

Readers who are new to YAML can find a primer in "Learn YAML in Y minutes."

version​

The version of the Spicepod manifest. The current version is v2.

kind​

The kind of Spicepod manifest. The kind is Spicepod.

name​

The name of the Spicepod.

secrets​

The secrets section in the Spicepod manifest is optional and is used to configure how secrets are stored and accessed by the Spicepod. For more information, see Secret Stores.

secrets.from​

The from field is a string that represents the Uniform Resource Identifier (URI) for the secret store. This URI is composed of two parts: a prefix indicating the Secret Store to use, and an optional selector that specifies the secret to retrieve.

The syntax for the from field is as follows:

from: <secret_store>:<selector>

Where:

  • <secret_store>: The Secret Store to use

    Currently supported secret stores:

    If no secret stores are explicitly specified, it defaults to env.

  • <selector>: The secret within the secret store to load.

The type of secret store for reading secrets.

Example

secrets:
- from: env
name: env

secrets.name​

The name of the secret store. This is used to reference the store in the secret replacement syntax, ${<secret_store_name>:<key_name>}.

runtime​

The runtime section specifies configuration settings for the Spice runtime. For detailed documentation, see the Runtime YAML reference.

metadata​

An optional map of metadata.

Example

metadata:
epoch_time: 1605312000
period: 72h
interval: 1m
granularity: 10s
episodes: 10

datasets​

A Spicepod can contain one or more datasets referenced by relative path.

Example

A datasets referenced by relative path.

datasets:
- ref: datasets/uniswap_v2_eth_usdc

A dataset defined inline.

datasets:
- from: spice.ai/spiceai/quickstart/datasets/taxi_trips
name: taxi_trips
acceleration:
enabled: true
refresh_mode: full
refresh_check_interval: 1h

snapshots​

Optional. Configure managed acceleration snapshots that Spice can use to bootstrap file-based accelerations. When enabled, datasets that opt in with acceleration.snapshots will download database files from the snapshot location if the local file is missing, and will optionally write new snapshots after each refresh. DuckDB, SQLite, Cayenne, and Turso accelerations running in mode: file are supported, and each dataset must write to its own file path.

snapshots:
enabled: true
location: s3://my_bucket/snapshots/
bootstrap_on_failure_behavior: warn # warn | retry | fallback
params:
s3_auth: iam_role

snapshots.enabled​

Enable or disable snapshot management globally. Defaults to true.

snapshots.location​

The folder where snapshots are stored. Supports S3 bucket URIs (s3://bucket/prefix/), Azure ADLS Gen2 URIs (abfss://[email protected]/path/), Google Cloud Storage URIs (gs://bucket/prefix/), and local filesystem URIs (file:///path/to/folder/). The value must be a URI with a scheme: a bare filesystem path such as /var/spice/snapshots is not a valid URI, so it fails to parse and snapshots are disabled with an error logged. The location must resolve to a single folder; Spice creates per-dataset folders underneath using Hive-style partitions (month=YYYY-MM/day=YYYY-MM-DD/dataset=<name>).

The location's store must enforce conditional writes (create only if absent, update only at the version read), because every publisher updates the location's shared metadata.json with them. Amazon S3, Azure ADLS Gen2, Google Cloud Storage, and local file:// directories do. Each Spice instance probes the store before it creates a dataset's first snapshot, and refuses a store that does not enforce them. The probe creates, updates, and then deletes a .spice-conditional-write-probe-<id> object under the location. The credentials of every instance that publishes snapshots need permission to create and update that object. Permission to delete it is needed only for cleanup: a failed delete does not change the probe result and leaves the probe object behind. See Conditional writes on the snapshot location.

Keep a cloud location on standard S3, GCS, or ADLS so bucket or object replication and readers in another region can use the prefix. Each snapshot written there is a complete copy of the acceleration file. This location is separate from Cayenne's S3 Express One Zone data tier (cayenne_file_path, cayenne_s3_*), which does not substitute for the snapshot bucket. See Snapshots.

snapshots.bootstrap_on_failure_behavior​

Controls what happens when Spice cannot load the most recent snapshot on startup. Valid values:

  • warn (default) – Log a warning and continue with an empty acceleration.
  • retry – Retry the newest snapshot until it loads successfully.
  • fallback – Attempt older snapshots in the same dataset folder until one works.

snapshots.params​

Optional key-value map passed to the snapshot storage layer. When location points to S3, params accepts these S3 parameters: s3_region, s3_endpoint, s3_auth (iam_role or key; default iam_role), s3_key, s3_secret, s3_session_token, s3_queue_url, client_timeout, and allow_http. Spice ignores other keys and logs a warning for each. Snapshots default to s3_auth: iam_role, which differs from the S3 dataset default of public. With s3_auth: key, both s3_key and s3_secret must be set, or the snapshot location is not used and Spice logs an error. Azure ADLS and GCS locations also accept their respective connector parameters for explicit credential overrides; when no overrides are supplied, Spice reads standard environment variables for each cloud provider. Values can reference secrets with ${secrets:<name>} for all three. An invalid S3 parameter value, such as s3_auth: public, disables snapshots for the dataset with an error; Spice does not fall back to default parameters. When s3_queue_url is set, the dataset fails to register with the parameter error instead.

models​

A Spicepod can contain one or more models referenced by relative path.

Example

A model referenced by path.

models:
- from: models/drive_stats

A model defined inline.

models:
- from: spice.ai:openai/gpt-4o
name: cloud_llm
params:
spiceai_api_key: ${secrets:SPICEAI_API_KEY}

embeddings​

A Spicepod can contain one or more embeddings referenced by relative path.

Example

An embeddings model referenced by path.

embeddings:
- from: embeddings/openai_text_embedding_3

An embedding defined inline.

embeddings:
- name: hf_baai_bge
from: huggingface:huggingface.co/BAAI/bge-small-en-v1.5

dependencies​

A list of dependent Spicepods.

dependencies:
- lukekim/demo
- spicehq/nfts

views​

A Spicepod can contain one or more views which are virtual tables defined by SQL queries.

Example

views:
- name: rankings
sql: |
WITH a AS (
SELECT products.id, SUM(count) AS count
FROM orders
INNER JOIN products ON orders.product_id = products.id
GROUP BY products.id
)
SELECT name, count
FROM products
LEFT JOIN a ON products.id = a.id
ORDER BY count DESC
LIMIT 5

workers​

A Spicepod can contain one or more workers defining configurable units of compute.

Example

workers:
- name: round-robin
description: |
Distributes requests between 'llama3_2' and 'gpt4_1' models in a round-robin fashion.
load_balance:
routing:
- from: llama3_2
- from: gpt4_1
- name: fallback
description: |
Attempts 'gpt4_1' first, then 'llama3_2', then 'anth_haiku' if previous models fail.
load_balance:
routing:
- from: llama3_2
order: 2
- from: gpt4_1
order: 1
- from: anth_haiku
order: 3
- name: weighted
description: |
Routes 80% of traffic to 'llama3_2'.
load_balance:
routing:
- from: llama3_2
weight: 4
- from: gpt4_1
weight: 1

For a complete specification of worker configuration, see the Workers Reference.